All tags
Catalogue tag

#pdf-to-text

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

6 records
Tagged “pdf-to-text”Ranked by health index
PyPI
87Excellenthealth index
Unstructured-IO/unstructured
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
HTML★ 15.2K↓ 5.1M/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
84Goodhealth index
docling-project/docling
Get your documents ready for gen AI
Python★ 63.1K↓ 7.4M/moJul 13, 2026
MITJul 13, 2026 · metrics 1.13.0
PyPI
83Goodhealth index
ds4sd/docling
Get your documents ready for gen AI
Python★ 63.4KJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
npm
78Goodhealth index
Fdawgs/node-poppler
Asynchronous Node.js wrapper for the Poppler PDF rendering utilities
JavaScript★ 247↓ 324.5K/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI · crates.io · npm +5
78Goodhealth index
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 886↓ 226.3K/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
npm
64Moderatehealth index
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8,650/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0