All tags
Catalogue tag

#pdf-parsing

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

5 records
Tagged “pdf-parsing”Ranked by health index
PyPI
98Exceptionalhealth index
py-pdf/pypdf
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Python★ 10.2KAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
npm · Maven · PyPI
95Exceptionalhealth index
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Java · Python★ 28.2K↓ 54.5K/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
crates.io · PyPI · npm +5
93Exceptionalhealth index
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1,015↓ 391.7K/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
npm
80Excellenthealth index
LibPDF-js/core
A modern PDF library for TypeScript. Parse, modify, and generate PDFs with a clean, intuitive API.
TypeScript★ 1,771↓ 352.6K/moJul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
PyPI
77Goodhealth index
jsvine/pdfplumber
Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
Python★ 10.7KAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0