All tags
Catalogue tag

#pdf-parser

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

11 records
Tagged “pdf-parser”Ranked by health index
PyPI
87Excellenthealth index
py-pdf/pypdf
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Python★ 10.1KJul 21, 2026
Custom licenseJul 21, 2026 · metrics 1.13.0
npm · PyPI
82Goodhealth index
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python★ 85.3K↓ 2.6M/moJul 13, 2026
Apache-2.0Jul 13, 2026 · metrics 1.13.0
Maven · npm
79Goodhealth index
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Java · Python★ 27.3KJul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 1.13.0
PyPI · crates.io · npm +5
78Goodhealth index
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 886↓ 226.3K/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
PyPI
75Goodhealth index
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 74.4K↓ 344K/moJul 13, 2026
Custom licenseJul 13, 2026 · metrics 1.13.0
crates.io · PyPI
74Goodhealth index
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 11.5K↓ 0/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
Packagist
73Goodhealth index
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3,925/moJul 15, 2026
MITJul 15, 2026 · metrics 1.13.0
crates.io
69Moderatehealth index
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 182↓ 6,045/moJul 13, 2026
MITJul 13, 2026 · metrics 1.13.0
PyPI
68Moderatehealth index
pdfminer/pdfminer.six
Community maintained fork of pdfminer - we fathom PDF
Python★ 7,001↓ 66.6M/moJul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
npm
64Moderatehealth index
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8,650/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
crates.io · npm · PyPI
61Moderatehealth index
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 1,577↓ 36.1K/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0