All tags
Catalogue tag

#pdf-parser

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

20 records
Tagged “pdf-parser”Ranked by health index
PyPI
98Exceptionalhealth index
py-pdf/pypdf
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Python★ 10.2KAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
PyPI · npm
96Exceptionalhealth index
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87KAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
npm · Maven · PyPI
95Exceptionalhealth index
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Java · Python★ 28.2K↓ 54.5K/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
npm · crates.io · PyPI
93Exceptionalhealth index
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 12.2K↓ 729.9K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
crates.io · PyPI · npm +5
93Exceptionalhealth index
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1,015↓ 391.7K/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/moAug 4, 2026
Custom licenseAug 4, 2026 · metrics 2.10.0
Packagist
86Excellenthealth index
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3,925/moJul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
PyPI · npm · crates.io
86Excellenthealth index
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
crates.io
84Excellenthealth index
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
83Excellenthealth index
pdfminer/pdfminer.six
Community maintained fork of pdfminer - we fathom PDF
Python★ 7,015↓ 69.6M/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
codereverser/casparser
Parser for Consolidated Account Statements (CAS) generated from CAMS/Karvy/Kfintech
Python★ 222↓ 23.6K/moAug 30, 2026
MITAug 30, 2026 · metrics 2.10.0
Go
80Excellenthealth index
coregx/gxpdf
GxPDF - Enterprise-grade PDF library for Go. Table extraction, text parsing, encryption, document creation.
Go★ 47Aug 7, 2026
MITAug 7, 2026 · metrics 2.10.0
npm
75Goodhealth index
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8,650/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
Go
73Goodhealth index
CASParser/cas-parser-go
CAS Parser allows you to track Consolidated Account Statement (CAS PDF) portfolios from NSDL, CDSL, CAMS, KFintech - CAS Parser API Client - GO
Go★ 0Aug 2, 2026
Apache-2.0Aug 2, 2026 · metrics 2.10.0
Go
73Goodhealth index
nmdra/notebrain-cli
A local-first CLI that makes your Obsidian vault (or Markdown Notes) searchable and AI-Agents queryable with semantic search and hidden connections.
Go★ 27Aug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm
71Goodhealth index
OpenSourceAGI/qwksearch-research-agent
🧠💻 Reimagine the Internet as Self-Organizing Mind Map 🤖🔎 STREAM: Search with Top Result Extraction & Answer Model 📈📝 REASON Docs Writing Agent 🚜📜 Tractor the Text Extractor 🔤📊 SEEKTOPIC
TypeScript★ 73↓ 79.4K/moJul 26, 2026
Custom licenseJul 26, 2026 · metrics 2.10.0
crates.io
60Moderatehealth index
AndyCappDev/stet
A PDF rendering engine and PostScript Level 3 interpreter written in pure Rust.
Rust★ 11↓ 91.6K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
Go
60Moderatehealth index
giraffesyo/pdf
Robust, zero-dependency PDF text extraction for Go, with positioned glyphs, reading- order reconstruction, and hardened parsing.
Go★ 0Sep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI · crates.io
51Moderatehealth index
4thel00z/pdfboss
From-scratch Rust PDF toolkit for Python: fast page rendering and text extraction via PyO3 — benchmarked faster than mainstream Python PDF libraries
Rust★ 0Jul 31, 2026
Apache-2.0Jul 31, 2026 · metrics 2.10.0
PyPI
12Criticalhealth index
euske/pdfminer
Python PDF Parser (Not actively maintained). Check out pdfminer.six.
Python★ 5,274Aug 12, 2026
MITAug 12, 2026 · metrics 2.10.0