Todas las etiquetas
Etiqueta del catálogo

#pdf-parser

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

20 registros
Con la etiqueta «pdf-parser»Ordenado por índice de salud
PyPI
98Excepcionalíndice de salud
py-pdf/pypdf
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Python★ 10.2K28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI · npm
96Excepcionalíndice de salud
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87K4 ago 2026
Apache-2.04 ago 2026 · métricas 2.10.0
npm · Maven · PyPI
95Excepcionalíndice de salud
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Java · Python★ 28.2K↓ 54.5K/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
npm · crates.io · PyPI
93Excepcionalíndice de salud
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 12.2K↓ 729.9K/mes22 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
crates.io · PyPI · npm +5
93Excepcionalíndice de salud
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1015↓ 391.7K/mes5 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
PyPI
89Excelenteíndice de salud
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/mes4 ago 2026
Licencia propia4 ago 2026 · métricas 2.10.0
Packagist
86Excelenteíndice de salud
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3925/mes15 jul 2026
MIT15 jul 2026 · métricas 2.10.0
PyPI · npm · crates.io
86Excelenteíndice de salud
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
crates.io
84Excelenteíndice de salud
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
PyPI
83Excelenteíndice de salud
pdfminer/pdfminer.six
Community maintained fork of pdfminer - we fathom PDF
Python★ 7015↓ 69.6M/mes4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
PyPI
80Excelenteíndice de salud
codereverser/casparser
Parser for Consolidated Account Statements (CAS) generated from CAMS/Karvy/Kfintech
Python★ 222↓ 23.6K/mes30 ago 2026
MIT30 ago 2026 · métricas 2.10.0
Go
80Excelenteíndice de salud
coregx/gxpdf
GxPDF - Enterprise-grade PDF library for Go. Table extraction, text parsing, encryption, document creation.
Go★ 477 ago 2026
MIT7 ago 2026 · métricas 2.10.0
npm
75Buenoíndice de salud
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8650/mes17 jul 2026
MIT17 jul 2026 · métricas 2.10.0
Go
73Buenoíndice de salud
CASParser/cas-parser-go
CAS Parser allows you to track Consolidated Account Statement (CAS PDF) portfolios from NSDL, CDSL, CAMS, KFintech - CAS Parser API Client - GO
Go★ 02 ago 2026
Apache-2.02 ago 2026 · métricas 2.10.0
Go
73Buenoíndice de salud
nmdra/notebrain-cli
A local-first CLI that makes your Obsidian vault (or Markdown Notes) searchable and AI-Agents queryable with semantic search and hidden connections.
Go★ 2722 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm
71Buenoíndice de salud
OpenSourceAGI/qwksearch-research-agent
🧠💻 Reimagine the Internet as Self-Organizing Mind Map 🤖🔎 STREAM: Search with Top Result Extraction & Answer Model 📈📝 REASON Docs Writing Agent 🚜📜 Tractor the Text Extractor 🔤📊 SEEKTOPIC
TypeScript★ 73↓ 79.4K/mes26 jul 2026
Licencia propia26 jul 2026 · métricas 2.10.0
crates.io
60Moderadoíndice de salud
AndyCappDev/stet
A PDF rendering engine and PostScript Level 3 interpreter written in pure Rust.
Rust★ 11↓ 91.6K/mes19 ago 2026
Apache-2.019 ago 2026 · métricas 2.10.0
Go
60Moderadoíndice de salud
giraffesyo/pdf
Robust, zero-dependency PDF text extraction for Go, with positioned glyphs, reading- order reconstruction, and hardened parsing.
Go★ 05 sept 2026
MIT5 sept 2026 · métricas 2.10.0
PyPI · crates.io
51Moderadoíndice de salud
4thel00z/pdfboss
From-scratch Rust PDF toolkit for Python: fast page rendering and text extraction via PyO3 — benchmarked faster than mainstream Python PDF libraries
Rust★ 031 jul 2026
Apache-2.031 jul 2026 · métricas 2.10.0
PyPI
12Críticoíndice de salud
euske/pdfminer
Python PDF Parser (Not actively maintained). Check out pdfminer.six.
Python★ 527412 ago 2026
MIT12 ago 2026 · métricas 2.10.0