Усі теги
Тег каталогу

#pdf-parser

Усі репозиторії публічного реєстру з цим тегом — із тем GitHub або ключових слів, які публікують їхні реєстри пакетів. Здоров'я вимірюється за тією ж версіонованою методологією, що й решта реєстру.

20 записів
З тегом «pdf-parser»Упорядковано за індексом здоров'я
PyPI
98Винятковийіндекс здоров'я
py-pdf/pypdf
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Python★ 10.2K28 серп. 2026 р.
Власна ліцензія28 серп. 2026 р. · метрики 2.10.0
PyPI · npm
96Винятковийіндекс здоров'я
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87K4 серп. 2026 р.
Apache-2.04 серп. 2026 р. · метрики 2.10.0
npm · Maven · PyPI
95Винятковийіндекс здоров'я
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Java · Python★ 28.2K↓ 54.5K/міс5 серп. 2026 р.
Apache-2.05 серп. 2026 р. · метрики 2.10.0
npm · crates.io · PyPI
93Винятковийіндекс здоров'я
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 12.2K↓ 729.9K/міс22 серп. 2026 р.
Apache-2.022 серп. 2026 р. · метрики 2.10.0
crates.io · PyPI · npm +5
93Винятковийіндекс здоров'я
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1 015↓ 391.7K/міс5 вер. 2026 р.
Apache-2.05 вер. 2026 р. · метрики 2.10.0
PyPI
89Відміннийіндекс здоров'я
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/міс4 серп. 2026 р.
Власна ліцензія4 серп. 2026 р. · метрики 2.10.0
Packagist
86Відміннийіндекс здоров'я
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3 925/міс15 лип. 2026 р.
MIT15 лип. 2026 р. · метрики 2.10.0
PyPI · npm · crates.io
86Відміннийіндекс здоров'я
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/міс22 серп. 2026 р.
MIT22 серп. 2026 р. · метрики 2.10.0
crates.io
84Відміннийіндекс здоров'я
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/міс22 серп. 2026 р.
MIT22 серп. 2026 р. · метрики 2.10.0
PyPI
83Відміннийіндекс здоров'я
pdfminer/pdfminer.six
Community maintained fork of pdfminer - we fathom PDF
Python★ 7 015↓ 69.6M/міс4 серп. 2026 р.
MIT4 серп. 2026 р. · метрики 2.10.0
PyPI
80Відміннийіндекс здоров'я
codereverser/casparser
Parser for Consolidated Account Statements (CAS) generated from CAMS/Karvy/Kfintech
Python★ 222↓ 23.6K/міс30 серп. 2026 р.
MIT30 серп. 2026 р. · метрики 2.10.0
Go
80Відміннийіндекс здоров'я
coregx/gxpdf
GxPDF - Enterprise-grade PDF library for Go. Table extraction, text parsing, encryption, document creation.
Go★ 477 серп. 2026 р.
MIT7 серп. 2026 р. · метрики 2.10.0
npm
75Добрийіндекс здоров'я
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8 650/міс17 лип. 2026 р.
MIT17 лип. 2026 р. · метрики 2.10.0
Go
73Добрийіндекс здоров'я
CASParser/cas-parser-go
CAS Parser allows you to track Consolidated Account Statement (CAS PDF) portfolios from NSDL, CDSL, CAMS, KFintech - CAS Parser API Client - GO
Go★ 02 серп. 2026 р.
Apache-2.02 серп. 2026 р. · метрики 2.10.0
Go
73Добрийіндекс здоров'я
nmdra/notebrain-cli
A local-first CLI that makes your Obsidian vault (or Markdown Notes) searchable and AI-Agents queryable with semantic search and hidden connections.
Go★ 2722 серп. 2026 р.
MIT22 серп. 2026 р. · метрики 2.10.0
npm
71Добрийіндекс здоров'я
OpenSourceAGI/qwksearch-research-agent
🧠💻 Reimagine the Internet as Self-Organizing Mind Map 🤖🔎 STREAM: Search with Top Result Extraction & Answer Model 📈📝 REASON Docs Writing Agent 🚜📜 Tractor the Text Extractor 🔤📊 SEEKTOPIC
TypeScript★ 73↓ 79.4K/міс26 лип. 2026 р.
Власна ліцензія26 лип. 2026 р. · метрики 2.10.0
crates.io
60Помірнийіндекс здоров'я
AndyCappDev/stet
A PDF rendering engine and PostScript Level 3 interpreter written in pure Rust.
Rust★ 11↓ 91.6K/міс19 серп. 2026 р.
Apache-2.019 серп. 2026 р. · метрики 2.10.0
Go
60Помірнийіндекс здоров'я
giraffesyo/pdf
Robust, zero-dependency PDF text extraction for Go, with positioned glyphs, reading- order reconstruction, and hardened parsing.
Go★ 05 вер. 2026 р.
MIT5 вер. 2026 р. · метрики 2.10.0
PyPI · crates.io
51Помірнийіндекс здоров'я
4thel00z/pdfboss
From-scratch Rust PDF toolkit for Python: fast page rendering and text extraction via PyO3 — benchmarked faster than mainstream Python PDF libraries
Rust★ 031 лип. 2026 р.
Apache-2.031 лип. 2026 р. · метрики 2.10.0
PyPI
12Критичнийіндекс здоров'я
euske/pdfminer
Python PDF Parser (Not actively maintained). Check out pdfminer.six.
Python★ 5 27412 серп. 2026 р.
MIT12 серп. 2026 р. · метрики 2.10.0