Alle Tags
Katalog-Tag

#pdf-parser

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

20 Einträge
Getaggt als „pdf-parser“Geordnet nach Gesundheitsindex
PyPI
98AußergewöhnlichGesundheitsindex
py-pdf/pypdf
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Python★ 10.2K28. Aug. 2026
Eigene Lizenz28. Aug. 2026 · Metriken 2.10.0
PyPI · npm
96AußergewöhnlichGesundheitsindex
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87K4. Aug. 2026
Apache-2.04. Aug. 2026 · Metriken 2.10.0
npm · Maven · PyPI
95AußergewöhnlichGesundheitsindex
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Java · Python★ 28.2K↓ 54.5K/Monat5. Aug. 2026
Apache-2.05. Aug. 2026 · Metriken 2.10.0
npm · crates.io · PyPI
93AußergewöhnlichGesundheitsindex
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 12.2K↓ 729.9K/Monat22. Aug. 2026
Apache-2.022. Aug. 2026 · Metriken 2.10.0
crates.io · PyPI · npm +5
93AußergewöhnlichGesundheitsindex
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1.015↓ 391.7K/Monat5. Sept. 2026
Apache-2.05. Sept. 2026 · Metriken 2.10.0
PyPI
89ExzellentGesundheitsindex
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/Monat4. Aug. 2026
Eigene Lizenz4. Aug. 2026 · Metriken 2.10.0
Packagist
86ExzellentGesundheitsindex
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3.925/Monat15. Juli 2026
MIT15. Juli 2026 · Metriken 2.10.0
PyPI · npm · crates.io
86ExzellentGesundheitsindex
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/Monat22. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
crates.io
84ExzellentGesundheitsindex
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/Monat22. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
PyPI
83ExzellentGesundheitsindex
pdfminer/pdfminer.six
Community maintained fork of pdfminer - we fathom PDF
Python★ 7.015↓ 69.6M/Monat4. Aug. 2026
MIT4. Aug. 2026 · Metriken 2.10.0
PyPI
80ExzellentGesundheitsindex
codereverser/casparser
Parser for Consolidated Account Statements (CAS) generated from CAMS/Karvy/Kfintech
Python★ 222↓ 23.6K/Monat30. Aug. 2026
MIT30. Aug. 2026 · Metriken 2.10.0
Go
80ExzellentGesundheitsindex
coregx/gxpdf
GxPDF - Enterprise-grade PDF library for Go. Table extraction, text parsing, encryption, document creation.
Go★ 477. Aug. 2026
MIT7. Aug. 2026 · Metriken 2.10.0
npm
75GutGesundheitsindex
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8.650/Monat17. Juli 2026
MIT17. Juli 2026 · Metriken 2.10.0
Go
73GutGesundheitsindex
CASParser/cas-parser-go
CAS Parser allows you to track Consolidated Account Statement (CAS PDF) portfolios from NSDL, CDSL, CAMS, KFintech - CAS Parser API Client - GO
Go★ 02. Aug. 2026
Apache-2.02. Aug. 2026 · Metriken 2.10.0
Go
73GutGesundheitsindex
nmdra/notebrain-cli
A local-first CLI that makes your Obsidian vault (or Markdown Notes) searchable and AI-Agents queryable with semantic search and hidden connections.
Go★ 2722. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
npm
71GutGesundheitsindex
OpenSourceAGI/qwksearch-research-agent
🧠💻 Reimagine the Internet as Self-Organizing Mind Map 🤖🔎 STREAM: Search with Top Result Extraction & Answer Model 📈📝 REASON Docs Writing Agent 🚜📜 Tractor the Text Extractor 🔤📊 SEEKTOPIC
TypeScript★ 73↓ 79.4K/Monat26. Juli 2026
Eigene Lizenz26. Juli 2026 · Metriken 2.10.0
crates.io
60MittelGesundheitsindex
AndyCappDev/stet
A PDF rendering engine and PostScript Level 3 interpreter written in pure Rust.
Rust★ 11↓ 91.6K/Monat19. Aug. 2026
Apache-2.019. Aug. 2026 · Metriken 2.10.0
Go
60MittelGesundheitsindex
giraffesyo/pdf
Robust, zero-dependency PDF text extraction for Go, with positioned glyphs, reading- order reconstruction, and hardened parsing.
Go★ 05. Sept. 2026
MIT5. Sept. 2026 · Metriken 2.10.0
PyPI · crates.io
51MittelGesundheitsindex
4thel00z/pdfboss
From-scratch Rust PDF toolkit for Python: fast page rendering and text extraction via PyO3 — benchmarked faster than mainstream Python PDF libraries
Rust★ 031. Juli 2026
Apache-2.031. Juli 2026 · Metriken 2.10.0
PyPI
12KritischGesundheitsindex
euske/pdfminer
Python PDF Parser (Not actively maintained). Check out pdfminer.six.
Python★ 5.27412. Aug. 2026
MIT12. Aug. 2026 · Metriken 2.10.0