All tags
Catalogue tag

#pdf-extractor-rag

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

2 records
Tagged “pdf-extractor-rag”Ranked by health index
PyPI · npm
96Exceptionalhealth index
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87KAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/moAug 4, 2026
Custom licenseAug 4, 2026 · metrics 2.10.0