All tags
Catalogue tag

#pdf-extractor-rag

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

2 records
Tagged “pdf-extractor-rag”Ranked by health index
npm · PyPI
82Goodhealth index
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python★ 85.3K↓ 2.6M/moJul 13, 2026
Apache-2.0Jul 13, 2026 · metrics 1.13.0
PyPI
75Goodhealth index
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 74.4K↓ 344K/moJul 13, 2026
Custom licenseJul 13, 2026 · metrics 1.13.0