All tags
Catalogue tag

#pdf-extractor-pretrain

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

1 record
Tagged “pdf-extractor-pretrain”Ranked by health index
PyPI
75Goodhealth index
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 74.4K↓ 344K/moJul 13, 2026
Custom licenseJul 13, 2026 · metrics 1.13.0