PyPI89优秀健康指数opendatalab/MinerUTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.Python★ 76.8K↓ 313.3K/月2026年8月4日自定义许可证2026年8月4日 · 指标 2.10.0
PyPI83优秀健康指数pdfminer/pdfminer.sixCommunity maintained fork of pdfminer - we fathom PDFPython★ 7,015↓ 69.6M/月2026年8月4日MIT2026年8月4日 · 指标 2.10.0
NuGet80优秀健康指数UglyToad/PdfPigRead and extract text and other content from PDFs in C# (port of PDFBox)C#★ 2,5012026年7月17日Apache-2.02026年7月17日 · 指标 2.10.0
PyPI · npm73良好健康指数u9401066/asset-aware-mcpAsset-Aware MCP Server — AI Agent precisely accesses tables, figures, sections from PDFs + .docx round-trip editing (DFM) with 46 tools / 13 resources, segmentation export, layout overlay, OCR preprocessing, knowledge graph (LightRAG)Python · TypeScript★ 02026年7月22日Apache-2.02026年7月22日 · 指标 2.10.0
PyPI34存在风险健康指数Layout-Parser/layout-parserA Unified Toolkit for Deep Learning Based Document Image AnalysisPython★ 5,7712026年8月12日Apache-2.02026年8月12日 · 指标 2.10.0
PyPI12危急健康指数euske/pdfminerPython PDF Parser (Not actively maintained). Check out pdfminer.six.Python★ 5,2742026年8月12日MIT2026年8月12日 · 指标 2.10.0