全部标签
目录标签

#document-parser

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

6 条记录
标签为“document-parser”按健康指数排序
PyPI
87优秀健康指数
Unstructured-IO/unstructured
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
HTML★ 15.2K↓ 5.1M/月2026年7月18日
Apache-2.02026年7月18日 · 指标 1.13.0
PyPI
84良好健康指数
docling-project/docling
Get your documents ready for gen AI
Python★ 63.1K↓ 7.4M/月2026年7月13日
MIT2026年7月13日 · 指标 1.13.0
PyPI
83良好健康指数
ds4sd/docling
Get your documents ready for gen AI
Python★ 63.4K2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
Packagist
73良好健康指数
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3,925/月2026年7月15日
MIT2026年7月15日 · 指标 1.13.0
npm
66中等健康指数
chrisryugj/kordoc
모두 파싱해버리겠다 — HWP3·HWP·HWPX·HWPML·PDF·XLS·XLSX·DOCX → Markdown. 신구대조·양식 자동 채우기·MCP 통합 (CLI + MCP Server) | Parse Korean documents (HWP3-5, HWPX, HWPML, PDF, Office) to Markdown — CLI + MCP Server with form-filler & diff
TypeScript★ 1,422↓ 39.2K/月2026年7月13日
MIT2026年7月13日 · 指标 1.13.0
npm
33存在风险健康指数
harshankur/officeParser
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 502↓ 2.1M/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0