Todas las etiquetas
Etiqueta del catálogo

#document-parser

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

8 registros
Con la etiqueta «document-parser»Ordenado por índice de salud
PyPI
98Excepcionalíndice de salud
docling-project/docling
Get your documents ready for gen AI
Python★ 64.2K4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
Unstructured-IO/unstructured
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
HTML★ 15.3K5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
crates.io · PyPI · npm +5
93Excepcionalíndice de salud
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1015↓ 391.7K/mes5 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
npm · PyPI
90Excelenteíndice de salud
Marker-Inc-Korea/AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4963↓ 309/mes2 ago 2026
Licencia propia2 ago 2026 · métricas 2.10.0
crates.io · PyPI · npm +1
87Excelenteíndice de salud
yfedoseev/office_oxide
The fastest Office document library for Python, Rust, Go, JS/TS, C# and WASM. DOCX, XLSX, PPTX, DOC, XLS, PPT. Up to 100× faster than python-docx/openpyxl/python-pptx. 100% pass rate on valid Office files. MIT/Apache-2.0.
Rust★ 108↓ 161.1K/mes22 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
Packagist
86Excelenteíndice de salud
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3925/mes15 jul 2026
MIT15 jul 2026 · métricas 2.10.0
npm
86Excelenteíndice de salud
harshankur/officeParser
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
npm
77Buenoíndice de salud
chrisryugj/kordoc
모두 파싱해버리겠다 — HWP·HWPX·PDF·Office 문서를 Markdown으로. 양식 자동 채우기와 신구대조를 갖춘 CLI·MCP 서버 | Convert Korean documents (HWP, HWPX, PDF, Office) to Markdown — CLI and MCP server with form filling and diff
TypeScript★ 1762↓ 65.4K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0