Todas las etiquetas
Etiqueta del catálogo

#text-extraction

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

28 registros
Con la etiqueta «text-extraction»Ordenado por índice de salud
npm · crates.io · PyPI
93Excepcionalíndice de salud
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 12.2K↓ 729.9K/mes22 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
crates.io · PyPI · npm +5
93Excepcionalíndice de salud
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1015↓ 391.7K/mes5 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
npm
91Excelenteíndice de salud
PT-Perkasa-Pilar-Utama/ppu-paddle-ocr
Lightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
PyPI
91Excelenteíndice de salud
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6718↓ 14M/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
npm · Packagist · crates.io +1
90Excelenteíndice de salud
xberg-io/xberg
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9228↓ 1625/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
PyPI
87Excelenteíndice de salud
chrismattmann/tika-python
Tika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.
Python★ 166613 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
Packagist · crates.io · npm +1
87Excelenteíndice de salud
xberg-io/html-to-markdown
High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 56+ document formats using streaming parsers and built-in OCR.
HTML · Rust★ 814↓ 11/mes27 jul 2026
MIT27 jul 2026 · métricas 2.10.0
crates.io · PyPI · npm +1
87Excelenteíndice de salud
yfedoseev/office_oxide
The fastest Office document library for Python, Rust, Go, JS/TS, C# and WASM. DOCX, XLSX, PPTX, DOC, XLS, PPT. Up to 100× faster than python-docx/openpyxl/python-pptx. 100% pass rate on valid Office files. MIT/Apache-2.0.
Rust★ 108↓ 161.1K/mes22 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
Packagist
86Excelenteíndice de salud
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3925/mes15 jul 2026
MIT15 jul 2026 · métricas 2.10.0
PyPI · npm · crates.io
86Excelenteíndice de salud
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm
86Excelenteíndice de salud
harshankur/officeParser
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
crates.io
84Excelenteíndice de salud
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
PyPI
80Excelenteíndice de salud
airmang/python-hwpx
Pure Python HWPX automation: read, edit, generate, and validate documents without Hancom Office.
Python★ 96↓ 75.8K/mes20 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
Go
80Excelenteíndice de salud
coregx/gxpdf
GxPDF - Enterprise-grade PDF library for Go. Table extraction, text parsing, encryption, document creation.
Go★ 477 ago 2026
MIT7 ago 2026 · métricas 2.10.0
npm
78Buenoíndice de salud
unjs/unpdf
📄 PDF extraction and rendering across all JavaScript runtimes
TypeScript · JavaScript★ 1204↓ 8.1M/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
npm
75Buenoíndice de salud
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8650/mes17 jul 2026
MIT17 jul 2026 · métricas 2.10.0
npm
71Buenoíndice de salud
happyvertical/ocr
El repositorio no publica descripción.
TypeScript★ 0↓ 3363/mes15 jul 2026
MIT15 jul 2026 · métricas 2.10.0
crates.io
69Buenoíndice de salud
iyulab/unpdf
High-performance Rust PDF extraction library with Markdown/JSON output, CJK/RTL support, multi-column layout detection, and Python/.NET/CLI bindings.
Rust★ 43↓ 2335/mes21 jul 2026
MIT21 jul 2026 · métricas 2.10.0
npm
67Buenoíndice de salud
happyvertical/pdf
El repositorio no publica descripción.
TypeScript★ 0↓ 6093/mes23 jul 2026
MIT23 jul 2026 · métricas 2.10.0
Packagist · Go · crates.io +2
67Buenoíndice de salud
kreuzberg-dev/kreuzberg-lts
Kreuzberg v4 LTS — long-term support for the v4 line (legacy; superseded by xberg for v5+). MIT-licensed.
Rust · HTML★ 10↓ 1/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
crates.io · PyPI
63Moderadoíndice de salud
developer0hye/pdfplumber-rs
Evidence-driven PDF extraction for Rust, with an alpha Python pdfplumber migration path.
Rust · Python★ 2↓ 23.8K/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
npm
63Moderadoíndice de salud
heey-global/deep-ocr-n8n
El repositorio no publica descripción.
TypeScript★ 0↓ 4479/mes15 jul 2026
MIT15 jul 2026 · métricas 2.10.0
npm
63Moderadoíndice de salud
johannschopplich/pdfjs-serverless
🪭 PDF.js redistributed as a single bundle for edge and serverless runtimes
TypeScript · JavaScript★ 97↓ 50.6K/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
Go
60Moderadoíndice de salud
giraffesyo/pdf
Robust, zero-dependency PDF text extraction for Go, with positioned glyphs, reading- order reconstruction, and hardened parsing.
Go★ 05 sept 2026
MIT5 sept 2026 · métricas 2.10.0
Packagist
57Moderadoíndice de salud
TYPO3-Solr/ext-tika
A TYPO3 CMS extension that provides Apache Tika functionality
PHP · HTML★ 9↓ 12.2K/mes17 jul 2026
GPL-3.017 jul 2026 · métricas 2.10.0
PyPI
54Moderadoíndice de salud
miso-belica/jusText
Heuristic based boilerplate removal tool
Python★ 824↓ 12.3M/mes27 ago 2026
BSD-2-Clause27 ago 2026 · métricas 2.10.0
PyPI · crates.io
51Moderadoíndice de salud
4thel00z/pdfboss
From-scratch Rust PDF toolkit for Python: fast page rendering and text extraction via PyO3 — benchmarked faster than mainstream Python PDF libraries
Rust★ 031 jul 2026
Apache-2.031 jul 2026 · métricas 2.10.0
PyPI
25En riesgoíndice de salud
cdown/srt
A simple library and set of tools for parsing, modifying, and composing SRT files.
Python★ 53619 ago 2026
MIT19 ago 2026 · métricas 2.10.0