All tags
Catalogue tag

#text-extraction

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

28 records
Tagged “text-extraction”Ranked by health index
npm · crates.io · PyPI
93Exceptionalhealth index
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 12.2K↓ 729.9K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
crates.io · PyPI · npm +5
93Exceptionalhealth index
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1,015↓ 391.7K/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
npm
91Excellenthealth index
PT-Perkasa-Pilar-Utama/ppu-paddle-ocr
Lightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
npm · Packagist · crates.io +1
90Excellenthealth index
xberg-io/xberg
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9,228↓ 1,625/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
PyPI
87Excellenthealth index
chrismattmann/tika-python
Tika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.
Python★ 1,666Aug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
Packagist · crates.io · npm +1
87Excellenthealth index
xberg-io/html-to-markdown
High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 56+ document formats using streaming parsers and built-in OCR.
HTML · Rust★ 814↓ 11/moJul 27, 2026
MITJul 27, 2026 · metrics 2.10.0
crates.io · PyPI · npm +1
87Excellenthealth index
yfedoseev/office_oxide
The fastest Office document library for Python, Rust, Go, JS/TS, C# and WASM. DOCX, XLSX, PPTX, DOC, XLS, PPT. Up to 100× faster than python-docx/openpyxl/python-pptx. 100% pass rate on valid Office files. MIT/Apache-2.0.
Rust★ 108↓ 161.1K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
Packagist
86Excellenthealth index
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3,925/moJul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
PyPI · npm · crates.io
86Excellenthealth index
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm
86Excellenthealth index
harshankur/officeParser
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
crates.io
84Excellenthealth index
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
airmang/python-hwpx
Pure Python HWPX automation: read, edit, generate, and validate documents without Hancom Office.
Python★ 96↓ 75.8K/moJul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 2.10.0
Go
80Excellenthealth index
coregx/gxpdf
GxPDF - Enterprise-grade PDF library for Go. Table extraction, text parsing, encryption, document creation.
Go★ 47Aug 7, 2026
MITAug 7, 2026 · metrics 2.10.0
npm
78Goodhealth index
unjs/unpdf
📄 PDF extraction and rendering across all JavaScript runtimes
TypeScript · JavaScript★ 1,204↓ 8.1M/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
npm
75Goodhealth index
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8,650/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
npm
71Goodhealth index
happyvertical/ocr
No repository description published.
TypeScript★ 0↓ 3,363/moJul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
crates.io
69Goodhealth index
iyulab/unpdf
High-performance Rust PDF extraction library with Markdown/JSON output, CJK/RTL support, multi-column layout detection, and Python/.NET/CLI bindings.
Rust★ 43↓ 2,335/moJul 21, 2026
MITJul 21, 2026 · metrics 2.10.0
npm
67Goodhealth index
happyvertical/pdf
No repository description published.
TypeScript★ 0↓ 6,093/moJul 23, 2026
MITJul 23, 2026 · metrics 2.10.0
Packagist · Go · crates.io +2
67Goodhealth index
kreuzberg-dev/kreuzberg-lts
Kreuzberg v4 LTS — long-term support for the v4 line (legacy; superseded by xberg for v5+). MIT-licensed.
Rust · HTML★ 10↓ 1/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
crates.io · PyPI
63Moderatehealth index
developer0hye/pdfplumber-rs
Evidence-driven PDF extraction for Rust, with an alpha Python pdfplumber migration path.
Rust · Python★ 2↓ 23.8K/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
npm
63Moderatehealth index
heey-global/deep-ocr-n8n
No repository description published.
TypeScript★ 0↓ 4,479/moJul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
npm
63Moderatehealth index
johannschopplich/pdfjs-serverless
🪭 PDF.js redistributed as a single bundle for edge and serverless runtimes
TypeScript · JavaScript★ 97↓ 50.6K/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
Go
60Moderatehealth index
giraffesyo/pdf
Robust, zero-dependency PDF text extraction for Go, with positioned glyphs, reading- order reconstruction, and hardened parsing.
Go★ 0Sep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
Packagist
57Moderatehealth index
TYPO3-Solr/ext-tika
A TYPO3 CMS extension that provides Apache Tika functionality
PHP · HTML★ 9↓ 12.2K/moJul 17, 2026
GPL-3.0Jul 17, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
miso-belica/jusText
Heuristic based boilerplate removal tool
Python★ 824↓ 12.3M/moAug 27, 2026
BSD-2-ClauseAug 27, 2026 · metrics 2.10.0
PyPI · crates.io
51Moderatehealth index
4thel00z/pdfboss
From-scratch Rust PDF toolkit for Python: fast page rendering and text extraction via PyO3 — benchmarked faster than mainstream Python PDF libraries
Rust★ 0Jul 31, 2026
Apache-2.0Jul 31, 2026 · metrics 2.10.0
PyPI
25At Riskhealth index
cdown/srt
A simple library and set of tools for parsing, modifying, and composing SRT files.
Python★ 536Aug 19, 2026
MITAug 19, 2026 · metrics 2.10.0