npm · crates.io · PyPI93Exceptionalhealth index
Rust★ 12.2K↓ 729.9K/moAug 22, 2026
crates.io · PyPI · npm +593Exceptionalhealth index

yfedoseev/pdf_oxideThe fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1,015↓ 391.7K/moSep 5, 2026
npm91Excellenthealth index
PT-Perkasa-Pilar-Utama/ppu-paddle-ocrLightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/moAug 1, 2026
PyPI91Excellenthealth index

adbar/trafilaturaPython & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/moAug 28, 2026
npm · Packagist · crates.io +190Excellenthealth index

xberg-io/xbergPolyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9,228↓ 1,625/moAug 28, 2026
PyPI87Excellenthealth index

chrismattmann/tika-pythonTika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.
Python★ 1,666Aug 13, 2026
Packagist · crates.io · npm +187Excellenthealth index
xberg-io/html-to-markdownHigh performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 56+ document formats using streaming parsers and built-in OCR.
HTML · Rust★ 814↓ 11/moJul 27, 2026
crates.io · PyPI · npm +187Excellenthealth index

yfedoseev/office_oxideThe fastest Office document library for Python, Rust, Go, JS/TS, C# and WASM. DOCX, XLSX, PPTX, DOC, XLS, PPT. Up to 100× faster than python-docx/openpyxl/python-pptx. 100% pass rate on valid Office files. MIT/Apache-2.0.
Rust★ 108↓ 161.1K/moAug 22, 2026
Packagist86Excellenthealth index
PrinsFrank/pdfparserPHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3,925/moJul 15, 2026
PyPI · npm · crates.io86Excellenthealth index

firecrawl/pdf-inspectorFast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/moAug 22, 2026
npm86Excellenthealth index

harshankur/officeParserA robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/moAug 27, 2026
crates.io84Excellenthealth index

bzsanti/oxidizePdfPure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/moAug 22, 2026
PyPI80Excellenthealth index
airmang/python-hwpxPure Python HWPX automation: read, edit, generate, and validate documents without Hancom Office.
Python★ 96↓ 75.8K/moJul 20, 2026
Go80Excellenthealth index

coregx/gxpdfGxPDF - Enterprise-grade PDF library for Go. Table extraction, text parsing, encryption, document creation.
Go★ 47Aug 7, 2026

unjs/unpdf📄 PDF extraction and rendering across all JavaScript runtimes
TypeScript · JavaScript★ 1,204↓ 8.1M/moAug 5, 2026
C++ · C · TypeScript★ 3↓ 8,650/moJul 17, 2026
TypeScript★ 0↓ 3,363/moJul 15, 2026
crates.io69Goodhealth index
iyulab/unpdfHigh-performance Rust PDF extraction library with Markdown/JSON output, CJK/RTL support, multi-column layout detection, and Python/.NET/CLI bindings.
Rust★ 43↓ 2,335/moJul 21, 2026
TypeScript★ 0↓ 6,093/moJul 23, 2026
Packagist · Go · crates.io +267Goodhealth index
Rust · HTML★ 10↓ 1/moAug 22, 2026
crates.io · PyPI63Moderatehealth index
Rust · Python★ 2↓ 23.8K/moAug 28, 2026
npm63Moderatehealth index
TypeScript★ 0↓ 4,479/moJul 15, 2026
npm63Moderatehealth index
TypeScript · JavaScript★ 97↓ 50.6K/moAug 5, 2026

giraffesyo/pdfRobust, zero-dependency PDF text extraction for Go, with positioned glyphs, reading- order reconstruction, and hardened parsing.
Go★ 0Sep 5, 2026
Packagist57Moderatehealth index
PHP · HTML★ 9↓ 12.2K/moJul 17, 2026
PyPI54Moderatehealth index
Python★ 824↓ 12.3M/moAug 27, 2026
PyPI · crates.io51Moderatehealth index
4thel00z/pdfbossFrom-scratch Rust PDF toolkit for Python: fast page rendering and text extraction via PyO3 — benchmarked faster than mainstream Python PDF libraries
Rust★ 0Jul 31, 2026

cdown/srtA simple library and set of tools for parsing, modifying, and composing SRT files.
Python★ 536Aug 19, 2026