PyPI97Exceptionalhealth index

Unstructured-IO/unstructuredConvert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
HTML★ 15.3KAug 5, 2026
PyPI · npm96Exceptionalhealth index

PaddlePaddle/PaddleOCRTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87KAug 4, 2026
npm · PyPI96Exceptionalhealth index
Python · TypeScript★ 44.1KAug 10, 2026
npm · Maven · PyPI95Exceptionalhealth index
Java · Python★ 28.2K↓ 54.5K/moAug 5, 2026
PyPI94Exceptionalhealth index

krkn-chaos/krknChaos and resiliency testing tool for Kubernetes with a focus on improving performance under failure conditions. A CNCF sandbox project.
Python★ 486↓ 16.5K/moAug 18, 2026
PyPI93Exceptionalhealth index

RapidAI/RapidOCR📄 Awesome OCR multiple programing languages toolkits based on ONNX Runtime, OpenVINO, MNN, PaddlePaddle, TensorRT and PyTorch.
Python · Jupyter Notebook★ 7,408↓ 3.8M/moAug 7, 2026
npm · crates.io · PyPI93Exceptionalhealth index
Rust★ 12.2K↓ 729.9K/moAug 22, 2026
npm · crates.io93Exceptionalhealth index

ruvnet/RuVectorRuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust · TypeScript★ 4,441↓ 432.2K/moAug 22, 2026
crates.io · PyPI · npm +593Exceptionalhealth index

yfedoseev/pdf_oxideThe fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1,015↓ 391.7K/moSep 5, 2026
PyPI92Excellenthealth index
Python★ 171.5K↓ 13.7M/moAug 4, 2026
PyPI92Excellenthealth index

mindee/doctrdocTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
Python★ 6,206↓ 326.1K/moAug 12, 2026
npm91Excellenthealth index
PT-Perkasa-Pilar-Utama/ppu-paddle-ocrLightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/moAug 1, 2026
PyPI91Excellenthealth index

ocrmypdf/OCRmyPDFOCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Python★ 34.4KAug 5, 2026
PyPI90Excellenthealth index

pymupdf/PyMuPDFPyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Python · SWIG★ 10.6KAug 28, 2026
npm · Packagist · crates.io +190Excellenthealth index

xberg-io/xbergPolyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9,228↓ 1,625/moAug 28, 2026
C · Rust★ 895Jul 31, 2026
NuGet89Excellenthealth index
C#★ 5Jul 22, 2026
PyPI89Excellenthealth index

opendatalab/MinerUTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/moAug 4, 2026
PyPI89Excellenthealth index
Python★ 57↓ 28.6M/moAug 8, 2026
Go89Excellenthealth index
y3owk1n/neruNavigate your entire screen without touching the mouse.
Go · Objective-C★ 465Jul 17, 2026
npm88Excellenthealth index
TypeScript★ 29↓ 248.7K/moAug 1, 2026
RubyGems87Excellenthealth index
Ruby★ 14Jul 15, 2026
PyPI87Excellenthealth index

robocorp/rpaframeworkCollection of open-source libraries and tools for Robotic Process Automation (RPA), designed to be used with both Robot Framework and Python
Python★ 1,545↓ 550.8K/moAug 13, 2026
PyPI · npm · crates.io86Excellenthealth index

firecrawl/pdf-inspectorFast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/moAug 22, 2026
npm86Excellenthealth index

harshankur/officeParserA robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/moAug 27, 2026
crates.io84Excellenthealth index

bzsanti/oxidizePdfPure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/moAug 22, 2026
PyPI84Excellenthealth index
Python★ 38.3K↓ 484.3K/moAug 5, 2026
PyPI83Excellenthealth index

datalab-to/suryaOCR, layout analysis, reading order, table recognition in 90+ languages
Python★ 21.2K↓ 968.8K/moAug 5, 2026
Go · npm83Excellenthealth index
icereed/paperless-gptUse LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
Go · TypeScript★ 2,596Aug 2, 2026
PyPI81Excellenthealth index
jztan/pdf-mcpAn MCP server that lets Claude Code and other AI agents work through large PDFs without overflowing their context — search by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts.
Python★ 91↓ 8,068/moAug 1, 2026