PyPI97Винятковийіндекс здоров'я

Unstructured-IO/unstructuredConvert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
HTML★ 15.3K5 серп. 2026 р.
PyPI · npm96Винятковийіндекс здоров'я

PaddlePaddle/PaddleOCRTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87K4 серп. 2026 р.
npm · PyPI96Винятковийіндекс здоров'я
Python · TypeScript★ 44.1K10 серп. 2026 р.
npm · Maven · PyPI95Винятковийіндекс здоров'я
Java · Python★ 28.2K↓ 54.5K/міс5 серп. 2026 р.
PyPI94Винятковийіндекс здоров'я

krkn-chaos/krknChaos and resiliency testing tool for Kubernetes with a focus on improving performance under failure conditions. A CNCF sandbox project.
Python★ 486↓ 16.5K/міс18 серп. 2026 р.
PyPI93Винятковийіндекс здоров'я

RapidAI/RapidOCR📄 Awesome OCR multiple programing languages toolkits based on ONNX Runtime, OpenVINO, MNN, PaddlePaddle, TensorRT and PyTorch.
Python · Jupyter Notebook★ 7 408↓ 3.8M/міс7 серп. 2026 р.
npm · crates.io · PyPI93Винятковийіндекс здоров'я
Rust★ 12.2K↓ 729.9K/міс22 серп. 2026 р.
npm · crates.io93Винятковийіндекс здоров'я

ruvnet/RuVectorRuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust · TypeScript★ 4 441↓ 432.2K/міс22 серп. 2026 р.
crates.io · PyPI · npm +593Винятковийіндекс здоров'я

yfedoseev/pdf_oxideThe fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1 015↓ 391.7K/міс5 вер. 2026 р.
PyPI92Відміннийіндекс здоров'я
Python★ 171.5K↓ 13.7M/міс4 серп. 2026 р.
PyPI92Відміннийіндекс здоров'я

mindee/doctrdocTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
Python★ 6 206↓ 326.1K/міс12 серп. 2026 р.
npm91Відміннийіндекс здоров'я
PT-Perkasa-Pilar-Utama/ppu-paddle-ocrLightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/міс1 серп. 2026 р.
PyPI91Відміннийіндекс здоров'я

ocrmypdf/OCRmyPDFOCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Python★ 34.4K5 серп. 2026 р.
PyPI90Відміннийіндекс здоров'я

pymupdf/PyMuPDFPyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Python · SWIG★ 10.6K28 серп. 2026 р.
npm · Packagist · crates.io +190Відміннийіндекс здоров'я

xberg-io/xbergPolyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9 228↓ 1 625/міс28 серп. 2026 р.
—89Відміннийіндекс здоров'я
C · Rust★ 89531 лип. 2026 р.
NuGet89Відміннийіндекс здоров'я
C#★ 522 лип. 2026 р.
PyPI89Відміннийіндекс здоров'я

opendatalab/MinerUTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/міс4 серп. 2026 р.
PyPI89Відміннийіндекс здоров'я
Python★ 57↓ 28.6M/міс8 серп. 2026 р.
Go89Відміннийіндекс здоров'я
y3owk1n/neruNavigate your entire screen without touching the mouse.
Go · Objective-C★ 46517 лип. 2026 р.
npm88Відміннийіндекс здоров'я
TypeScript★ 29↓ 248.7K/міс1 серп. 2026 р.
RubyGems87Відміннийіндекс здоров'я
Ruby★ 1415 лип. 2026 р.
PyPI87Відміннийіндекс здоров'я

robocorp/rpaframeworkCollection of open-source libraries and tools for Robotic Process Automation (RPA), designed to be used with both Robot Framework and Python
Python★ 1 545↓ 550.8K/міс13 серп. 2026 р.
PyPI · npm · crates.io86Відміннийіндекс здоров'я

firecrawl/pdf-inspectorFast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/міс22 серп. 2026 р.
npm86Відміннийіндекс здоров'я

harshankur/officeParserA robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/міс27 серп. 2026 р.
crates.io84Відміннийіндекс здоров'я

bzsanti/oxidizePdfPure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/міс22 серп. 2026 р.
PyPI84Відміннийіндекс здоров'я
Python★ 38.3K↓ 484.3K/міс5 серп. 2026 р.
PyPI83Відміннийіндекс здоров'я

datalab-to/suryaOCR, layout analysis, reading order, table recognition in 90+ languages
Python★ 21.2K↓ 968.8K/міс5 серп. 2026 р.
Go · npm83Відміннийіндекс здоров'я
icereed/paperless-gptUse LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
Go · TypeScript★ 2 5962 серп. 2026 р.
PyPI81Відміннийіндекс здоров'я
jztan/pdf-mcpAn MCP server that lets Claude Code and other AI agents work through large PDFs without overflowing their context — search by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts.
Python★ 91↓ 8 068/міс1 серп. 2026 р.