
Unstructured-IO/unstructuredConvert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
HTML★ 15.3K2026年8月5日

PaddlePaddle/PaddleOCRTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87K2026年8月4日
Python · TypeScript★ 44.1K2026年8月10日
npm · Maven · PyPI95卓越健康指数
Java · Python★ 28.2K↓ 54.5K/月2026年8月5日

krkn-chaos/krknChaos and resiliency testing tool for Kubernetes with a focus on improving performance under failure conditions. A CNCF sandbox project.
Python★ 486↓ 16.5K/月2026年8月18日

RapidAI/RapidOCR📄 Awesome OCR multiple programing languages toolkits based on ONNX Runtime, OpenVINO, MNN, PaddlePaddle, TensorRT and PyTorch.
Python · Jupyter Notebook★ 7,408↓ 3.8M/月2026年8月7日
npm · crates.io · PyPI93卓越健康指数
Rust★ 12.2K↓ 729.9K/月2026年8月22日

ruvnet/RuVectorRuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust · TypeScript★ 4,441↓ 432.2K/月2026年8月22日
crates.io · PyPI · npm +593卓越健康指数

yfedoseev/pdf_oxideThe fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1,015↓ 391.7K/月2026年9月5日
Python★ 171.5K↓ 13.7M/月2026年8月4日

mindee/doctrdocTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
Python★ 6,206↓ 326.1K/月2026年8月12日
PT-Perkasa-Pilar-Utama/ppu-paddle-ocrLightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/月2026年8月1日

ocrmypdf/OCRmyPDFOCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Python★ 34.4K2026年8月5日

pymupdf/PyMuPDFPyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Python · SWIG★ 10.6K2026年8月28日
npm · Packagist · crates.io +190优秀健康指数

xberg-io/xbergPolyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9,228↓ 1,625/月2026年8月28日
C · Rust★ 8952026年7月31日
C#★ 52026年7月22日

opendatalab/MinerUTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/月2026年8月4日
Python★ 57↓ 28.6M/月2026年8月8日
y3owk1n/neruNavigate your entire screen without touching the mouse.
Go · Objective-C★ 4652026年7月17日
TypeScript★ 29↓ 248.7K/月2026年8月1日
Ruby★ 142026年7月15日

robocorp/rpaframeworkCollection of open-source libraries and tools for Robotic Process Automation (RPA), designed to be used with both Robot Framework and Python
Python★ 1,545↓ 550.8K/月2026年8月13日
PyPI · npm · crates.io86优秀健康指数

firecrawl/pdf-inspectorFast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/月2026年8月22日

harshankur/officeParserA robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/月2026年8月27日

bzsanti/oxidizePdfPure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/月2026年8月22日
Python★ 38.3K↓ 484.3K/月2026年8月5日

datalab-to/suryaOCR, layout analysis, reading order, table recognition in 90+ languages
Python★ 21.2K↓ 968.8K/月2026年8月5日
icereed/paperless-gptUse LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
Go · TypeScript★ 2,5962026年8月2日
jztan/pdf-mcpAn MCP server that lets Claude Code and other AI agents work through large PDFs without overflowing their context — search by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts.
Python★ 91↓ 8,068/月2026年8月1日