全部标签
目录标签

#ocr

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

84 条记录
标签为“ocr”按健康指数排序
PyPI
97卓越健康指数
Unstructured-IO/unstructured
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
HTML★ 15.3K2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI · npm
96卓越健康指数
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87K2026年8月4日
Apache-2.02026年8月4日 · 指标 2.10.0
npm · PyPI
96卓越健康指数
paperless-ngx/paperless-ngx
A community-supported supercharged document management system: scan, index and archive all your documents
Python · TypeScript★ 44.1K2026年8月10日
GPL-3.02026年8月10日 · 指标 2.10.0
npm · Maven · PyPI
95卓越健康指数
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Java · Python★ 28.2K↓ 54.5K/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI
94卓越健康指数
krkn-chaos/krkn
Chaos and resiliency testing tool for Kubernetes with a focus on improving performance under failure conditions. A CNCF sandbox project.
Python★ 486↓ 16.5K/月2026年8月18日
Apache-2.02026年8月18日 · 指标 2.10.0
PyPI
93卓越健康指数
RapidAI/RapidOCR
📄 Awesome OCR multiple programing languages toolkits based on ONNX Runtime, OpenVINO, MNN, PaddlePaddle, TensorRT and PyTorch.
Python · Jupyter Notebook★ 7,408↓ 3.8M/月2026年8月7日
Apache-2.02026年8月7日 · 指标 2.10.0
npm · crates.io · PyPI
93卓越健康指数
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 12.2K↓ 729.9K/月2026年8月22日
Apache-2.02026年8月22日 · 指标 2.10.0
npm · crates.io
93卓越健康指数
ruvnet/RuVector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust · TypeScript★ 4,441↓ 432.2K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
crates.io · PyPI · npm +5
93卓越健康指数
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1,015↓ 391.7K/月2026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
PyPI
92优秀健康指数
microsoft/markitdown
Python tool for converting files and office documents to Markdown.
Python★ 171.5K↓ 13.7M/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
PyPI
92优秀健康指数
mindee/doctr
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
Python★ 6,206↓ 326.1K/月2026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
npm
91优秀健康指数
PT-Perkasa-Pilar-Utama/ppu-paddle-ocr
Lightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
PyPI
91优秀健康指数
ocrmypdf/OCRmyPDF
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Python★ 34.4K2026年8月5日
MPL-2.02026年8月5日 · 指标 2.10.0
PyPI
90优秀健康指数
pymupdf/PyMuPDF
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Python · SWIG★ 10.6K2026年8月28日
AGPL-3.02026年8月28日 · 指标 2.10.0
npm · Packagist · crates.io +1
90优秀健康指数
xberg-io/xberg
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9,228↓ 1,625/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
89优秀健康指数
CCExtractor/ccextractor
CCExtractor - Official version maintained by the core team
C · Rust★ 8952026年7月31日
GPL-2.02026年7月31日 · 指标 2.10.0
NuGet
89优秀健康指数
mindee/mindee-api-dotnet
Mindee API Helper Library for .NET ecosystem
C#★ 52026年7月22日
MIT2026年7月22日 · 指标 2.10.0
PyPI
89优秀健康指数
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/月2026年8月4日
自定义许可证2026年8月4日 · 指标 2.10.0
PyPI
89优秀健康指数
run-llama/llama-parse-py
Python SDK for OCR and document parsing in the cloud with LlamaParse
Python★ 57↓ 28.6M/月2026年8月8日
MIT2026年8月8日 · 指标 2.10.0
Go
89优秀健康指数
y3owk1n/neru
Navigate your entire screen without touching the mouse.
Go · Objective-C★ 4652026年7月17日
MIT2026年7月17日 · 指标 2.10.0
npm
88优秀健康指数
run-llama/llama-parse-ts
Typescript SDK for OCR and document parsing in the cloud with LlamaParse
TypeScript★ 29↓ 248.7K/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
RubyGems
87优秀健康指数
mindee/mindee-api-ruby
Mindee API Helper Library for Ruby
Ruby★ 142026年7月15日
MIT2026年7月15日 · 指标 2.10.0
PyPI
87优秀健康指数
robocorp/rpaframework
Collection of open-source libraries and tools for Robotic Process Automation (RPA), designed to be used with both Robot Framework and Python
Python★ 1,545↓ 550.8K/月2026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
PyPI · npm · crates.io
86优秀健康指数
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm
86优秀健康指数
harshankur/officeParser
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
crates.io
84优秀健康指数
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
PyPI
84优秀健康指数
datalab-to/marker
Convert PDF to markdown + JSON quickly with high accuracy
Python★ 38.3K↓ 484.3K/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI
83优秀健康指数
datalab-to/surya
OCR, layout analysis, reading order, table recognition in 90+ languages
Python★ 21.2K↓ 968.8K/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
Go · npm
83优秀健康指数
icereed/paperless-gpt
Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
Go · TypeScript★ 2,5962026年8月2日
MIT2026年8月2日 · 指标 2.10.0
PyPI
81优秀健康指数
jztan/pdf-mcp
An MCP server that lets Claude Code and other AI agents work through large PDFs without overflowing their context — search by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts.
Python★ 91↓ 8,068/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0