Todas las etiquetas
Etiqueta del catálogo

#ocr

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

84 registros
Con la etiqueta «ocr»Ordenado por índice de salud
PyPI
97Excepcionalíndice de salud
Unstructured-IO/unstructured
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
HTML★ 15.3K5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI · npm
96Excepcionalíndice de salud
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Python · C++★ 87K4 ago 2026
Apache-2.04 ago 2026 · métricas 2.10.0
npm · PyPI
96Excepcionalíndice de salud
paperless-ngx/paperless-ngx
A community-supported supercharged document management system: scan, index and archive all your documents
Python · TypeScript★ 44.1K10 ago 2026
GPL-3.010 ago 2026 · métricas 2.10.0
npm · Maven · PyPI
95Excepcionalíndice de salud
opendataloader-project/opendataloader-pdf
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Java · Python★ 28.2K↓ 54.5K/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
krkn-chaos/krkn
Chaos and resiliency testing tool for Kubernetes with a focus on improving performance under failure conditions. A CNCF sandbox project.
Python★ 486↓ 16.5K/mes18 ago 2026
Apache-2.018 ago 2026 · métricas 2.10.0
PyPI
93Excepcionalíndice de salud
RapidAI/RapidOCR
📄 Awesome OCR multiple programing languages toolkits based on ONNX Runtime, OpenVINO, MNN, PaddlePaddle, TensorRT and PyTorch.
Python · Jupyter Notebook★ 7408↓ 3.8M/mes7 ago 2026
Apache-2.07 ago 2026 · métricas 2.10.0
npm · crates.io · PyPI
93Excepcionalíndice de salud
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 12.2K↓ 729.9K/mes22 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
npm · crates.io
93Excepcionalíndice de salud
ruvnet/RuVector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust · TypeScript★ 4441↓ 432.2K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
crates.io · PyPI · npm +5
93Excepcionalíndice de salud
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1015↓ 391.7K/mes5 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
PyPI
92Excelenteíndice de salud
microsoft/markitdown
Python tool for converting files and office documents to Markdown.
Python★ 171.5K↓ 13.7M/mes4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
PyPI
92Excelenteíndice de salud
mindee/doctr
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
Python★ 6206↓ 326.1K/mes12 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
npm
91Excelenteíndice de salud
PT-Perkasa-Pilar-Utama/ppu-paddle-ocr
Lightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
PyPI
91Excelenteíndice de salud
ocrmypdf/OCRmyPDF
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Python★ 34.4K5 ago 2026
MPL-2.05 ago 2026 · métricas 2.10.0
PyPI
90Excelenteíndice de salud
pymupdf/PyMuPDF
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Python · SWIG★ 10.6K28 ago 2026
AGPL-3.028 ago 2026 · métricas 2.10.0
npm · Packagist · crates.io +1
90Excelenteíndice de salud
xberg-io/xberg
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9228↓ 1625/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
89Excelenteíndice de salud
CCExtractor/ccextractor
CCExtractor - Official version maintained by the core team
C · Rust★ 89531 jul 2026
GPL-2.031 jul 2026 · métricas 2.10.0
NuGet
89Excelenteíndice de salud
mindee/mindee-api-dotnet
Mindee API Helper Library for .NET ecosystem
C#★ 522 jul 2026
MIT22 jul 2026 · métricas 2.10.0
PyPI
89Excelenteíndice de salud
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/mes4 ago 2026
Licencia propia4 ago 2026 · métricas 2.10.0
PyPI
89Excelenteíndice de salud
run-llama/llama-parse-py
Python SDK for OCR and document parsing in the cloud with LlamaParse
Python★ 57↓ 28.6M/mes8 ago 2026
MIT8 ago 2026 · métricas 2.10.0
Go
89Excelenteíndice de salud
y3owk1n/neru
Navigate your entire screen without touching the mouse.
Go · Objective-C★ 46517 jul 2026
MIT17 jul 2026 · métricas 2.10.0
npm
88Excelenteíndice de salud
run-llama/llama-parse-ts
Typescript SDK for OCR and document parsing in the cloud with LlamaParse
TypeScript★ 29↓ 248.7K/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
RubyGems
87Excelenteíndice de salud
mindee/mindee-api-ruby
Mindee API Helper Library for Ruby
Ruby★ 1415 jul 2026
MIT15 jul 2026 · métricas 2.10.0
PyPI
87Excelenteíndice de salud
robocorp/rpaframework
Collection of open-source libraries and tools for Robotic Process Automation (RPA), designed to be used with both Robot Framework and Python
Python★ 1545↓ 550.8K/mes13 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI · npm · crates.io
86Excelenteíndice de salud
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm
86Excelenteíndice de salud
harshankur/officeParser
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
crates.io
84Excelenteíndice de salud
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
PyPI
84Excelenteíndice de salud
datalab-to/marker
Convert PDF to markdown + JSON quickly with high accuracy
Python★ 38.3K↓ 484.3K/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI
83Excelenteíndice de salud
datalab-to/surya
OCR, layout analysis, reading order, table recognition in 90+ languages
Python★ 21.2K↓ 968.8K/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
Go · npm
83Excelenteíndice de salud
icereed/paperless-gpt
Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
Go · TypeScript★ 25962 ago 2026
MIT2 ago 2026 · métricas 2.10.0
PyPI
81Excelenteíndice de salud
jztan/pdf-mcp
An MCP server that lets Claude Code and other AI agents work through large PDFs without overflowing their context — search by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts.
Python★ 91↓ 8068/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0