All tags
Catalogue tag

#document-analysis

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

6 records
Tagged “document-analysis”Ranked by health index
npm
91Excellenthealth index
PT-Perkasa-Pilar-Utama/ppu-paddle-ocr
Lightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/moAug 4, 2026
Custom licenseAug 4, 2026 · metrics 2.10.0
NuGet
80Excellenthealth index
UglyToad/PdfPig
Read and extract text and other content from PDFs in C# (port of PDFBox)
C#★ 2,501Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
npm · crates.io · RubyGems +5
78Goodhealth index
retab-dev/retab
The developper starter pack for document processing
Go · Ruby · Python★ 45↓ 822/moAug 7, 2026
No licenseAug 7, 2026 · metrics 2.10.0
npm
67Goodhealth index
cesarandreslopez/occ
Document metrics, structure extraction, and code exploration for real repositories
TypeScript★ 6↓ 2,263/moJul 14, 2026
MITJul 14, 2026 · metrics 2.10.0
npm
63Moderatehealth index
mysleekdesigns/crawlforge-mcp
28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5,179/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0