全部标签
目录标签

#document-analysis

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

6 条记录
标签为“document-analysis”按健康指数排序
npm
91优秀健康指数
PT-Perkasa-Pilar-Utama/ppu-paddle-ocr
Lightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
PyPI
89优秀健康指数
opendatalab/MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Python★ 76.8K↓ 313.3K/月2026年8月4日
自定义许可证2026年8月4日 · 指标 2.10.0
NuGet
80优秀健康指数
UglyToad/PdfPig
Read and extract text and other content from PDFs in C# (port of PDFBox)
C#★ 2,5012026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
npm · crates.io · RubyGems +5
78良好健康指数
retab-dev/retab
The developper starter pack for document processing
Go · Ruby · Python★ 45↓ 822/月2026年8月7日
无许可证2026年8月7日 · 指标 2.10.0
npm
67良好健康指数
cesarandreslopez/occ
Document metrics, structure extraction, and code exploration for real repositories
TypeScript★ 6↓ 2,263/月2026年7月14日
MIT2026年7月14日 · 指标 2.10.0
npm
63中等健康指数
mysleekdesigns/crawlforge-mcp
28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5,179/月2026年9月5日
MIT2026年9月5日 · 指标 2.10.0