全部标签
目录标签

#tesseract

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

7 条记录
标签为“tesseract”按健康指数排序
PyPI
76良好健康指数
pymupdf/pymupdf
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Python · SWIG★ 10.2K2026年7月17日
AGPL-3.02026年7月17日 · 指标 1.13.0
PyPI
75良好健康指数
ocrmypdf/OCRmyPDF
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Python★ 34.2K↓ 990.1K/月2026年7月20日
MPL-2.02026年7月20日 · 指标 1.13.0
Packagist · crates.io · npm +1
73良好健康指数
xberg-io/xberg
A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured information from PDFs, Office documents, images, and 97+ formats. Available for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, TypeScript (Node/Bun/Wasm/Deno)- or use via CLI, REST API, or MCP server.
Rust★ 8,675↓ 40/月2026年7月20日
MIT2026年7月20日 · 指标 1.13.0
npm
61中等健康指数
happyvertical/ocr
该仓库未发布描述。
TypeScript★ 0↓ 3,363/月2026年7月15日
MIT2026年7月15日 · 指标 1.13.0
crates.io · Go · npm +2
61中等健康指数
kreuzberg-dev/kreuzberg-lts
Kreuzberg v4 LTS — long-term support for the v4 line (legacy; superseded by xberg for v5+). MIT-licensed.
Rust★ 2↓ 0/月2026年7月13日
MIT2026年7月13日 · 指标 1.13.0
RubyGems
53中等健康指数
dannnylo/rtesseract
Ruby library for working with the Tesseract OCR.
Ruby★ 8812026年7月22日
MIT2026年7月22日 · 指标 1.13.0
npm
33存在风险健康指数
harshankur/officeParser
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 502↓ 2.1M/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0