全部标签
目录标签

#table-extraction

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

7 条记录
标签为“table-extraction”按健康指数排序
PyPI
90优秀健康指数
pymupdf/PyMuPDF
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Python · SWIG★ 10.6K2026年8月28日
AGPL-3.02026年8月28日 · 指标 2.10.0
npm · Packagist · crates.io +1
90优秀健康指数
xberg-io/xberg
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9,228↓ 1,625/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
crates.io
84优秀健康指数
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
PyPI
81优秀健康指数
jztan/pdf-mcp
An MCP server that lets Claude Code and other AI agents work through large PDFs without overflowing their context — search by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts.
Python★ 91↓ 8,068/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
Go
80优秀健康指数
coregx/gxpdf
GxPDF - Enterprise-grade PDF library for Go. Table extraction, text parsing, encryption, document creation.
Go★ 472026年8月7日
MIT2026年8月7日 · 指标 2.10.0
PyPI
77良好健康指数
jsvine/pdfplumber
Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
Python★ 10.7K2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
Packagist · Go · crates.io +2
67良好健康指数
kreuzberg-dev/kreuzberg-lts
Kreuzberg v4 LTS — long-term support for the v4 line (legacy; superseded by xberg for v5+). MIT-licensed.
Rust · HTML★ 10↓ 1/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0