All tags
Catalogue tag

#table-extraction

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

4 records
Tagged “table-extraction”Ranked by health index
PyPI
76Goodhealth index
pymupdf/pymupdf
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Python · SWIG★ 10.2KJul 17, 2026
AGPL-3.0Jul 17, 2026 · metrics 1.13.0
Packagist · crates.io · npm +1
73Goodhealth index
xberg-io/xberg
A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured information from PDFs, Office documents, images, and 97+ formats. Available for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, TypeScript (Node/Bun/Wasm/Deno)- or use via CLI, REST API, or MCP server.
Rust★ 8,675↓ 40/moJul 20, 2026
MITJul 20, 2026 · metrics 1.13.0
crates.io
69Moderatehealth index
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 182↓ 6,045/moJul 13, 2026
MITJul 13, 2026 · metrics 1.13.0
PyPI
63Moderatehealth index
jsvine/pdfplumber
Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
Python★ 10.6KJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0