全部标签
目录标签

#text-extraction

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

28 条记录
标签为“text-extraction”按健康指数排序
npm · crates.io · PyPI
93卓越健康指数
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 12.2K↓ 729.9K/月2026年8月22日
Apache-2.02026年8月22日 · 指标 2.10.0
crates.io · PyPI · npm +5
93卓越健康指数
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1,015↓ 391.7K/月2026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
npm
91优秀健康指数
PT-Perkasa-Pilar-Utama/ppu-paddle-ocr
Lightweight, probably the fastest PaddleOCR SDK in TypeScript. Multilingual Support. Runs anywhere JavaScript runs: Node.js, Bun, Deno, mobile react-native, web browsers, and browser extensions. Docker & CLI supported. The official SDK is browser-only.
TypeScript · HTML★ 113↓ 33.9K/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
PyPI
91优秀健康指数
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
npm · Packagist · crates.io +1
90优秀健康指数
xberg-io/xberg
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Rust★ 9,228↓ 1,625/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
PyPI
87优秀健康指数
chrismattmann/tika-python
Tika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.
Python★ 1,6662026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
Packagist · crates.io · npm +1
87优秀健康指数
xberg-io/html-to-markdown
High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 56+ document formats using streaming parsers and built-in OCR.
HTML · Rust★ 814↓ 11/月2026年7月27日
MIT2026年7月27日 · 指标 2.10.0
crates.io · PyPI · npm +1
87优秀健康指数
yfedoseev/office_oxide
The fastest Office document library for Python, Rust, Go, JS/TS, C# and WASM. DOCX, XLSX, PPTX, DOC, XLS, PPT. Up to 100× faster than python-docx/openpyxl/python-pptx. 100% pass rate on valid Office files. MIT/Apache-2.0.
Rust★ 108↓ 161.1K/月2026年8月22日
Apache-2.02026年8月22日 · 指标 2.10.0
Packagist
86优秀健康指数
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3,925/月2026年7月15日
MIT2026年7月15日 · 指标 2.10.0
PyPI · npm · crates.io
86优秀健康指数
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 16.5K↓ 606.5K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm
86优秀健康指数
harshankur/officeParser
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 533↓ 2.9M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
crates.io
84优秀健康指数
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
PyPI
80优秀健康指数
airmang/python-hwpx
Pure Python HWPX automation: read, edit, generate, and validate documents without Hancom Office.
Python★ 96↓ 75.8K/月2026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
Go
80优秀健康指数
coregx/gxpdf
GxPDF - Enterprise-grade PDF library for Go. Table extraction, text parsing, encryption, document creation.
Go★ 472026年8月7日
MIT2026年8月7日 · 指标 2.10.0
npm
78良好健康指数
unjs/unpdf
📄 PDF extraction and rendering across all JavaScript runtimes
TypeScript · JavaScript★ 1,204↓ 8.1M/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
npm
75良好健康指数
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8,650/月2026年7月17日
MIT2026年7月17日 · 指标 2.10.0
npm
71良好健康指数
happyvertical/ocr
该仓库未发布描述。
TypeScript★ 0↓ 3,363/月2026年7月15日
MIT2026年7月15日 · 指标 2.10.0
crates.io
69良好健康指数
iyulab/unpdf
High-performance Rust PDF extraction library with Markdown/JSON output, CJK/RTL support, multi-column layout detection, and Python/.NET/CLI bindings.
Rust★ 43↓ 2,335/月2026年7月21日
MIT2026年7月21日 · 指标 2.10.0
npm
67良好健康指数
happyvertical/pdf
该仓库未发布描述。
TypeScript★ 0↓ 6,093/月2026年7月23日
MIT2026年7月23日 · 指标 2.10.0
Packagist · Go · crates.io +2
67良好健康指数
kreuzberg-dev/kreuzberg-lts
Kreuzberg v4 LTS — long-term support for the v4 line (legacy; superseded by xberg for v5+). MIT-licensed.
Rust · HTML★ 10↓ 1/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
crates.io · PyPI
63中等健康指数
developer0hye/pdfplumber-rs
Evidence-driven PDF extraction for Rust, with an alpha Python pdfplumber migration path.
Rust · Python★ 2↓ 23.8K/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
npm
63中等健康指数
heey-global/deep-ocr-n8n
该仓库未发布描述。
TypeScript★ 0↓ 4,479/月2026年7月15日
MIT2026年7月15日 · 指标 2.10.0
npm
63中等健康指数
johannschopplich/pdfjs-serverless
🪭 PDF.js redistributed as a single bundle for edge and serverless runtimes
TypeScript · JavaScript★ 97↓ 50.6K/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
Go
60中等健康指数
giraffesyo/pdf
Robust, zero-dependency PDF text extraction for Go, with positioned glyphs, reading- order reconstruction, and hardened parsing.
Go★ 02026年9月5日
MIT2026年9月5日 · 指标 2.10.0
Packagist
57中等健康指数
TYPO3-Solr/ext-tika
A TYPO3 CMS extension that provides Apache Tika functionality
PHP · HTML★ 9↓ 12.2K/月2026年7月17日
GPL-3.02026年7月17日 · 指标 2.10.0
PyPI
54中等健康指数
miso-belica/jusText
Heuristic based boilerplate removal tool
Python★ 824↓ 12.3M/月2026年8月27日
BSD-2-Clause2026年8月27日 · 指标 2.10.0
PyPI · crates.io
51中等健康指数
4thel00z/pdfboss
From-scratch Rust PDF toolkit for Python: fast page rendering and text extraction via PyO3 — benchmarked faster than mainstream Python PDF libraries
Rust★ 02026年7月31日
Apache-2.02026年7月31日 · 指标 2.10.0
PyPI
25存在风险健康指数
cdown/srt
A simple library and set of tools for parsing, modifying, and composing SRT files.
Python★ 5362026年8月19日
MIT2026年8月19日 · 指标 2.10.0