Alle Tags
Katalog-Tag

#text-extraction

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

18 Einträge
Getaggt als „text-extraction“Geordnet nach Gesundheitsindex
PyPI · crates.io · npm +5
78GutGesundheitsindex
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 886↓ 226.3K/Monat14. Juli 2026
Apache-2.014. Juli 2026 · Metriken 1.13.0
crates.io · PyPI
74GutGesundheitsindex
run-llama/liteparse
A fast, helpful, and open-source document parser
Rust★ 11.5K↓ 0/Monat14. Juli 2026
Apache-2.014. Juli 2026 · Metriken 1.13.0
Packagist
73GutGesundheitsindex
PrinsFrank/pdfparser
PHP library to read and extract text & images from PDFs - Fast & Low memory - Built from scratch
PHP★ 163↓ 3.925/Monat15. Juli 2026
MIT15. Juli 2026 · Metriken 1.13.0
Packagist · crates.io · npm +1
73GutGesundheitsindex
xberg-io/xberg
A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured information from PDFs, Office documents, images, and 97+ formats. Available for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, TypeScript (Node/Bun/Wasm/Deno)- or use via CLI, REST API, or MCP server.
Rust★ 8.675↓ 40/Monat20. Juli 2026
MIT20. Juli 2026 · Metriken 1.13.0
PyPI
72GutGesundheitsindex
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6.31218. Juli 2026
Apache-2.018. Juli 2026 · Metriken 1.13.0
crates.io · npm · Packagist +1
71GutGesundheitsindex
xberg-io/html-to-markdown
High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 56+ document formats using streaming parsers and built-in OCR.
HTML★ 806↓ 6/Monat14. Juli 2026
MIT14. Juli 2026 · Metriken 1.13.0
PyPI
69MittelGesundheitsindex
airmang/python-hwpx
Pure Python HWPX automation: read, edit, generate, and validate documents without Hancom Office.
Python★ 96↓ 75.8K/Monat20. Juli 2026
Apache-2.020. Juli 2026 · Metriken 1.13.0
crates.io
69MittelGesundheitsindex
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 182↓ 6.045/Monat13. Juli 2026
MIT13. Juli 2026 · Metriken 1.13.0
crates.io · Go · npm +1
69MittelGesundheitsindex
yfedoseev/office_oxide
The fastest Office document library for Python, Rust, Go, JS/TS, C# and WASM. DOCX, XLSX, PPTX, DOC, XLS, PPT. Up to 100× faster than python-docx/openpyxl/python-pptx. 100% pass rate on valid Office files. MIT/Apache-2.0.
Rust★ 68↓ 71K/Monat13. Juli 2026
Apache-2.013. Juli 2026 · Metriken 1.13.0
npm
64MittelGesundheitsindex
xonaman/nodejs-pdfium-native
Native Node.js bindings for PDFium
C++ · C · TypeScript★ 3↓ 8.650/Monat17. Juli 2026
MIT17. Juli 2026 · Metriken 1.13.0
crates.io · npm · PyPI
61MittelGesundheitsindex
firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust★ 1.577↓ 36.1K/Monat14. Juli 2026
MIT14. Juli 2026 · Metriken 1.13.0
npm
61MittelGesundheitsindex
happyvertical/ocr
Keine Repository-Beschreibung veröffentlicht.
TypeScript★ 0↓ 3.363/Monat15. Juli 2026
MIT15. Juli 2026 · Metriken 1.13.0
crates.io
61MittelGesundheitsindex
iyulab/unpdf
High-performance Rust PDF extraction library with Markdown/JSON output, CJK/RTL support, multi-column layout detection, and Python/.NET/CLI bindings.
Rust★ 43↓ 2.335/Monat21. Juli 2026
MIT21. Juli 2026 · Metriken 1.13.0
crates.io · Go · npm +2
61MittelGesundheitsindex
kreuzberg-dev/kreuzberg-lts
Kreuzberg v4 LTS — long-term support for the v4 line (legacy; superseded by xberg for v5+). MIT-licensed.
Rust★ 2↓ 0/Monat13. Juli 2026
MIT13. Juli 2026 · Metriken 1.13.0
npm
58MittelGesundheitsindex
heey-global/deep-ocr-n8n
Keine Repository-Beschreibung veröffentlicht.
TypeScript★ 0↓ 4.479/Monat15. Juli 2026
MIT15. Juli 2026 · Metriken 1.13.0
Packagist
56MittelGesundheitsindex
TYPO3-Solr/ext-tika
A TYPO3 CMS extension that provides Apache Tika functionality
PHP · HTML★ 9↓ 12.2K/Monat17. Juli 2026
GPL-3.017. Juli 2026 · Metriken 1.13.0
PyPI
49GefährdetGesundheitsindex
miso-belica/jusText
Heuristic based boilerplate removal tool
Python★ 818↓ 10.2M/Monat21. Juli 2026
BSD-2-Clause21. Juli 2026 · Metriken 1.13.0
npm
33GefährdetGesundheitsindex
harshankur/officeParser
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
Rich Text Format · HTML★ 502↓ 2.1M/Monat17. Juli 2026
MIT17. Juli 2026 · Metriken 1.13.0