全部标签
目录标签

#text-mining

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

9 条记录
标签为“text-mining”按健康指数排序
PyPI
91优秀健康指数
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
PyPI
86优秀健康指数
deanmalmgren/textract
extract text from any document. no muss. no fuss.
HTML · Python · Rich Text Format★ 4,6872026年8月12日
MIT2026年8月12日 · 指标 2.10.0
PyPI
83优秀健康指数
pdfminer/pdfminer.six
Community maintained fork of pdfminer - we fathom PDF
Python★ 7,015↓ 69.6M/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
Hex
47薄弱健康指数
kupolak/afinn
Sentiment analysis in Elixir.
Elixir★ 3↓ 3,944/月2026年7月17日
MIT2026年7月17日 · 指标 2.10.0
PyPI
45薄弱健康指数
kpwhri/konsepy
Framework for build NLP information extraction systems using regular expressions.
Python★ 1↓ 194/月2026年9月5日
无许可证2026年9月5日 · 指标 2.10.0
PyPI
31存在风险健康指数
lum-ai/odinson
Odinson is a powerful and highly optimized open-source framework for rule-based information extraction. Odinson couples a simple, yet powerful pattern language that can operate over multiple representations of text, with a runtime system that operates in near real time.
Scala★ 742026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
PyPI
27存在风险健康指数
csurfer/rake-nltk
Python implementation of the Rapid Automatic Keyword Extraction algorithm using NLTK.
Python★ 1,083↓ 273.5K/月2026年8月13日
MIT2026年8月13日 · 指标 2.10.0
Maven
14危急健康指数
BMDSoftware/neji
Flexible and powerful platform for biomedical information extraction from text
JavaScript · Java★ 392026年8月1日
无许可证2026年8月1日 · 指标 2.10.0
PyPI
12危急健康指数
euske/pdfminer
Python PDF Parser (Not actively maintained). Check out pdfminer.six.
Python★ 5,2742026年8月12日
MIT2026年8月12日 · 指标 2.10.0