All tags
Catalogue tag

#nlp

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

144 records
Tagged “nlp”Ranked by health index
PyPI · crates.io
57Moderatehealth index
typangaa/canto-hk-g2p
Fast Cantonese text-to-Jyutping (G2P) — Rust core, Python (PyO3). Numbers, dates, and HK English code-switching.
Python · Rust★ 0↓ 2,726/moAug 2, 2026
Custom licenseAug 2, 2026 · metrics 2.10.0
PyPI
56Moderatehealth index
buchwandler/abbr2words
Multilingual, context-aware abbreviation expansion for text normalization and speech.
Python★ 0↓ 11K/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
ryanpavlicek/pyaegean
A specialist Python toolkit for Ancient Greek - alphabetic Greek NLP (incl. a neural pipeline) and the Aegean syllabic scripts (Linear A, Linear B, Cypriot, Cypro-Minoan)
Python★ 1↓ 11.3K/moJul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 2.10.0
RubyGems
54Moderatehealth index
yohasebe/engtagger
English Part-of-Speech Tagger Library; a Ruby port of Lingua::EN::Tagger
Ruby★ 276Jul 27, 2026
GPL-2.0Jul 27, 2026 · metrics 2.10.0
npm
53Moderatehealth index
VisualText/npm-package-nlpengine
Node.js bindings for the NLP++ text-analysis engine (npm peer of the Python package NLPPlus)
JavaScript · C++ · CMake★ 0↓ 3,348/moJul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
crates.io
53Moderatehealth index
laisuk/opencc-jieba-rs
A high performance Rust-based Chinese text converter that performs word segmentation using Jieba and OpenCC lexicons.
Rust · C★ 2↓ 2,822/moAug 29, 2026
MITAug 29, 2026 · metrics 2.10.0
PyPI
53Moderatehealth index
rmovva/HypotheSAEs
HypotheSAEs: hypothesizing interpretable relationships in text datasets using sparse autoencoders. https://arxiv.org/abs/2502.04382
Jupyter Notebook★ 91↓ 240/moJul 27, 2026
Apache-2.0Jul 27, 2026 · metrics 2.10.0
RubyGems
53Moderatehealth index
yohasebe/wp2txt
A command-line tool to extract plain text from Wikipedia dumps with category and section filtering
Ruby★ 195Jul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
PyPI
51Moderatehealth index
686f6c61/pypi-legal-expand
Smart expansion of Spanish legal acronyms. 647 verified acronyms from RAE, BOE and DPEJ. | Expansión inteligente de siglas legales españolas.
Python★ 3↓ 2,237/moJul 18, 2026
Custom licenseJul 18, 2026 · metrics 2.10.0
crates.io
51Moderatehealth index
Xuanwo/frostem
Pre-built Snowball stemmers for Rust (tracks snowball main)
Rust★ 3↓ 2,078/moAug 11, 2026
BSD-3-ClauseAug 11, 2026 · metrics 2.10.0
Go
51Moderatehealth index
bachtiarpanjaitan/ihandai-go
Provider-agnostic Go AI framework for building RAG, agents, and AI-powered applications. Swap LLM providers without code changes — supports Ollama, OpenAI, and more.
Go★ 1Jul 18, 2026
No licenseJul 18, 2026 · metrics 2.10.0
crates.io
51Moderatehealth index
jorge-menjivar/tekken-rs
Rust implementation of the Mistral Tekken tokenizer
Rust★ 7Aug 23, 2026
Apache-2.0Aug 23, 2026 · metrics 2.10.0
crates.io
51Moderatehealth index
jqueguiner/presidio-rs
A Rust port of Microsoft Presidio — PII detection & anonymization (analyzer + anonymizer + Python bindings)
Rust★ 2Jul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
Hex
51Moderatehealth index
minibikini/paasaa
🔤 Natural language detection for Elixir without AI
Elixir★ 144↓ 2,060/moSep 3, 2026
MITSep 3, 2026 · metrics 2.10.0
PyPI
51Moderatehealth index
pemistahl/lingua-py
The most accurate natural language detection library for Python, suitable for short text and mixed-language text
Python★ 1,775Aug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI · crates.io
50Moderatehealth index
marcelroed/gigatoken
Language model tokenization at GB/s
Rust · Python★ 115↓ 2,741/moJul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
PyPI
50Moderatehealth index
mholtzscher/syllapy
Calculate syllable count for English words.
Python · Just★ 42↓ 1.2M/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
PyPI
48Weakhealth index
aki0ka/meisai-checker
特許明細書の自動方式チェックツール(CLI/GUI/MCP対応)
Python · HTML★ 0↓ 4,088/moJul 19, 2026
No licenseJul 19, 2026 · metrics 2.10.0
PyPI
48Weakhealth index
hplt-project/sacremoses
Python port of Moses tokenizer, truecaser and normalizer
Python★ 497↓ 2.7M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
Packagist
48Weakhealth index
nitotm/efficient-language-detector
Fast and accurate natural language detection. Detector written in PHP. Nito-ELD, ELD.
PHP★ 63↓ 33.6K/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
crates.io
47Weakhealth index
MicheleYin/misaki-rs
Rust port of Misaki
Rust★ 10↓ 16.7K/moSep 4, 2026
MITSep 4, 2026 · metrics 2.10.0
crates.io
47Weakhealth index
allo-media/text2num-rs
Parse and convert numbers written in English, Dutch, Spanish, German, Italian or French into their digit representation.
Rust★ 9↓ 2,277/moJul 29, 2026
MITJul 29, 2026 · metrics 2.10.0
PyPI
45Weakhealth index
kpwhri/konsepy
Framework for build NLP information extraction systems using regular expressions.
Python★ 1↓ 194/moSep 5, 2026
No licenseSep 5, 2026 · metrics 2.10.0
PyPI
44Weakhealth index
polm/fugashi
A Cython MeCab wrapper for fast, pythonic Japanese tokenization and morphological analysis.
C++ · Cython · Python★ 533Jul 21, 2026
MITJul 21, 2026 · metrics 2.10.0
PyPI
44Weakhealth index
promplate/partial-json-parser
Parse partial JSON generated by LLM
Python★ 137↓ 7.1M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
RubyGems
44Weakhealth index
yohasebe/ruby-spacy
A wrapper module for using spaCy natural language processing library from the Ruby programming language via PyCall
Ruby★ 67Jul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
PyPI
42Weakhealth index
explosion/spacy-loggers
📟 Logging utilities for spaCy
Python★ 12↓ 23.1M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
Hex
41Weakhealth index
nshkrdotcom/agent_session_manager
Agent Session Manager - A comprehensive Elixir library for managing AI agent sessions, state persistence, conversation context, and multi-agent orchestration workflows
Elixir★ 9Jul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
PyPI
39Weakhealth index
Tiiiger/bert_score
BERT score for text generation
Jupyter Notebook · Python★ 1,908Jul 21, 2026
MITJul 21, 2026 · metrics 2.10.0
Go
39Weakhealth index
xDarkicex/steady
Zero-allocation text classification in pure Go. Byte-level n-gram embeddings, OVA logistic regression, Platt scaling, conformal prediction. Sub-millisecond, GC-free.
Go★ 1Aug 12, 2026
MITAug 12, 2026 · metrics 2.10.0