All tags
Catalogue tag

#tokenization

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

10 records
Tagged “tokenization”Ranked by health index
PyPI
98Exceptionalhealth index
PyThaiNLP/pythainlp
Thai natural language processing in Python
Python★ 1,148↓ 1.5M/moAug 8, 2026
Apache-2.0Aug 8, 2026 · metrics 2.10.0
npm
88Excellenthealth index
toon-format/toon
🎒 Token-Oriented Object Notation (TOON) – compact, human-readable serialization of JSON data for LLM prompts. TypeScript SDK, CLI, benchmarks.
TypeScript★ 25.1K↓ 4.4M/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI · npm
80Excellenthealth index
explosion/spaCy
💫 Industrial-strength Natural Language Processing (NLP) in Python
Python · MDX · Cython★ 33.8KAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
Go
59Moderatehealth index
clipperhouse/uax29
A tokenizer based on Unicode text segmentation (UAX #29), for Go. Split graphemes, words, sentences.
Go★ 119Jul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
npm
59Moderatehealth index
privent-ai/n8n-nodes-privent
Open-source data protection framework for AI agent workflows in n8n. Detects, tokenizes, and reversibly anonymizes sensitive data before it reaches your models.
TypeScript★ 3↓ 2,871/moJul 27, 2026
MITJul 27, 2026 · metrics 2.10.0
Go
57Moderatehealth index
chadsr/ollamatokenizer
Ollama's internal tokenization as API endpoints.
Go · Makefile★ 1Aug 23, 2026
MITAug 23, 2026 · metrics 2.10.0
Go
57Moderatehealth index
ron2111/omnitoken
Pure-Go LLM tokenizer and tiktoken-compatible token counter for OpenAI BPE, WordPiece, SentencePiece, Gemini, Llama, Mistral, and Hugging Face adapters.
Go★ 6Sep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
crates.io
56Moderatehealth index
bio-rs/bio-rs
AI-ready biological data I/O, validation, and tokenization engine
Rust · Python★ 3Jul 15, 2026
Custom licenseJul 15, 2026 · metrics 2.10.0
PyPI · crates.io
50Moderatehealth index
marcelroed/gigatoken
Language model tokenization at GB/s
Rust · Python★ 115↓ 2,741/moJul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
PyPI
47Weakhealth index
OpenVoiceOS/quebra_frases
chunks strings into byte sized pieces
Python★ 1Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0