全部标签
目录标签

#tokenization

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

10 条记录
标签为“tokenization”按健康指数排序
PyPI
98卓越健康指数
PyThaiNLP/pythainlp
Thai natural language processing in Python
Python★ 1,148↓ 1.5M/月2026年8月8日
Apache-2.02026年8月8日 · 指标 2.10.0
npm
88优秀健康指数
toon-format/toon
🎒 Token-Oriented Object Notation (TOON) – compact, human-readable serialization of JSON data for LLM prompts. TypeScript SDK, CLI, benchmarks.
TypeScript★ 25.1K↓ 4.4M/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI · npm
80优秀健康指数
explosion/spaCy
💫 Industrial-strength Natural Language Processing (NLP) in Python
Python · MDX · Cython★ 33.8K2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
Go
59中等健康指数
clipperhouse/uax29
A tokenizer based on Unicode text segmentation (UAX #29), for Go. Split graphemes, words, sentences.
Go★ 1192026年7月28日
MIT2026年7月28日 · 指标 2.10.0
npm
59中等健康指数
privent-ai/n8n-nodes-privent
Open-source data protection framework for AI agent workflows in n8n. Detects, tokenizes, and reversibly anonymizes sensitive data before it reaches your models.
TypeScript★ 3↓ 2,871/月2026年7月27日
MIT2026年7月27日 · 指标 2.10.0
Go
57中等健康指数
chadsr/ollamatokenizer
Ollama's internal tokenization as API endpoints.
Go · Makefile★ 12026年8月23日
MIT2026年8月23日 · 指标 2.10.0
Go
57中等健康指数
ron2111/omnitoken
Pure-Go LLM tokenizer and tiktoken-compatible token counter for OpenAI BPE, WordPiece, SentencePiece, Gemini, Llama, Mistral, and Hugging Face adapters.
Go★ 62026年9月5日
MIT2026年9月5日 · 指标 2.10.0
crates.io
56中等健康指数
bio-rs/bio-rs
AI-ready biological data I/O, validation, and tokenization engine
Rust · Python★ 32026年7月15日
自定义许可证2026年7月15日 · 指标 2.10.0
PyPI · crates.io
50中等健康指数
marcelroed/gigatoken
Language model tokenization at GB/s
Rust · Python★ 115↓ 2,741/月2026年7月22日
MIT2026年7月22日 · 指标 2.10.0
PyPI
47薄弱健康指数
OpenVoiceOS/quebra_frases
chunks strings into byte sized pieces
Python★ 12026年7月21日
Apache-2.02026年7月21日 · 指标 2.10.0