Todas las etiquetas
Etiqueta del catálogo

#kv-cache

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

11 registros
Con la etiqueta «kv-cache»Ordenado por índice de salud
PyPI · Go
98Excepcionalíndice de salud
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/mes15 ago 2026
Apache-2.015 ago 2026 · métricas 2.10.0
PyPI
98Excepcionalíndice de salud
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 625612 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI
95Excepcionalíndice de salud
MemTensor/MemOS
Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings and DeepSeek Harness support.
TypeScript · Python★ 11K28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI · npm
84Excelenteíndice de salud
rajveer43/VeloxQuant-MLX
Fast KV-cache quantization for Apple Silicon (MLX) — 43 research-adapted compression methods with Metal kernels
Python★ 15↓ 7549/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
PyPI
80Excelenteíndice de salud
manjunathshiva/turboquant-mlx
Extreme weight + KV cache compression for LLMs on Apple Silicon (MLX implementation of Google's TurboQuant)
Python★ 83↓ 1116/mes28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI · crates.io
77Buenoíndice de salud
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/mes22 jul 2026
MIT22 jul 2026 · métricas 2.10.0
Go · PyPI
65Buenoíndice de salud
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516 jul 2026
Licencia propia16 jul 2026 · métricas 2.10.0
npm
51Moderadoíndice de salud
jiangge/pi-cache-optimizer
Improve Pi prompt/KV cache hit rates with stable prompts, OpenAI-compatible cache keys, proxy compat warnings, and footer cache stats.
TypeScript · Python★ 40↓ 4809/mes24 jul 2026
MIT24 jul 2026 · métricas 2.10.0
crates.io
44Débilíndice de salud
RecursiveIntell/turbo-quant
Rust implementation of TurboQuant, PolarQuant, and QJL — zero-overhead vector quantization for semantic search and KV cache compression (ICLR 2026)
Rust · Python★ 26↓ 3048/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0