All tags
Catalogue tag

#kv-cache

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

5 records
Tagged “kv-cache”Ranked by health index
PyPI
63Moderatehealth index
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
PyPI
60Moderatehealth index
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
Custom licenseJul 16, 2026 · metrics 1.13.0
Go · PyPI
59Moderatehealth index
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
npm · PyPI
56Moderatehealth index
rajveer43/VeloxQuant-MLX
TurboQuant MLX implementation for Apple Silicon Faster KV cache quantization optimized for MLX
Python★ 10↓ 4,108/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
crates.io
39At riskhealth index
RecursiveIntell/turbo-quant
Rust implementation of TurboQuant, PolarQuant, and QJL — zero-overhead vector quantization for semantic search and KV cache compression (ICLR 2026)
Rust★ 29↓ 2,256/moJul 13, 2026
MITJul 13, 2026 · metrics 1.13.0