All tags
Catalogue tag

#quantization

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

17 records
Tagged “quantization”Ranked by health index
PyPI
85Excellenthealth index
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,520Jul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 1.13.0
PyPI
84Goodhealth index
bitsandbytes-foundation/bitsandbytes
Accessible large language models via k-bit quantization for PyTorch.
Python · Cuda★ 8,329↓ 5.7M/moJul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
PyPI
84Goodhealth index
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 2,909Jul 21, 2026
Custom licenseJul 21, 2026 · metrics 1.13.0
PyPI
82Goodhealth index
huggingface/optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 3,448Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI
76Goodhealth index
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,207Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
PyPI
75Goodhealth index
Xilinx/brevitas
Brevitas: neural network quantization in PyTorch
Python★ 1,554Jul 22, 2026
Custom licenseJul 22, 2026 · metrics 1.13.0
PyPI
71Goodhealth index
OpenNMT/CTranslate2
Fast inference engine for Transformer models
C++ · Python★ 4,577↓ 9.4M/moJul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
PyPI
63Moderatehealth index
Comfy-Org/comfy-kitchen
Fast kernel library for Diffusion inference with multiple compute backends.
Python · Cuda★ 117Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
63Moderatehealth index
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
PyPI
60Moderatehealth index
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
Custom licenseJul 16, 2026 · metrics 1.13.0
PyPI
59Moderatehealth index
google-ai-edge/LiteRT-CLI
A convenient CLI to streamline LiteRT related development workflows, including converting, quantizing, compiling, managing, running, benchmarking and visualizing LiteRT (TFLite) models on various hardwares (CPU / GPU / NPU) across platforms (desktop, mobile or cloud).
Python★ 34↓ 80/moJul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 1.13.0
PyPI
57Moderatehealth index
SYSTRAN/faster-whisper
Faster Whisper transcription with CTranslate2
Python★ 24.4KJul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
npm
56Moderatehealth index
mkbabb/value.js
CSS value units for color, length, angles, & c.
TypeScript★ 0↓ 2,468/moJul 20, 2026
MITJul 20, 2026 · metrics 1.13.0
crates.io · PyPI
56Moderatehealth index
ohdearquant/lattice
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust★ 31↓ 0/moJul 13, 2026
Apache-2.0Jul 13, 2026 · metrics 1.13.0
npm · PyPI
56Moderatehealth index
rajveer43/VeloxQuant-MLX
TurboQuant MLX implementation for Apple Silicon Faster KV cache quantization optimized for MLX
Python★ 10↓ 4,108/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
PyPI
54Moderatehealth index
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 214Jul 22, 2026
No licenseJul 22, 2026 · metrics 1.13.0
PyPI
19Criticalhealth index
blue-oil/blueoil
Bring Deep Learning to small devices
Python · C++★ 248Jul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 1.13.0