PyPI99Exceptionalhealth index
Python★ 1,189↓ 1.3M/moAug 13, 2026
PyPI97Exceptionalhealth index
intel/auto-roundA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,520Jul 16, 2026
PyPI97Exceptionalhealth index
pytorch/aoPyTorch native quantization and sparsity for training and inference
Python · C++★ 2,909Jul 21, 2026
PyPI96Exceptionalhealth index
Python · Cuda★ 8,442↓ 5.6M/moAug 27, 2026
PyPI96Exceptionalhealth index
vllm-project/llm-compressorTransformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Python★ 3,580↓ 190.2K/moJul 25, 2026
PyPI94Exceptionalhealth index
huggingface/optimum🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 3,448Jul 21, 2026
PyPI89Excellenthealth index
ModelCloud/GPTQModelLLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,207Jul 17, 2026
PyPI89Excellenthealth index
Python★ 1,554Jul 22, 2026
PyPI · crates.io88Excellenthealth index
Rust · Python★ 16.2K↓ 76K/moAug 23, 2026
PyPI86Excellenthealth index
C++ · Python★ 4,645↓ 12.4M/moAug 27, 2026
PyPI · npm84Excellenthealth index

rajveer43/VeloxQuant-MLXFast KV-cache quantization for Apple Silicon (MLX) — 43 research-adapted compression methods with Metal kernels
Python★ 15↓ 7,549/moSep 5, 2026
PyPI83Excellenthealth index
Python★ 192↓ 228.3K/moAug 28, 2026
PyPI80Excellenthealth index
Python★ 83↓ 1,116/moAug 28, 2026
PyPI80Excellenthealth index
Python★ 23Aug 1, 2026
Python★ 2↓ 2,506/moAug 22, 2026
crates.io · npm · PyPI78Goodhealth index

ohdearquant/latticeRun, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust · Python★ 40↓ 19.1K/moAug 22, 2026
PyPI · crates.io77Goodhealth index

ahb-sjsu/turboquant-proConsumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/moSep 5, 2026
Python · Cuda★ 1,054↓ 300.3K/moAug 13, 2026
Python · Cuda★ 117Jul 18, 2026

asher/mlx-kquantNative K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/moAug 23, 2026
varjoranta/turboquant-vllmTurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
PyPI · crates.io63Moderatehealth index

FedericoTs/quantprobeRun a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/moAug 20, 2026
PyPI63Moderatehealth index

RobTand/gridbookOut-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026
PyPI63Moderatehealth index
Python★ 24.7KAug 5, 2026
PyPI63Moderatehealth index
google-ai-edge/LiteRT-CLIA convenient CLI to streamline LiteRT related development workflows, including converting, quantizing, compiling, managing, running, benchmarking and visualizing LiteRT (TFLite) models on various hardwares (CPU / GPU / NPU) across platforms (desktop, mobile or cloud).
Python★ 34↓ 80/moJul 15, 2026
PyPI63Moderatehealth index
jagmarques/nexusquantTraining-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
npm60Moderatehealth index
TypeScript★ 0↓ 2,468/moJul 20, 2026
PyPI54Moderatehealth index
jjang-ai/jangqJANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 214Jul 22, 2026
PyPI · npm · crates.io53Moderatehealth index
JavaScript · Rust★ 2↓ 3,094/moSep 4, 2026
PyPI50Moderatehealth index
Python★ 1↓ 2,064/moAug 1, 2026