All tags
Catalogue tag

#quantization

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

34 records
Tagged “quantization”Ranked by health index
PyPI
99Exceptionalhealth index
openvinotoolkit/nncf
Neural Network Compression Framework for enhanced OpenVINO™ inference
Python★ 1,189↓ 1.3M/moAug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
97Exceptionalhealth index
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,520Jul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 2.10.0
PyPI
97Exceptionalhealth index
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 2,909Jul 21, 2026
Custom licenseJul 21, 2026 · metrics 2.10.0
PyPI
96Exceptionalhealth index
bitsandbytes-foundation/bitsandbytes
Accessible large language models via k-bit quantization for PyTorch.
Python · Cuda★ 8,442↓ 5.6M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI
96Exceptionalhealth index
vllm-project/llm-compressor
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Python★ 3,580↓ 190.2K/moJul 25, 2026
Apache-2.0Jul 25, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
huggingface/optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 3,448Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,207Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
Xilinx/brevitas
Brevitas: neural network quantization in PyTorch
Python★ 1,554Jul 22, 2026
Custom licenseJul 22, 2026 · metrics 2.10.0
PyPI · crates.io
88Excellenthealth index
RyanCodrai/turbovec
A vector index built on TurboQuant, written in Rust with Python bindings
Rust · Python★ 16.2K↓ 76K/moAug 23, 2026
MITAug 23, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
OpenNMT/CTranslate2
Fast inference engine for Transformer models
C++ · Python★ 4,645↓ 12.4M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI · npm
84Excellenthealth index
rajveer43/VeloxQuant-MLX
Fast KV-cache quantization for Apple Silicon (MLX) — 43 research-adapted compression methods with Metal kernels
Python★ 15↓ 7,549/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI
83Excellenthealth index
google-ai-edge/ai-edge-quantizer
AI Edge Quantizer: flexible post training quantization for LiteRT models.
Python★ 192↓ 228.3K/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
manjunathshiva/turboquant-mlx
Extreme weight + KV cache compression for LLMs on Apple Silicon (MLX implementation of Google's TurboQuant)
Python★ 83↓ 1,116/moAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
mobilint/mblt-model-zoo
Mobilint Model Zoo Project
Python★ 23Aug 1, 2026
BSD-3-ClauseAug 1, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2,506/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
crates.io · npm · PyPI
78Goodhealth index
ohdearquant/lattice
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust · Python★ 40↓ 19.1K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
PyPI · crates.io
77Goodhealth index
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI
77Goodhealth index
huggingface/optimum-quanto
A pytorch quantization backend for optimum
Python · Cuda★ 1,054↓ 300.3K/moAug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
Comfy-Org/comfy-kitchen
Fast kernel library for Diffusion inference with multiple compute backends.
Python · Cuda★ 117Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
asher/mlx-kquant
Native K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/moAug 23, 2026
MITAug 23, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
PyPI · crates.io
63Moderatehealth index
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/moAug 20, 2026
MITAug 20, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026
Apache-2.0Aug 15, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
SYSTRAN/faster-whisper
Faster Whisper transcription with CTranslate2
Python★ 24.7KAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
google-ai-edge/LiteRT-CLI
A convenient CLI to streamline LiteRT related development workflows, including converting, quantizing, compiling, managing, running, benchmarking and visualizing LiteRT (TFLite) models on various hardwares (CPU / GPU / NPU) across platforms (desktop, mobile or cloud).
Python★ 34↓ 80/moJul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
Custom licenseJul 16, 2026 · metrics 2.10.0
npm
60Moderatehealth index
mkbabb/value.js
CSS value units for color, length, angles, & c.
TypeScript★ 0↓ 2,468/moJul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 214Jul 22, 2026
No licenseJul 22, 2026 · metrics 2.10.0
PyPI · npm · crates.io
53Moderatehealth index
JunHwan-Kwon/deepbom
Local static analysis and evidence generation for deployed AI model artifacts
JavaScript · Rust★ 2↓ 3,094/moSep 4, 2026
Apache-2.0Sep 4, 2026 · metrics 2.10.0
PyPI
50Moderatehealth index
pjordanandrsn/experts4bit-qlora
QLoRA fine-tuning of fused 4-bit Mixture-of-Experts on a single small GPU (bitsandbytes Experts4bit)
Python★ 1↓ 2,064/moAug 1, 2026
Custom licenseAug 1, 2026 · metrics 2.10.0