Python★ 1,189↓ 1.3M/月2026年8月13日
intel/auto-roundA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,5202026年7月16日
pytorch/aoPyTorch native quantization and sparsity for training and inference
Python · C++★ 2,9092026年7月21日
Python · Cuda★ 8,442↓ 5.6M/月2026年8月27日
vllm-project/llm-compressorTransformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Python★ 3,580↓ 190.2K/月2026年7月25日
huggingface/optimum🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 3,4482026年7月21日
ModelCloud/GPTQModelLLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,2072026年7月17日
Python★ 1,5542026年7月22日
Rust · Python★ 16.2K↓ 76K/月2026年8月23日
C++ · Python★ 4,645↓ 12.4M/月2026年8月27日

rajveer43/VeloxQuant-MLXFast KV-cache quantization for Apple Silicon (MLX) — 43 research-adapted compression methods with Metal kernels
Python★ 15↓ 7,549/月2026年9月5日
Python★ 192↓ 228.3K/月2026年8月28日
Python★ 83↓ 1,116/月2026年8月28日
Python★ 232026年8月1日
Python★ 2↓ 2,506/月2026年8月22日
crates.io · npm · PyPI78良好健康指数

ohdearquant/latticeRun, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust · Python★ 40↓ 19.1K/月2026年8月22日

ahb-sjsu/turboquant-proConsumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/月2026年9月5日
Python · Cuda★ 1,054↓ 300.3K/月2026年8月13日
Python · Cuda★ 1172026年7月18日

asher/mlx-kquantNative K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/月2026年8月23日
varjoranta/turboquant-vllmTurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/月2026年7月22日

FedericoTs/quantprobeRun a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/月2026年8月20日

RobTand/gridbookOut-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/月2026年8月15日
Python★ 24.7K2026年8月5日
google-ai-edge/LiteRT-CLIA convenient CLI to streamline LiteRT related development workflows, including converting, quantizing, compiling, managing, running, benchmarking and visualizing LiteRT (TFLite) models on various hardwares (CPU / GPU / NPU) across platforms (desktop, mobile or cloud).
Python★ 34↓ 80/月2026年7月15日
jagmarques/nexusquantTraining-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 252026年7月16日
TypeScript★ 0↓ 2,468/月2026年7月20日
jjang-ai/jangqJANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 2142026年7月22日
PyPI · npm · crates.io53中等健康指数
JavaScript · Rust★ 2↓ 3,094/月2026年9月4日
Python★ 1↓ 2,064/月2026年8月1日