全部标签
目录标签

#quantization

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

17 条记录
标签为“quantization”按健康指数排序
PyPI
85优秀健康指数
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,5202026年7月16日
Apache-2.02026年7月16日 · 指标 1.13.0
PyPI
84良好健康指数
bitsandbytes-foundation/bitsandbytes
Accessible large language models via k-bit quantization for PyTorch.
Python · Cuda★ 8,329↓ 5.7M/月2026年7月18日
MIT2026年7月18日 · 指标 1.13.0
PyPI
84良好健康指数
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 2,9092026年7月21日
自定义许可证2026年7月21日 · 指标 1.13.0
PyPI
82良好健康指数
huggingface/optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 3,4482026年7月21日
Apache-2.02026年7月21日 · 指标 1.13.0
PyPI
76良好健康指数
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,2072026年7月17日
自定义许可证2026年7月17日 · 指标 1.13.0
PyPI
75良好健康指数
Xilinx/brevitas
Brevitas: neural network quantization in PyTorch
Python★ 1,5542026年7月22日
自定义许可证2026年7月22日 · 指标 1.13.0
PyPI
71良好健康指数
OpenNMT/CTranslate2
Fast inference engine for Transformer models
C++ · Python★ 4,577↓ 9.4M/月2026年7月21日
MIT2026年7月21日 · 指标 1.13.0
PyPI
63中等健康指数
Comfy-Org/comfy-kitchen
Fast kernel library for Diffusion inference with multiple compute backends.
Python · Cuda★ 1172026年7月18日
Apache-2.02026年7月18日 · 指标 1.13.0
PyPI
63中等健康指数
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/月2026年7月22日
MIT2026年7月22日 · 指标 1.13.0
PyPI
60中等健康指数
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 252026年7月16日
自定义许可证2026年7月16日 · 指标 1.13.0
PyPI
59中等健康指数
google-ai-edge/LiteRT-CLI
A convenient CLI to streamline LiteRT related development workflows, including converting, quantizing, compiling, managing, running, benchmarking and visualizing LiteRT (TFLite) models on various hardwares (CPU / GPU / NPU) across platforms (desktop, mobile or cloud).
Python★ 34↓ 80/月2026年7月15日
Apache-2.02026年7月15日 · 指标 1.13.0
PyPI
57中等健康指数
SYSTRAN/faster-whisper
Faster Whisper transcription with CTranslate2
Python★ 24.4K2026年7月18日
MIT2026年7月18日 · 指标 1.13.0
npm
56中等健康指数
mkbabb/value.js
CSS value units for color, length, angles, & c.
TypeScript★ 0↓ 2,468/月2026年7月20日
MIT2026年7月20日 · 指标 1.13.0
crates.io · PyPI
56中等健康指数
ohdearquant/lattice
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust★ 31↓ 0/月2026年7月13日
Apache-2.02026年7月13日 · 指标 1.13.0
npm · PyPI
56中等健康指数
rajveer43/VeloxQuant-MLX
TurboQuant MLX implementation for Apple Silicon Faster KV cache quantization optimized for MLX
Python★ 10↓ 4,108/月2026年7月14日
MIT2026年7月14日 · 指标 1.13.0
PyPI
54中等健康指数
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 2142026年7月22日
无许可证2026年7月22日 · 指标 1.13.0
PyPI
19危急健康指数
blue-oil/blueoil
Bring Deep Learning to small devices
Python · C++★ 2482026年7月20日
Apache-2.02026年7月20日 · 指标 1.13.0