Todas las etiquetas
Etiqueta del catálogo

#quantization

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

17 registros
Con la etiqueta «quantization»Ordenado por índice de salud
PyPI
85Excelenteíndice de salud
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 152016 jul 2026
Apache-2.016 jul 2026 · métricas 1.13.0
PyPI
84Buenoíndice de salud
bitsandbytes-foundation/bitsandbytes
Accessible large language models via k-bit quantization for PyTorch.
Python · Cuda★ 8329↓ 5.7M/mes18 jul 2026
MIT18 jul 2026 · métricas 1.13.0
PyPI
84Buenoíndice de salud
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 290921 jul 2026
Licencia propia21 jul 2026 · métricas 1.13.0
PyPI
82Buenoíndice de salud
huggingface/optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 344821 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI
76Buenoíndice de salud
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 120717 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
PyPI
75Buenoíndice de salud
Xilinx/brevitas
Brevitas: neural network quantization in PyTorch
Python★ 155422 jul 2026
Licencia propia22 jul 2026 · métricas 1.13.0
PyPI
71Buenoíndice de salud
OpenNMT/CTranslate2
Fast inference engine for Transformer models
C++ · Python★ 4577↓ 9.4M/mes21 jul 2026
MIT21 jul 2026 · métricas 1.13.0
PyPI
63Moderadoíndice de salud
Comfy-Org/comfy-kitchen
Fast kernel library for Diffusion inference with multiple compute backends.
Python · Cuda★ 11718 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
63Moderadoíndice de salud
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/mes22 jul 2026
MIT22 jul 2026 · métricas 1.13.0
PyPI
60Moderadoíndice de salud
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516 jul 2026
Licencia propia16 jul 2026 · métricas 1.13.0
PyPI
59Moderadoíndice de salud
google-ai-edge/LiteRT-CLI
A convenient CLI to streamline LiteRT related development workflows, including converting, quantizing, compiling, managing, running, benchmarking and visualizing LiteRT (TFLite) models on various hardwares (CPU / GPU / NPU) across platforms (desktop, mobile or cloud).
Python★ 34↓ 80/mes15 jul 2026
Apache-2.015 jul 2026 · métricas 1.13.0
PyPI
57Moderadoíndice de salud
SYSTRAN/faster-whisper
Faster Whisper transcription with CTranslate2
Python★ 24.4K18 jul 2026
MIT18 jul 2026 · métricas 1.13.0
npm
56Moderadoíndice de salud
mkbabb/value.js
CSS value units for color, length, angles, & c.
TypeScript★ 0↓ 2468/mes20 jul 2026
MIT20 jul 2026 · métricas 1.13.0
crates.io · PyPI
56Moderadoíndice de salud
ohdearquant/lattice
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust★ 31↓ 0/mes13 jul 2026
Apache-2.013 jul 2026 · métricas 1.13.0
npm · PyPI
56Moderadoíndice de salud
rajveer43/VeloxQuant-MLX
TurboQuant MLX implementation for Apple Silicon Faster KV cache quantization optimized for MLX
Python★ 10↓ 4108/mes14 jul 2026
MIT14 jul 2026 · métricas 1.13.0
PyPI
54Moderadoíndice de salud
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 21422 jul 2026
Sin licencia22 jul 2026 · métricas 1.13.0
PyPI
19Críticoíndice de salud
blue-oil/blueoil
Bring Deep Learning to small devices
Python · C++★ 24820 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0