Todas las etiquetas
Etiqueta del catálogo

#quantization

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

34 registros
Con la etiqueta «quantization»Ordenado por índice de salud
PyPI
99Excepcionalíndice de salud
openvinotoolkit/nncf
Neural Network Compression Framework for enhanced OpenVINO™ inference
Python★ 1189↓ 1.3M/mes13 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 152016 jul 2026
Apache-2.016 jul 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 290921 jul 2026
Licencia propia21 jul 2026 · métricas 2.10.0
PyPI
96Excepcionalíndice de salud
bitsandbytes-foundation/bitsandbytes
Accessible large language models via k-bit quantization for PyTorch.
Python · Cuda★ 8442↓ 5.6M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
PyPI
96Excepcionalíndice de salud
vllm-project/llm-compressor
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Python★ 3580↓ 190.2K/mes25 jul 2026
Apache-2.025 jul 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
huggingface/optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 344821 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
PyPI
89Excelenteíndice de salud
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 120717 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
PyPI
89Excelenteíndice de salud
Xilinx/brevitas
Brevitas: neural network quantization in PyTorch
Python★ 155422 jul 2026
Licencia propia22 jul 2026 · métricas 2.10.0
PyPI · crates.io
88Excelenteíndice de salud
RyanCodrai/turbovec
A vector index built on TurboQuant, written in Rust with Python bindings
Rust · Python★ 16.2K↓ 76K/mes23 ago 2026
MIT23 ago 2026 · métricas 2.10.0
PyPI
86Excelenteíndice de salud
OpenNMT/CTranslate2
Fast inference engine for Transformer models
C++ · Python★ 4645↓ 12.4M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
PyPI · npm
84Excelenteíndice de salud
rajveer43/VeloxQuant-MLX
Fast KV-cache quantization for Apple Silicon (MLX) — 43 research-adapted compression methods with Metal kernels
Python★ 15↓ 7549/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
PyPI
83Excelenteíndice de salud
google-ai-edge/ai-edge-quantizer
AI Edge Quantizer: flexible post training quantization for LiteRT models.
Python★ 192↓ 228.3K/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
80Excelenteíndice de salud
manjunathshiva/turboquant-mlx
Extreme weight + KV cache compression for LLMs on Apple Silicon (MLX implementation of Google's TurboQuant)
Python★ 83↓ 1116/mes28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI
80Excelenteíndice de salud
mobilint/mblt-model-zoo
Mobilint Model Zoo Project
Python★ 231 ago 2026
BSD-3-Clause1 ago 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2506/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
crates.io · npm · PyPI
78Buenoíndice de salud
ohdearquant/lattice
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust · Python★ 40↓ 19.1K/mes22 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
PyPI · crates.io
77Buenoíndice de salud
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
PyPI
77Buenoíndice de salud
huggingface/optimum-quanto
A pytorch quantization backend for optimum
Python · Cuda★ 1054↓ 300.3K/mes13 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
Comfy-Org/comfy-kitchen
Fast kernel library for Diffusion inference with multiple compute backends.
Python · Cuda★ 11718 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
asher/mlx-kquant
Native K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3360/mes23 ago 2026
MIT23 ago 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/mes22 jul 2026
MIT22 jul 2026 · métricas 2.10.0
PyPI · crates.io
63Moderadoíndice de salud
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6494/mes20 ago 2026
MIT20 ago 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2218/mes15 ago 2026
Apache-2.015 ago 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
SYSTRAN/faster-whisper
Faster Whisper transcription with CTranslate2
Python★ 24.7K5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
google-ai-edge/LiteRT-CLI
A convenient CLI to streamline LiteRT related development workflows, including converting, quantizing, compiling, managing, running, benchmarking and visualizing LiteRT (TFLite) models on various hardwares (CPU / GPU / NPU) across platforms (desktop, mobile or cloud).
Python★ 34↓ 80/mes15 jul 2026
Apache-2.015 jul 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516 jul 2026
Licencia propia16 jul 2026 · métricas 2.10.0
npm
60Moderadoíndice de salud
mkbabb/value.js
CSS value units for color, length, angles, & c.
TypeScript★ 0↓ 2468/mes20 jul 2026
MIT20 jul 2026 · métricas 2.10.0
PyPI
54Moderadoíndice de salud
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 21422 jul 2026
Sin licencia22 jul 2026 · métricas 2.10.0
PyPI · npm · crates.io
53Moderadoíndice de salud
JunHwan-Kwon/deepbom
Local static analysis and evidence generation for deployed AI model artifacts
JavaScript · Rust★ 2↓ 3094/mes4 sept 2026
Apache-2.04 sept 2026 · métricas 2.10.0
PyPI
50Moderadoíndice de salud
pjordanandrsn/experts4bit-qlora
QLoRA fine-tuning of fused 4-bit Mixture-of-Experts on a single small GPU (bitsandbytes Experts4bit)
Python★ 1↓ 2064/mes1 ago 2026
Licencia propia1 ago 2026 · métricas 2.10.0