Todas las etiquetas
Etiqueta del catálogo

#moe

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

11 registros
Con la etiqueta «moe»Ordenado por índice de salud
PyPI · crates.io
99Excepcionalíndice de salud
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2K4 ago 2026
Apache-2.04 ago 2026 · métricas 2.10.0
PyPI · crates.io
98Excepcionalíndice de salud
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 31.3K↓ 144M/mes4 ago 2026
Apache-2.04 ago 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda★ 6263↓ 7.1M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
modelscope/ms-swift
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
Python★ 15.4K↓ 111K/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
PyPI
89Excelenteíndice de salud
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 120717 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
PyPI
88Excelenteíndice de salud
NVIDIA/cudnn-frontend
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Python · C++★ 88621 jul 2026
MIT21 jul 2026 · métricas 2.10.0
PyPI
80Excelenteíndice de salud
manjunathshiva/turboquant-mlx
Extreme weight + KV cache compression for LLMs on Apple Silicon (MLX implementation of Google's TurboQuant)
Python★ 83↓ 1116/mes28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI · crates.io
63Moderadoíndice de salud
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6494/mes20 ago 2026
MIT20 ago 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2218/mes15 ago 2026
Apache-2.015 ago 2026 · métricas 2.10.0
PyPI
54Moderadoíndice de salud
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 21422 jul 2026
Sin licencia22 jul 2026 · métricas 2.10.0
PyPI
50Moderadoíndice de salud
pjordanandrsn/experts4bit-qlora
QLoRA fine-tuning of fused 4-bit Mixture-of-Experts on a single small GPU (bitsandbytes Experts4bit)
Python★ 1↓ 2064/mes1 ago 2026
Licencia propia1 ago 2026 · métricas 2.10.0