All tags
Catalogue tag

#moe

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

11 records
Tagged “moe”Ranked by health index
PyPI · crates.io
99Exceptionalhealth index
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2KAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI · crates.io
98Exceptionalhealth index
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 31.3K↓ 144M/moAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI
97Exceptionalhealth index
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda★ 6,263↓ 7.1M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
modelscope/ms-swift
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
Python★ 15.4K↓ 111K/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,207Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
PyPI
88Excellenthealth index
NVIDIA/cudnn-frontend
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Python · C++★ 886Jul 21, 2026
MITJul 21, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
manjunathshiva/turboquant-mlx
Extreme weight + KV cache compression for LLMs on Apple Silicon (MLX implementation of Google's TurboQuant)
Python★ 83↓ 1,116/moAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
PyPI · crates.io
63Moderatehealth index
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/moAug 20, 2026
MITAug 20, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026
Apache-2.0Aug 15, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 214Jul 22, 2026
No licenseJul 22, 2026 · metrics 2.10.0
PyPI
50Moderatehealth index
pjordanandrsn/experts4bit-qlora
QLoRA fine-tuning of fused 4-bit Mixture-of-Experts on a single small GPU (bitsandbytes Experts4bit)
Python★ 1↓ 2,064/moAug 1, 2026
Custom licenseAug 1, 2026 · metrics 2.10.0