All tags
Catalogue tag

#vllm

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

38 records
Tagged “vllm”Ranked by health index
PyPI · crates.io
99Exceptionalhealth index
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,886↓ 59.4K/moAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
PyPI · Go
98Exceptionalhealth index
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/moAug 15, 2026
Apache-2.0Aug 15, 2026 · metrics 2.10.0
PyPI
98Exceptionalhealth index
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6,256Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI
97Exceptionalhealth index
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,520Jul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 2.10.0
Go · PyPI
97Exceptionalhealth index
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,837Aug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
95Exceptionalhealth index
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5,422↓ 2,132/moAug 2, 2026
Apache-2.0Aug 2, 2026 · metrics 2.10.0
PyPI · npm
94Exceptionalhealth index
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6KAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
crates.io · PyPI
92Excellenthealth index
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/moJul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 2.10.0
Go
91Excellenthealth index
praetorian-inc/julius
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 175Aug 8, 2026
Apache-2.0Aug 8, 2026 · metrics 2.10.0
crates.io · PyPI
91Excellenthealth index
smg-project/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 432Aug 2, 2026
Apache-2.0Aug 2, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,207Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
Go
89Excellenthealth index
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 207Sep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
Go · npm
89Excellenthealth index
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 256Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
Go · npm
87Excellenthealth index
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 481Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
npm
87Excellenthealth index
pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 307↓ 7,921/moSep 6, 2026
Apache-2.0Sep 6, 2026 · metrics 2.10.0
Go · PyPI
86Excellenthealth index
kubeai-project/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 1,235Jul 31, 2026
Apache-2.0Jul 31, 2026 · metrics 2.10.0
Go · npm
86Excellenthealth index
raids-lab/crater
Crater is a cloud-native AI training & inference platform.
TypeScript · Go · MDX★ 542Jul 28, 2026
Apache-2.0Jul 28, 2026 · metrics 2.10.0
Go · npm
83Excellenthealth index
voidmind-io/voidllm
Privacy-first LLM proxy and AI gateway - load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts.
Go · TypeScript★ 123Aug 2, 2026
Custom licenseAug 2, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2,506/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm
77Goodhealth index
0xsero/vllm-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
TypeScript★ 1,368Jul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 2.10.0
PyPI · crates.io
77Goodhealth index
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
crates.io · npm
77Goodhealth index
soapbucket/sbproxy
Self-hosted AI gateway and LLM proxy. OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock and 60+ providers, or serve vLLM/llama.cpp on your GPUs. Keys, budgets, guardrails, semantic cache, MCP
Rust★ 47Jul 30, 2026
Apache-2.0Jul 30, 2026 · metrics 2.10.0
PyPI
75Goodhealth index
freesolo-co/flash
LoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.
Python★ 2↓ 5,834/moAug 29, 2026
Apache-2.0Aug 29, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
CocoRoF/geny-executor
Manifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9,957/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
npm
71Goodhealth index
yeemio/owlcoda
OwlCoda — independent local-first AI coding workbench. Native REPL, 42+ tools, learned skills, GPL-3.0-or-later.
TypeScript★ 5↓ 4,505/moJul 25, 2026
GPL-3.0Jul 25, 2026 · metrics 2.10.0
PyPI
69Goodhealth index
Aitherium/awdk
Build AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 9Sep 5, 2026
Custom licenseSep 5, 2026 · metrics 2.10.0
PyPI · npm
69Goodhealth index
jjang-ai/vmlx
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4,850/moJul 24, 2026
Apache-2.0Jul 24, 2026 · metrics 2.10.0
PyPI
67Goodhealth index
hackspaces/blueshark-forge
Model-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 473/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026
Apache-2.0Aug 15, 2026 · metrics 2.10.0