All tags
Catalogue tag

#vllm

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

21 records
Tagged “vllm”Ranked by health index
crates.io · PyPI
89Excellenthealth index
dynemo-ai/dynemo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,526Jul 19, 2026
Custom licenseJul 19, 2026 · metrics 1.13.0
PyPI · crates.io
88Excellenthealth index
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,514Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
Go · PyPI
87Excellenthealth index
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,705Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
PyPI
85Excellenthealth index
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,520Jul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 1.13.0
Go · PyPI
85Excellenthealth index
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,705Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
crates.io · PyPI
78Goodhealth index
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/moJul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 1.13.0
PyPI · npm
78Goodhealth index
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.4KJul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
npm · PyPI
77Goodhealth index
alibaba-damo-academy/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3KJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
npm · PyPI
77Goodhealth index
alibaba/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3KJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI
76Goodhealth index
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,207Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
Go · npm
75Goodhealth index
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 256Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
Go
74Goodhealth index
defilantech/llmkube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
Go · npm
73Goodhealth index
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 481Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
Go · PyPI
72Goodhealth index
kubeai-project/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 1,226Jul 22, 2026
Apache-2.0Jul 22, 2026 · metrics 1.13.0
npm
69Moderatehealth index
pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 297↓ 10.1K/moJul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 1.13.0
npm
65Moderatehealth index
0xsero/vllm-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
TypeScript★ 1,368Jul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 1.13.0
PyPI
64Moderatehealth index
hackspaces/blueshark-forge
Model-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 3,513/moJul 15, 2026
MITJul 15, 2026 · metrics 1.13.0
PyPI
63Moderatehealth index
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
PyPI
60Moderatehealth index
CocoRoF/geny-executor
Manifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9,957/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
PyPI
60Moderatehealth index
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
Custom licenseJul 16, 2026 · metrics 1.13.0
PyPI
58Moderatehealth index
Aitherium/aither-adk
Build AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 9Jul 15, 2026
Custom licenseJul 15, 2026 · metrics 1.13.0