Todas las etiquetas
Etiqueta del catálogo

#vllm

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

38 registros
Con la etiqueta «vllm»Ordenado por índice de salud
PyPI · crates.io
99Excepcionalíndice de salud
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7886↓ 59.4K/mes28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI · Go
98Excepcionalíndice de salud
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/mes15 ago 2026
Apache-2.015 ago 2026 · métricas 2.10.0
PyPI
98Excepcionalíndice de salud
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 625612 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 152016 jul 2026
Apache-2.016 jul 2026 · métricas 2.10.0
Go · PyPI
97Excepcionalíndice de salud
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 583728 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
95Excepcionalíndice de salud
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5422↓ 2132/mes2 ago 2026
Apache-2.02 ago 2026 · métricas 2.10.0
PyPI · npm
94Excepcionalíndice de salud
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6K5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
crates.io · PyPI
92Excelenteíndice de salud
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/mes16 jul 2026
Apache-2.016 jul 2026 · métricas 2.10.0
Go
91Excelenteíndice de salud
praetorian-inc/julius
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 1758 ago 2026
Apache-2.08 ago 2026 · métricas 2.10.0
crates.io · PyPI
91Excelenteíndice de salud
smg-project/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 4322 ago 2026
Apache-2.02 ago 2026 · métricas 2.10.0
PyPI
89Excelenteíndice de salud
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 120717 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
Go
89Excelenteíndice de salud
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 2075 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
Go · npm
89Excelenteíndice de salud
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
Go · npm
87Excelenteíndice de salud
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
npm
87Excelenteíndice de salud
pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 307↓ 7921/mes6 sept 2026
Apache-2.06 sept 2026 · métricas 2.10.0
Go · PyPI
86Excelenteíndice de salud
kubeai-project/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 123531 jul 2026
Apache-2.031 jul 2026 · métricas 2.10.0
Go · npm
86Excelenteíndice de salud
raids-lab/crater
Crater is a cloud-native AI training & inference platform.
TypeScript · Go · MDX★ 54228 jul 2026
Apache-2.028 jul 2026 · métricas 2.10.0
Go · npm
83Excelenteíndice de salud
voidmind-io/voidllm
Privacy-first LLM proxy and AI gateway - load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts.
Go · TypeScript★ 1232 ago 2026
Licencia propia2 ago 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2506/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm
77Buenoíndice de salud
0xsero/vllm-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
TypeScript★ 136816 jul 2026
Apache-2.016 jul 2026 · métricas 2.10.0
PyPI · crates.io
77Buenoíndice de salud
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
crates.io · npm
77Buenoíndice de salud
soapbucket/sbproxy
Self-hosted AI gateway and LLM proxy. OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock and 60+ providers, or serve vLLM/llama.cpp on your GPUs. Keys, budgets, guardrails, semantic cache, MCP
Rust★ 4730 jul 2026
Apache-2.030 jul 2026 · métricas 2.10.0
PyPI
75Buenoíndice de salud
freesolo-co/flash
LoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.
Python★ 2↓ 5834/mes29 ago 2026
Apache-2.029 ago 2026 · métricas 2.10.0
PyPI
73Buenoíndice de salud
CocoRoF/geny-executor
Manifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9957/mes5 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/mes22 jul 2026
MIT22 jul 2026 · métricas 2.10.0
npm
71Buenoíndice de salud
yeemio/owlcoda
OwlCoda — independent local-first AI coding workbench. Native REPL, 42+ tools, learned skills, GPL-3.0-or-later.
TypeScript★ 5↓ 4505/mes25 jul 2026
GPL-3.025 jul 2026 · métricas 2.10.0
PyPI
69Buenoíndice de salud
Aitherium/awdk
Build AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 95 sept 2026
Licencia propia5 sept 2026 · métricas 2.10.0
PyPI · npm
69Buenoíndice de salud
jjang-ai/vmlx
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4850/mes24 jul 2026
Apache-2.024 jul 2026 · métricas 2.10.0
PyPI
67Buenoíndice de salud
hackspaces/blueshark-forge
Model-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 473/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2218/mes15 ago 2026
Apache-2.015 ago 2026 · métricas 2.10.0