Todas las etiquetas
Etiqueta del catálogo

#vllm

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

21 registros
Con la etiqueta «vllm»Ordenado por índice de salud
crates.io · PyPI
89Excelenteíndice de salud
dynemo-ai/dynemo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 752619 jul 2026
Licencia propia19 jul 2026 · métricas 1.13.0
PyPI · crates.io
88Excelenteíndice de salud
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 751417 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
Go · PyPI
87Excelenteíndice de salud
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 570517 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI
85Excelenteíndice de salud
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 152016 jul 2026
Apache-2.016 jul 2026 · métricas 1.13.0
Go · PyPI
85Excelenteíndice de salud
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 570517 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
crates.io · PyPI
78Buenoíndice de salud
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/mes16 jul 2026
Apache-2.016 jul 2026 · métricas 1.13.0
PyPI · npm
78Buenoíndice de salud
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.4K21 jul 2026
MIT21 jul 2026 · métricas 1.13.0
npm · PyPI
77Buenoíndice de salud
alibaba-damo-academy/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3K17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
npm · PyPI
77Buenoíndice de salud
alibaba/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3K17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
PyPI
76Buenoíndice de salud
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 120717 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
Go · npm
75Buenoíndice de salud
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
Go
74Buenoíndice de salud
defilantech/llmkube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
Go · npm
73Buenoíndice de salud
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
Go · PyPI
72Buenoíndice de salud
kubeai-project/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 122622 jul 2026
Apache-2.022 jul 2026 · métricas 1.13.0
npm
69Moderadoíndice de salud
pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 297↓ 10.1K/mes15 jul 2026
Apache-2.015 jul 2026 · métricas 1.13.0
npm
65Moderadoíndice de salud
0xsero/vllm-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
TypeScript★ 136816 jul 2026
Apache-2.016 jul 2026 · métricas 1.13.0
PyPI
64Moderadoíndice de salud
hackspaces/blueshark-forge
Model-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 3513/mes15 jul 2026
MIT15 jul 2026 · métricas 1.13.0
PyPI
63Moderadoíndice de salud
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/mes22 jul 2026
MIT22 jul 2026 · métricas 1.13.0
PyPI
60Moderadoíndice de salud
CocoRoF/geny-executor
Manifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9957/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
PyPI
60Moderadoíndice de salud
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516 jul 2026
Licencia propia16 jul 2026 · métricas 1.13.0
PyPI
58Moderadoíndice de salud
Aitherium/aither-adk
Build AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 915 jul 2026
Licencia propia15 jul 2026 · métricas 1.13.0