全部标签
目录标签

#vllm

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

21 条记录
标签为“vllm”按健康指数排序
crates.io · PyPI
89优秀健康指数
dynemo-ai/dynemo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,5262026年7月19日
自定义许可证2026年7月19日 · 指标 1.13.0
PyPI · crates.io
88优秀健康指数
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,5142026年7月17日
自定义许可证2026年7月17日 · 指标 1.13.0
Go · PyPI
87优秀健康指数
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,7052026年7月17日
Apache-2.02026年7月17日 · 指标 1.13.0
PyPI
85优秀健康指数
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,5202026年7月16日
Apache-2.02026年7月16日 · 指标 1.13.0
Go · PyPI
85优秀健康指数
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,7052026年7月17日
Apache-2.02026年7月17日 · 指标 1.13.0
crates.io · PyPI
78良好健康指数
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/月2026年7月16日
Apache-2.02026年7月16日 · 指标 1.13.0
PyPI · npm
78良好健康指数
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.4K2026年7月21日
MIT2026年7月21日 · 指标 1.13.0
npm · PyPI
77良好健康指数
alibaba-damo-academy/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3K2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
npm · PyPI
77良好健康指数
alibaba/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3K2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
PyPI
76良好健康指数
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,2072026年7月17日
自定义许可证2026年7月17日 · 指标 1.13.0
Go · npm
75良好健康指数
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 2562026年7月17日
Apache-2.02026年7月17日 · 指标 1.13.0
Go
74良好健康指数
defilantech/llmkube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/月2026年7月14日
Apache-2.02026年7月14日 · 指标 1.13.0
Go · npm
73良好健康指数
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 4812026年7月21日
Apache-2.02026年7月21日 · 指标 1.13.0
Go · PyPI
72良好健康指数
kubeai-project/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 1,2262026年7月22日
Apache-2.02026年7月22日 · 指标 1.13.0
npm
69中等健康指数
pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 297↓ 10.1K/月2026年7月15日
Apache-2.02026年7月15日 · 指标 1.13.0
npm
65中等健康指数
0xsero/vllm-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
TypeScript★ 1,3682026年7月16日
Apache-2.02026年7月16日 · 指标 1.13.0
PyPI
64中等健康指数
hackspaces/blueshark-forge
Model-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 3,513/月2026年7月15日
MIT2026年7月15日 · 指标 1.13.0
PyPI
63中等健康指数
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/月2026年7月22日
MIT2026年7月22日 · 指标 1.13.0
PyPI
60中等健康指数
CocoRoF/geny-executor
Manifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9,957/月2026年7月14日
Apache-2.02026年7月14日 · 指标 1.13.0
PyPI
60中等健康指数
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 252026年7月16日
自定义许可证2026年7月16日 · 指标 1.13.0
PyPI
58中等健康指数
Aitherium/aither-adk
Build AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 92026年7月15日
自定义许可证2026年7月15日 · 指标 1.13.0