全部标签
目录标签

#vllm

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

38 条记录
标签为“vllm”按健康指数排序
PyPI · crates.io
99卓越健康指数
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,886↓ 59.4K/月2026年8月28日
自定义许可证2026年8月28日 · 指标 2.10.0
PyPI · Go
98卓越健康指数
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/月2026年8月15日
Apache-2.02026年8月15日 · 指标 2.10.0
PyPI
98卓越健康指数
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6,2562026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI
97卓越健康指数
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,5202026年7月16日
Apache-2.02026年7月16日 · 指标 2.10.0
Go · PyPI
97卓越健康指数
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,8372026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
PyPI
95卓越健康指数
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5,422↓ 2,132/月2026年8月2日
Apache-2.02026年8月2日 · 指标 2.10.0
PyPI · npm
94卓越健康指数
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6K2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
crates.io · PyPI
92优秀健康指数
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/月2026年7月16日
Apache-2.02026年7月16日 · 指标 2.10.0
Go
91优秀健康指数
praetorian-inc/julius
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 1752026年8月8日
Apache-2.02026年8月8日 · 指标 2.10.0
crates.io · PyPI
91优秀健康指数
smg-project/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 4322026年8月2日
Apache-2.02026年8月2日 · 指标 2.10.0
PyPI
89优秀健康指数
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,2072026年7月17日
自定义许可证2026年7月17日 · 指标 2.10.0
Go
89优秀健康指数
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 2072026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
Go · npm
89优秀健康指数
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 2562026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
Go · npm
87优秀健康指数
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 4812026年7月21日
Apache-2.02026年7月21日 · 指标 2.10.0
npm
87优秀健康指数
pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 307↓ 7,921/月2026年9月6日
Apache-2.02026年9月6日 · 指标 2.10.0
Go · PyPI
86优秀健康指数
kubeai-project/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 1,2352026年7月31日
Apache-2.02026年7月31日 · 指标 2.10.0
Go · npm
86优秀健康指数
raids-lab/crater
Crater is a cloud-native AI training & inference platform.
TypeScript · Go · MDX★ 5422026年7月28日
Apache-2.02026年7月28日 · 指标 2.10.0
Go · npm
83优秀健康指数
voidmind-io/voidllm
Privacy-first LLM proxy and AI gateway - load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts.
Go · TypeScript★ 1232026年8月2日
自定义许可证2026年8月2日 · 指标 2.10.0
PyPI
78良好健康指数
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2,506/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm
77良好健康指数
0xsero/vllm-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
TypeScript★ 1,3682026年7月16日
Apache-2.02026年7月16日 · 指标 2.10.0
PyPI · crates.io
77良好健康指数
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/月2026年9月5日
MIT2026年9月5日 · 指标 2.10.0
crates.io · npm
77良好健康指数
soapbucket/sbproxy
Self-hosted AI gateway and LLM proxy. OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock and 60+ providers, or serve vLLM/llama.cpp on your GPUs. Keys, budgets, guardrails, semantic cache, MCP
Rust★ 472026年7月30日
Apache-2.02026年7月30日 · 指标 2.10.0
PyPI
75良好健康指数
freesolo-co/flash
LoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.
Python★ 2↓ 5,834/月2026年8月29日
Apache-2.02026年8月29日 · 指标 2.10.0
PyPI
73良好健康指数
CocoRoF/geny-executor
Manifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9,957/月2026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
PyPI
71良好健康指数
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/月2026年7月22日
MIT2026年7月22日 · 指标 2.10.0
npm
71良好健康指数
yeemio/owlcoda
OwlCoda — independent local-first AI coding workbench. Native REPL, 42+ tools, learned skills, GPL-3.0-or-later.
TypeScript★ 5↓ 4,505/月2026年7月25日
GPL-3.02026年7月25日 · 指标 2.10.0
PyPI
69良好健康指数
Aitherium/awdk
Build AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 92026年9月5日
自定义许可证2026年9月5日 · 指标 2.10.0
PyPI · npm
69良好健康指数
jjang-ai/vmlx
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4,850/月2026年7月24日
Apache-2.02026年7月24日 · 指标 2.10.0
PyPI
67良好健康指数
hackspaces/blueshark-forge
Model-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 473/月2026年9月5日
MIT2026年9月5日 · 指标 2.10.0
PyPI
63中等健康指数
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/月2026年8月15日
Apache-2.02026年8月15日 · 指标 2.10.0