Alle Tags
Katalog-Tag

#vllm

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

38 Einträge
Getaggt als „vllm“Geordnet nach Gesundheitsindex
PyPI · crates.io
99AußergewöhnlichGesundheitsindex
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7.886↓ 59.4K/Monat28. Aug. 2026
Eigene Lizenz28. Aug. 2026 · Metriken 2.10.0
PyPI · Go
98AußergewöhnlichGesundheitsindex
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/Monat15. Aug. 2026
Apache-2.015. Aug. 2026 · Metriken 2.10.0
PyPI
98AußergewöhnlichGesundheitsindex
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6.25612. Aug. 2026
Apache-2.012. Aug. 2026 · Metriken 2.10.0
PyPI
97AußergewöhnlichGesundheitsindex
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1.52016. Juli 2026
Apache-2.016. Juli 2026 · Metriken 2.10.0
Go · PyPI
97AußergewöhnlichGesundheitsindex
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5.83728. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
PyPI
95AußergewöhnlichGesundheitsindex
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5.422↓ 2.132/Monat2. Aug. 2026
Apache-2.02. Aug. 2026 · Metriken 2.10.0
PyPI · npm
94AußergewöhnlichGesundheitsindex
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6K5. Aug. 2026
MIT5. Aug. 2026 · Metriken 2.10.0
crates.io · PyPI
92ExzellentGesundheitsindex
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/Monat16. Juli 2026
Apache-2.016. Juli 2026 · Metriken 2.10.0
Go
91ExzellentGesundheitsindex
praetorian-inc/julius
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 1758. Aug. 2026
Apache-2.08. Aug. 2026 · Metriken 2.10.0
crates.io · PyPI
91ExzellentGesundheitsindex
smg-project/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 4322. Aug. 2026
Apache-2.02. Aug. 2026 · Metriken 2.10.0
PyPI
89ExzellentGesundheitsindex
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1.20717. Juli 2026
Eigene Lizenz17. Juli 2026 · Metriken 2.10.0
Go
89ExzellentGesundheitsindex
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 2075. Sept. 2026
Apache-2.05. Sept. 2026 · Metriken 2.10.0
Go · npm
89ExzellentGesundheitsindex
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617. Juli 2026
Apache-2.017. Juli 2026 · Metriken 2.10.0
Go · npm
87ExzellentGesundheitsindex
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121. Juli 2026
Apache-2.021. Juli 2026 · Metriken 2.10.0
npm
87ExzellentGesundheitsindex
pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 307↓ 7.921/Monat6. Sept. 2026
Apache-2.06. Sept. 2026 · Metriken 2.10.0
Go · PyPI
86ExzellentGesundheitsindex
kubeai-project/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 1.23531. Juli 2026
Apache-2.031. Juli 2026 · Metriken 2.10.0
Go · npm
86ExzellentGesundheitsindex
raids-lab/crater
Crater is a cloud-native AI training & inference platform.
TypeScript · Go · MDX★ 54228. Juli 2026
Apache-2.028. Juli 2026 · Metriken 2.10.0
Go · npm
83ExzellentGesundheitsindex
voidmind-io/voidllm
Privacy-first LLM proxy and AI gateway - load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts.
Go · TypeScript★ 1232. Aug. 2026
Eigene Lizenz2. Aug. 2026 · Metriken 2.10.0
PyPI
78GutGesundheitsindex
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2.506/Monat22. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
npm
77GutGesundheitsindex
0xsero/vllm-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
TypeScript★ 1.36816. Juli 2026
Apache-2.016. Juli 2026 · Metriken 2.10.0
PyPI · crates.io
77GutGesundheitsindex
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/Monat5. Sept. 2026
MIT5. Sept. 2026 · Metriken 2.10.0
crates.io · npm
77GutGesundheitsindex
soapbucket/sbproxy
Self-hosted AI gateway and LLM proxy. OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock and 60+ providers, or serve vLLM/llama.cpp on your GPUs. Keys, budgets, guardrails, semantic cache, MCP
Rust★ 4730. Juli 2026
Apache-2.030. Juli 2026 · Metriken 2.10.0
PyPI
75GutGesundheitsindex
freesolo-co/flash
LoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.
Python★ 2↓ 5.834/Monat29. Aug. 2026
Apache-2.029. Aug. 2026 · Metriken 2.10.0
PyPI
73GutGesundheitsindex
CocoRoF/geny-executor
Manifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9.957/Monat5. Sept. 2026
Apache-2.05. Sept. 2026 · Metriken 2.10.0
PyPI
71GutGesundheitsindex
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/Monat22. Juli 2026
MIT22. Juli 2026 · Metriken 2.10.0
npm
71GutGesundheitsindex
yeemio/owlcoda
OwlCoda — independent local-first AI coding workbench. Native REPL, 42+ tools, learned skills, GPL-3.0-or-later.
TypeScript★ 5↓ 4.505/Monat25. Juli 2026
GPL-3.025. Juli 2026 · Metriken 2.10.0
PyPI
69GutGesundheitsindex
Aitherium/awdk
Build AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 95. Sept. 2026
Eigene Lizenz5. Sept. 2026 · Metriken 2.10.0
PyPI · npm
69GutGesundheitsindex
jjang-ai/vmlx
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4.850/Monat24. Juli 2026
Apache-2.024. Juli 2026 · Metriken 2.10.0
PyPI
67GutGesundheitsindex
hackspaces/blueshark-forge
Model-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 473/Monat5. Sept. 2026
MIT5. Sept. 2026 · Metriken 2.10.0
PyPI
63MittelGesundheitsindex
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2.218/Monat15. Aug. 2026
Apache-2.015. Aug. 2026 · Metriken 2.10.0