Alle Tags
Katalog-Tag

#vllm

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

21 Einträge
Getaggt als „vllm“Geordnet nach Gesundheitsindex
crates.io · PyPI
89ExzellentGesundheitsindex
dynemo-ai/dynemo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7.52619. Juli 2026
Eigene Lizenz19. Juli 2026 · Metriken 1.13.0
PyPI · crates.io
88ExzellentGesundheitsindex
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7.51417. Juli 2026
Eigene Lizenz17. Juli 2026 · Metriken 1.13.0
Go · PyPI
87ExzellentGesundheitsindex
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5.70517. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
PyPI
85ExzellentGesundheitsindex
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1.52016. Juli 2026
Apache-2.016. Juli 2026 · Metriken 1.13.0
Go · PyPI
85ExzellentGesundheitsindex
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5.70517. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
crates.io · PyPI
78GutGesundheitsindex
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/Monat16. Juli 2026
Apache-2.016. Juli 2026 · Metriken 1.13.0
PyPI · npm
78GutGesundheitsindex
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.4K21. Juli 2026
MIT21. Juli 2026 · Metriken 1.13.0
npm · PyPI
77GutGesundheitsindex
alibaba-damo-academy/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3K17. Juli 2026
MIT17. Juli 2026 · Metriken 1.13.0
npm · PyPI
77GutGesundheitsindex
alibaba/funasr
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.3K17. Juli 2026
MIT17. Juli 2026 · Metriken 1.13.0
PyPI
76GutGesundheitsindex
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1.20717. Juli 2026
Eigene Lizenz17. Juli 2026 · Metriken 1.13.0
Go · npm
75GutGesundheitsindex
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
Go
74GutGesundheitsindex
defilantech/llmkube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/Monat14. Juli 2026
Apache-2.014. Juli 2026 · Metriken 1.13.0
Go · npm
73GutGesundheitsindex
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121. Juli 2026
Apache-2.021. Juli 2026 · Metriken 1.13.0
Go · PyPI
72GutGesundheitsindex
kubeai-project/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 1.22622. Juli 2026
Apache-2.022. Juli 2026 · Metriken 1.13.0
npm
69MittelGesundheitsindex
pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 297↓ 10.1K/Monat15. Juli 2026
Apache-2.015. Juli 2026 · Metriken 1.13.0
npm
65MittelGesundheitsindex
0xsero/vllm-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
TypeScript★ 1.36816. Juli 2026
Apache-2.016. Juli 2026 · Metriken 1.13.0
PyPI
64MittelGesundheitsindex
hackspaces/blueshark-forge
Model-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 3.513/Monat15. Juli 2026
MIT15. Juli 2026 · Metriken 1.13.0
PyPI
63MittelGesundheitsindex
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/Monat22. Juli 2026
MIT22. Juli 2026 · Metriken 1.13.0
PyPI
60MittelGesundheitsindex
CocoRoF/geny-executor
Manifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9.957/Monat14. Juli 2026
Apache-2.014. Juli 2026 · Metriken 1.13.0
PyPI
60MittelGesundheitsindex
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516. Juli 2026
Eigene Lizenz16. Juli 2026 · Metriken 1.13.0
PyPI
58MittelGesundheitsindex
Aitherium/aither-adk
Build AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 915. Juli 2026
Eigene Lizenz15. Juli 2026 · Metriken 1.13.0