Усі теги
Тег каталогу

#vllm

Усі репозиторії публічного реєстру з цим тегом — із тем GitHub або ключових слів, які публікують їхні реєстри пакетів. Здоров'я вимірюється за тією ж версіонованою методологією, що й решта реєстру.

38 записів
З тегом «vllm»Упорядковано за індексом здоров'я
PyPI · crates.io
99Винятковийіндекс здоров'я
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7 886↓ 59.4K/міс28 серп. 2026 р.
Власна ліцензія28 серп. 2026 р. · метрики 2.10.0
PyPI · Go
98Винятковийіндекс здоров'я
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/міс15 серп. 2026 р.
Apache-2.015 серп. 2026 р. · метрики 2.10.0
PyPI
98Винятковийіндекс здоров'я
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6 25612 серп. 2026 р.
Apache-2.012 серп. 2026 р. · метрики 2.10.0
PyPI
97Винятковийіндекс здоров'я
intel/auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1 52016 лип. 2026 р.
Apache-2.016 лип. 2026 р. · метрики 2.10.0
Go · PyPI
97Винятковийіндекс здоров'я
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5 83728 серп. 2026 р.
Apache-2.028 серп. 2026 р. · метрики 2.10.0
PyPI
95Винятковийіндекс здоров'я
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5 422↓ 2 132/міс2 серп. 2026 р.
Apache-2.02 серп. 2026 р. · метрики 2.10.0
PyPI · npm
94Винятковийіндекс здоров'я
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6K5 серп. 2026 р.
MIT5 серп. 2026 р. · метрики 2.10.0
crates.io · PyPI
92Відміннийіндекс здоров'я
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/міс16 лип. 2026 р.
Apache-2.016 лип. 2026 р. · метрики 2.10.0
Go
91Відміннийіндекс здоров'я
praetorian-inc/julius
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 1758 серп. 2026 р.
Apache-2.08 серп. 2026 р. · метрики 2.10.0
crates.io · PyPI
91Відміннийіндекс здоров'я
smg-project/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 4322 серп. 2026 р.
Apache-2.02 серп. 2026 р. · метрики 2.10.0
PyPI
89Відміннийіндекс здоров'я
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1 20717 лип. 2026 р.
Власна ліцензія17 лип. 2026 р. · метрики 2.10.0
Go
89Відміннийіндекс здоров'я
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 2075 вер. 2026 р.
Apache-2.05 вер. 2026 р. · метрики 2.10.0
Go · npm
89Відміннийіндекс здоров'я
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 2.10.0
Go · npm
87Відміннийіндекс здоров'я
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121 лип. 2026 р.
Apache-2.021 лип. 2026 р. · метрики 2.10.0
npm
87Відміннийіндекс здоров'я
pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 307↓ 7 921/міс6 вер. 2026 р.
Apache-2.06 вер. 2026 р. · метрики 2.10.0
Go · PyPI
86Відміннийіндекс здоров'я
kubeai-project/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 1 23531 лип. 2026 р.
Apache-2.031 лип. 2026 р. · метрики 2.10.0
Go · npm
86Відміннийіндекс здоров'я
raids-lab/crater
Crater is a cloud-native AI training & inference platform.
TypeScript · Go · MDX★ 54228 лип. 2026 р.
Apache-2.028 лип. 2026 р. · метрики 2.10.0
Go · npm
83Відміннийіндекс здоров'я
voidmind-io/voidllm
Privacy-first LLM proxy and AI gateway - load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts.
Go · TypeScript★ 1232 серп. 2026 р.
Власна ліцензія2 серп. 2026 р. · метрики 2.10.0
PyPI
78Добрийіндекс здоров'я
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2 506/міс22 серп. 2026 р.
MIT22 серп. 2026 р. · метрики 2.10.0
npm
77Добрийіндекс здоров'я
0xsero/vllm-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
TypeScript★ 1 36816 лип. 2026 р.
Apache-2.016 лип. 2026 р. · метрики 2.10.0
PyPI · crates.io
77Добрийіндекс здоров'я
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/міс5 вер. 2026 р.
MIT5 вер. 2026 р. · метрики 2.10.0
crates.io · npm
77Добрийіндекс здоров'я
soapbucket/sbproxy
Self-hosted AI gateway and LLM proxy. OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock and 60+ providers, or serve vLLM/llama.cpp on your GPUs. Keys, budgets, guardrails, semantic cache, MCP
Rust★ 4730 лип. 2026 р.
Apache-2.030 лип. 2026 р. · метрики 2.10.0
PyPI
75Добрийіндекс здоров'я
freesolo-co/flash
LoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.
Python★ 2↓ 5 834/міс29 серп. 2026 р.
Apache-2.029 серп. 2026 р. · метрики 2.10.0
PyPI
73Добрийіндекс здоров'я
CocoRoF/geny-executor
Manifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9 957/міс5 вер. 2026 р.
Apache-2.05 вер. 2026 р. · метрики 2.10.0
PyPI
71Добрийіндекс здоров'я
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/міс22 лип. 2026 р.
MIT22 лип. 2026 р. · метрики 2.10.0
npm
71Добрийіндекс здоров'я
yeemio/owlcoda
OwlCoda — independent local-first AI coding workbench. Native REPL, 42+ tools, learned skills, GPL-3.0-or-later.
TypeScript★ 5↓ 4 505/міс25 лип. 2026 р.
GPL-3.025 лип. 2026 р. · метрики 2.10.0
PyPI
69Добрийіндекс здоров'я
Aitherium/awdk
Build AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 95 вер. 2026 р.
Власна ліцензія5 вер. 2026 р. · метрики 2.10.0
PyPI · npm
69Добрийіндекс здоров'я
jjang-ai/vmlx
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4 850/міс24 лип. 2026 р.
Apache-2.024 лип. 2026 р. · метрики 2.10.0
PyPI
67Добрийіндекс здоров'я
hackspaces/blueshark-forge
Model-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 473/міс5 вер. 2026 р.
MIT5 вер. 2026 р. · метрики 2.10.0
PyPI
63Помірнийіндекс здоров'я
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2 218/міс15 серп. 2026 р.
Apache-2.015 серп. 2026 р. · метрики 2.10.0