Rust · Python · Go★ 7,886↓ 59.4K/月2026年8月28日

LMCache/LMCacheLMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/月2026年8月15日

kvcache-ai/MooncakeMooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6,2562026年8月12日
intel/auto-roundA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,5202026年7月16日

kserve/kserveStandardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,8372026年8月28日
gpustack/gpustackA GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5,422↓ 2,132/月2026年8月2日

modelscope/FunASROpen-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6K2026年8月5日
lightseekorg/smgEngine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/月2026年7月16日

praetorian-inc/juliusSimple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 1752026年8月8日
smg-project/smgEngine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 4322026年8月2日
ModelCloud/GPTQModelLLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,2072026年7月17日

defilantech/LLMKubeKubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 2072026年9月5日
matrixhub-ai/matrixhubAn Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 2562026年7月17日
ome-projects/omeOpen Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 4812026年7月21日

pensarai/apexAI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 307↓ 7,921/月2026年9月6日
kubeai-project/kubeaiAI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 1,2352026年7月31日
TypeScript · Go · MDX★ 5422026年7月28日
voidmind-io/voidllmPrivacy-first LLM proxy and AI gateway - load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts.
Go · TypeScript★ 1232026年8月2日
Python★ 2↓ 2,506/月2026年8月22日
TypeScript★ 1,3682026年7月16日

ahb-sjsu/turboquant-proConsumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/月2026年9月5日
soapbucket/sbproxySelf-hosted AI gateway and LLM proxy. OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock and 60+ providers, or serve vLLM/llama.cpp on your GPUs. Keys, budgets, guardrails, semantic cache, MCP
Rust★ 472026年7月30日

freesolo-co/flashLoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.
Python★ 2↓ 5,834/月2026年8月29日

CocoRoF/geny-executorManifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9,957/月2026年9月5日
varjoranta/turboquant-vllmTurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/月2026年7月22日
yeemio/owlcodaOwlCoda — independent local-first AI coding workbench. Native REPL, 42+ tools, learned skills, GPL-3.0-or-later.
TypeScript★ 5↓ 4,505/月2026年7月25日

Aitherium/awdkBuild AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 92026年9月5日
jjang-ai/vmlxvMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4,850/月2026年7月24日

hackspaces/blueshark-forgeModel-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 473/月2026年9月5日

RobTand/gridbookOut-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/月2026年8月15日