PyPI · crates.io99Exceptionalhealth index
Rust · Python · Go★ 7,886↓ 59.4K/moAug 28, 2026
PyPI · Go98Exceptionalhealth index

LMCache/LMCacheLMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/moAug 15, 2026
PyPI98Exceptionalhealth index

kvcache-ai/MooncakeMooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6,256Aug 12, 2026
PyPI97Exceptionalhealth index
intel/auto-roundA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,520Jul 16, 2026
Go · PyPI97Exceptionalhealth index

kserve/kserveStandardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,837Aug 28, 2026
PyPI95Exceptionalhealth index
gpustack/gpustackA GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5,422↓ 2,132/moAug 2, 2026
PyPI · npm94Exceptionalhealth index

modelscope/FunASROpen-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6KAug 5, 2026
crates.io · PyPI92Excellenthealth index
lightseekorg/smgEngine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/moJul 16, 2026
Go91Excellenthealth index

praetorian-inc/juliusSimple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 175Aug 8, 2026
crates.io · PyPI91Excellenthealth index
smg-project/smgEngine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 432Aug 2, 2026
PyPI89Excellenthealth index
ModelCloud/GPTQModelLLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Python · Cuda★ 1,207Jul 17, 2026
Go89Excellenthealth index

defilantech/LLMKubeKubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 207Sep 5, 2026
Go · npm89Excellenthealth index
matrixhub-ai/matrixhubAn Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 256Jul 17, 2026
Go · npm87Excellenthealth index
ome-projects/omeOpen Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 481Jul 21, 2026
npm87Excellenthealth index

pensarai/apexAI-powered offensive security testing using autonomous agents, directly in your terminal.
TypeScript★ 307↓ 7,921/moSep 6, 2026
Go · PyPI86Excellenthealth index
kubeai-project/kubeaiAI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
Go · Jupyter Notebook★ 1,235Jul 31, 2026
Go · npm86Excellenthealth index
TypeScript · Go · MDX★ 542Jul 28, 2026
Go · npm83Excellenthealth index
voidmind-io/voidllmPrivacy-first LLM proxy and AI gateway - load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts.
Go · TypeScript★ 123Aug 2, 2026
Python★ 2↓ 2,506/moAug 22, 2026
TypeScript★ 1,368Jul 16, 2026
PyPI · crates.io77Goodhealth index

ahb-sjsu/turboquant-proConsumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/moSep 5, 2026
crates.io · npm77Goodhealth index
soapbucket/sbproxySelf-hosted AI gateway and LLM proxy. OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock and 60+ providers, or serve vLLM/llama.cpp on your GPUs. Keys, budgets, guardrails, semantic cache, MCP
Rust★ 47Jul 30, 2026

freesolo-co/flashLoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.
Python★ 2↓ 5,834/moAug 29, 2026

CocoRoF/geny-executorManifest-driven 21-stage agent pipeline for Python — 5 LLM backends (Anthropic / OpenAI / Google / vLLM / Claude Code CLI), tools, skills, memory, sandboxed CLI runs, MCP. The engine behind Geny. Apache-2.0.
Python★ 2↓ 9,957/moSep 5, 2026
varjoranta/turboquant-vllmTurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
yeemio/owlcodaOwlCoda — independent local-first AI coding workbench. Native REPL, 42+ tools, learned skills, GPL-3.0-or-later.
TypeScript★ 5↓ 4,505/moJul 25, 2026

Aitherium/awdkBuild AI agent fleets. 3 lines, any backend, local or cloud. Effort-based model routing, 48 identities, knowledge graph memory, fleet orchestration.
Python★ 9Sep 5, 2026
PyPI · npm69Goodhealth index
jjang-ai/vmlxvMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4,850/moJul 24, 2026

hackspaces/blueshark-forgeModel-agnostic agentic runtime for the terminal — any local model becomes a capable agent. The intelligence lives in the harness, not the weights.
Python★ 1↓ 473/moSep 5, 2026
PyPI63Moderatehealth index

RobTand/gridbookOut-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026