Maven · PyPI100Exceptionalhealth index

ray-project/rayRay is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.4KAug 5, 2026
PyPI · crates.io99Exceptionalhealth index
Rust · Python · Go★ 7,886↓ 59.4K/moAug 28, 2026
PyPI98Exceptionalhealth index

kvcache-ai/MooncakeMooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6,256Aug 12, 2026
PyPI97Exceptionalhealth index
Python · Cuda★ 6,263↓ 7.1M/moAug 27, 2026
Go · PyPI97Exceptionalhealth index

kserve/kserveStandardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,837Aug 28, 2026
PyPI96Exceptionalhealth index

bentoml/BentoMLThe easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8,782↓ 256.2K/moAug 12, 2026
PyPI95Exceptionalhealth index
gpustack/gpustackA GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5,422↓ 2,132/moAug 2, 2026
PyPI94Exceptionalhealth index

lemonade-sdk/lemonadeLemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
C++ · Python · TypeScript★ 5,510Aug 28, 2026
PyPI94Exceptionalhealth index

monocle2ai/monocleMonocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written in Python.
Python★ 337↓ 50.8K/moAug 28, 2026
PyPI93Exceptionalhealth index

Lightning-AI/litgpt20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Python★ 13.6K↓ 15.5K/moAug 27, 2026
npm93Exceptionalhealth index

Nano-Collective/nanocoderAn open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe nothing to anyone.
TypeScript★ 2,376↓ 9,541/moAug 25, 2026
npm · crates.io93Exceptionalhealth index

ruvnet/RuVectorRuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust · TypeScript★ 4,441↓ 432.2K/moAug 22, 2026
Packagist92Excellenthealth index
neuron-core/neuron-aiThe Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that can interact with your data.
PHP★ 2,012↓ 168.1K/moJul 15, 2026
Go · npm89Excellenthealth index
matrixhub-ai/matrixhubAn Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 256Jul 17, 2026
Go · npm87Excellenthealth index
ome-projects/omeOpen Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 481Jul 21, 2026
PyPI86Excellenthealth index

Nayjest/lm-proxyOpenAI-compatible HTTP LLM proxy / gateway for multi-provider inference (Google, Anthropic, OpenAI, PyTorch). Lightweight, extensible Python/FastAPI—use as library or standalone service.
Python★ 150↓ 1,822/moAug 26, 2026
PyPI86Excellenthealth index
Python · JavaScript★ 7,278Aug 28, 2026
Python★ 2↓ 2,506/moAug 22, 2026
TypeScript★ 167↓ 6,355/moSep 6, 2026
TypeScript★ 8↓ 15.5K/moJul 22, 2026

asher/mlx-kquantNative K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/moAug 23, 2026
Go · PyPI65Goodhealth index
anthony-chaudhary/fakfak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
PyPI63Moderatehealth index

RobTand/gridbookOut-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026
flexigpt/inference-goA single interface in Go to get inference from multiple llm/ai providers using their official SDKs
Go★ 2Aug 3, 2026
PyPI63Moderatehealth index
jagmarques/nexusquantTraining-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
PyPI · crates.io62Moderatehealth index
Python★ 25↓ 3,845/moAug 23, 2026
eastriverlee/LLM.swiftLLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
Swift★ 866Jul 28, 2026
C++★ 1,307Aug 4, 2026
PyPI59Moderatehealth index
ai-hypercomputer/jetstreamJetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
Go★ 2Aug 20, 2026