All tags
Catalogue tag

#llm-inference

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

35 records
Tagged “llm-inference”Ranked by health index
Maven · PyPI
100Exceptionalhealth index
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.4KAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · crates.io
99Exceptionalhealth index
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,886↓ 59.4K/moAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
PyPI
98Exceptionalhealth index
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6,256Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI
97Exceptionalhealth index
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda★ 6,263↓ 7.1M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
Go · PyPI
97Exceptionalhealth index
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,837Aug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
96Exceptionalhealth index
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8,782↓ 256.2K/moAug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI
95Exceptionalhealth index
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5,422↓ 2,132/moAug 2, 2026
Apache-2.0Aug 2, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
lemonade-sdk/lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
C++ · Python · TypeScript★ 5,510Aug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
monocle2ai/monocle
Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written in Python.
Python★ 337↓ 50.8K/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
93Exceptionalhealth index
Lightning-AI/litgpt
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Python★ 13.6K↓ 15.5K/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
npm
93Exceptionalhealth index
Nano-Collective/nanocoder
An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe nothing to anyone.
TypeScript★ 2,376↓ 9,541/moAug 25, 2026
Custom licenseAug 25, 2026 · metrics 2.10.0
npm · crates.io
93Exceptionalhealth index
ruvnet/RuVector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust · TypeScript★ 4,441↓ 432.2K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
Packagist
92Excellenthealth index
neuron-core/neuron-ai
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that can interact with your data.
PHP★ 2,012↓ 168.1K/moJul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
Go · npm
89Excellenthealth index
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 256Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
Go · npm
87Excellenthealth index
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 481Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
Nayjest/lm-proxy
OpenAI-compatible HTTP LLM proxy / gateway for multi-provider inference (Google, Anthropic, OpenAI, PyTorch). Lightweight, extensible Python/FastAPI—use as library or standalone service.
Python★ 150↓ 1,822/moAug 26, 2026
MITAug 26, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2,506/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm
78Goodhealth index
webgptorg/promptbook
Turn your company's scattered knowledge into AI ready Books ✨
TypeScript★ 167↓ 6,355/moSep 6, 2026
Custom licenseSep 6, 2026 · metrics 2.10.0
npm
75Goodhealth index
requestyai/ai-sdk-requesty
The Requesty integration package for the Vercel AI SDK.
TypeScript★ 8↓ 15.5K/moJul 22, 2026
Apache-2.0Jul 22, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
asher/mlx-kquant
Native K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/moAug 23, 2026
MITAug 23, 2026 · metrics 2.10.0
Go · PyPI
65Goodhealth index
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026
Apache-2.0Aug 15, 2026 · metrics 2.10.0
Go
63Moderatehealth index
flexigpt/inference-go
A single interface in Go to get inference from multiple llm/ai providers using their official SDKs
Go★ 2Aug 3, 2026
MITAug 3, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
Custom licenseJul 16, 2026 · metrics 2.10.0
PyPI · crates.io
62Moderatehealth index
theoddden/Terradev
An imperative command-line-interface for AI workload orchestration
Python★ 25↓ 3,845/moAug 23, 2026
Apache-2.0Aug 23, 2026 · metrics 2.10.0
60Moderatehealth index
eastriverlee/LLM.swift
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
Swift★ 866Jul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
60Moderatehealth index
lean-dojo/LeanCopilot
LLMs as Copilots for Theorem Proving in Lean
C++★ 1,307Aug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
PyPI
59Moderatehealth index
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
Go
59Moderatehealth index
flexigpt/llmtools-go
LLM Tool implementations for Golang
Go★ 2Aug 20, 2026
MITAug 20, 2026 · metrics 2.10.0