All tags
Catalogue tag

#llm-inference

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

19 records
Tagged “llm-inference”Ranked by health index
Maven · PyPI
90Excellenthealth index
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.3KJul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
crates.io · PyPI
89Excellenthealth index
dynemo-ai/dynemo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,526Jul 19, 2026
Custom licenseJul 19, 2026 · metrics 1.13.0
PyPI · crates.io
88Excellenthealth index
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,514Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
Go · PyPI
87Excellenthealth index
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,705Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
PyPI
85Excellenthealth index
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda · C++★ 5,988↓ 3M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
Go · PyPI
85Excellenthealth index
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,705Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
PyPI
79Goodhealth index
lemonade-sdk/lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
C++ · Python · TypeScript★ 4,964Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
Packagist
78Goodhealth index
neuron-core/neuron-ai
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that can interact with your data.
PHP★ 2,012↓ 168.1K/moJul 15, 2026
MITJul 15, 2026 · metrics 1.13.0
Go · npm
75Goodhealth index
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 256Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
crates.io · npm
75Goodhealth index
ruvnet/ruvector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust★ 4,351↓ 0/moJul 13, 2026
MITJul 13, 2026 · metrics 1.13.0
Go · npm
73Goodhealth index
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 481Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI
70Goodhealth index
Nayjest/lm-proxy
OpenAI-compatible HTTP LLM proxy / gateway for multi-provider inference (Google, Anthropic, OpenAI, PyTorch). Lightweight, extensible Python/FastAPI—use as library or standalone service.
Python★ 141Jul 15, 2026
MITJul 15, 2026 · metrics 1.13.0
npm
65Moderatehealth index
requestyai/ai-sdk-requesty
The Requesty integration package for the Vercel AI SDK.
TypeScript★ 8↓ 15.5K/moJul 22, 2026
Apache-2.0Jul 22, 2026 · metrics 1.13.0
npm
62Moderatehealth index
webgptorg/promptbook
Turn your company's scattered knowledge into AI ready Books ✨
TypeScript★ 164↓ 11.5K/moJul 15, 2026
Custom licenseJul 15, 2026 · metrics 1.13.0
PyPI
60Moderatehealth index
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 25Jul 16, 2026
Custom licenseJul 16, 2026 · metrics 1.13.0
Go · PyPI
59Moderatehealth index
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
PyPI
55Moderatehealth index
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
55Moderatehealth index
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 451Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
npm
48At riskhealth index
empirical-run/empirical
Test and evaluate LLMs and model configurations, across all the scenarios that matter for your application
TypeScript★ 167Jul 15, 2026
MITJul 15, 2026 · metrics 1.13.0