All tags
Catalogue tag

#llm-serving

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

10 records
Tagged “llm-serving”Ranked by health index
Maven · PyPI
100Exceptionalhealth index
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.4KAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · crates.io
99Exceptionalhealth index
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2KAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI
96Exceptionalhealth index
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8,782↓ 256.2K/moAug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI
95Exceptionalhealth index
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5,422↓ 2,132/moAug 2, 2026
Apache-2.0Aug 2, 2026 · metrics 2.10.0
PyPI · Go
95Exceptionalhealth index
skypilot-org/skypilot
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Python★ 10.5K↓ 2M/moAug 26, 2026
Apache-2.0Aug 26, 2026 · metrics 2.10.0
Go
86Excellenthealth index
NexusGPU/tensor-fusion
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 158Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
Go
84Excellenthealth index
nexusgpu/tensor-fusion-operator
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 158Jul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 2.10.0
Go
65Goodhealth index
aimd54/palan
Pull, push, pack, and serve GGUF models as OCI ModelPack artifacts — daemonless, air-gap-first, one binary
Go★ 0Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
Go · PyPI
65Goodhealth index
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/moAug 15, 2026
Apache-2.0Aug 15, 2026 · metrics 2.10.0