全部标签
目录标签

#llm-serving

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

10 条记录
标签为“llm-serving”按健康指数排序
Maven · PyPI
100卓越健康指数
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.4K2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI · crates.io
99卓越健康指数
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2K2026年8月4日
Apache-2.02026年8月4日 · 指标 2.10.0
PyPI
96卓越健康指数
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8,782↓ 256.2K/月2026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI
95卓越健康指数
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5,422↓ 2,132/月2026年8月2日
Apache-2.02026年8月2日 · 指标 2.10.0
PyPI · Go
95卓越健康指数
skypilot-org/skypilot
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Python★ 10.5K↓ 2M/月2026年8月26日
Apache-2.02026年8月26日 · 指标 2.10.0
Go
86优秀健康指数
NexusGPU/tensor-fusion
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 1582026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
Go
84优秀健康指数
nexusgpu/tensor-fusion-operator
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 1582026年7月15日
Apache-2.02026年7月15日 · 指标 2.10.0
Go
65良好健康指数
aimd54/palan
Pull, push, pack, and serve GGUF models as OCI ModelPack artifacts — daemonless, air-gap-first, one binary
Go★ 02026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
Go · PyPI
65良好健康指数
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 152026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
PyPI
63中等健康指数
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2,218/月2026年8月15日
Apache-2.02026年8月15日 · 指标 2.10.0