全部标签
目录标签

#model-serving

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

10 条记录
标签为“model-serving”按健康指数排序
PyPI · crates.io
99卓越健康指数
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2K2026年8月4日
Apache-2.02026年8月4日 · 指标 2.10.0
PyPI · crates.io
98卓越健康指数
basetenlabs/truss
The simplest way to serve AI/ML models in production
Python · Rust★ 1,195↓ 966.4K/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
Go · PyPI
97卓越健康指数
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5,8372026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
PyPI
96卓越健康指数
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8,782↓ 256.2K/月2026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI
95卓越健康指数
mlrun/mlrun
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.
Python★ 1,6902026年8月2日
Apache-2.02026年8月2日 · 指标 2.10.0
Go · npm
87优秀健康指数
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 4812026年7月21日
Apache-2.02026年7月21日 · 指标 2.10.0
PyPI
59中等健康指数
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 4512026年7月18日
Apache-2.02026年7月18日 · 指标 2.10.0
PyPI · npm · Maven
59中等健康指数
hpnkv/a11
A streaming action runtime for AI agents, model serving, and multimodal APIs
C++ · Python · TypeScript★ 1↓ 17K/月2026年9月2日
无许可证2026年9月2日 · 指标 2.10.0
PyPI
57中等健康指数
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 4512026年7月18日
Apache-2.02026年7月18日 · 指标 2.10.0
PyPI
19危急健康指数
Lightning-Universe/stable-diffusion-deploy
Learn to serve Stable Diffusion models on cloud infrastructure at scale. This Lightning App shows load-balancing, orchestrating, pre-provisioning, dynamic batching, GPU-inference, micro-services working together via the Lightning Apps framework.
Python · TypeScript★ 3912026年7月21日
Apache-2.02026年7月21日 · 指标 2.10.0