Усі теги
Тег каталогу

#model-serving

Усі репозиторії публічного реєстру з цим тегом — із тем GitHub або ключових слів, які публікують їхні реєстри пакетів. Здоров'я вимірюється за тією ж версіонованою методологією, що й решта реєстру.

7 записів
З тегом «model-serving»Упорядковано за індексом здоров'я
crates.io · PyPI
88Відміннийіндекс здоров'я
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 86.1K↓ 0/міс13 лип. 2026 р.
Apache-2.013 лип. 2026 р. · метрики 1.13.0
Go · PyPI
87Відміннийіндекс здоров'я
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5 70517 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 1.13.0
Go · PyPI
85Відміннийіндекс здоров'я
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5 70517 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 1.13.0
Go · npm
73Добрийіндекс здоров'я
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121 лип. 2026 р.
Apache-2.021 лип. 2026 р. · метрики 1.13.0
PyPI
55Помірнийіндекс здоров'я
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 лип. 2026 р.
Apache-2.018 лип. 2026 р. · метрики 1.13.0
PyPI
55Помірнийіндекс здоров'я
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 лип. 2026 р.
Apache-2.018 лип. 2026 р. · метрики 1.13.0
PyPI
22Критичнийіндекс здоров'я
Lightning-Universe/stable-diffusion-deploy
Learn to serve Stable Diffusion models on cloud infrastructure at scale. This Lightning App shows load-balancing, orchestrating, pre-provisioning, dynamic batching, GPU-inference, micro-services working together via the Lightning Apps framework.
Python · TypeScript★ 39121 лип. 2026 р.
Apache-2.021 лип. 2026 р. · метрики 1.13.0