Todas las etiquetas
Etiqueta del catálogo

#model-serving

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

7 registros
Con la etiqueta «model-serving»Ordenado por índice de salud
crates.io · PyPI
88Excelenteíndice de salud
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 86.1K↓ 0/mes13 jul 2026
Apache-2.013 jul 2026 · métricas 1.13.0
Go · PyPI
87Excelenteíndice de salud
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 570517 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
Go · PyPI
85Excelenteíndice de salud
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 570517 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
Go · npm
73Buenoíndice de salud
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI
55Moderadoíndice de salud
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
55Moderadoíndice de salud
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
22Críticoíndice de salud
Lightning-Universe/stable-diffusion-deploy
Learn to serve Stable Diffusion models on cloud infrastructure at scale. This Lightning App shows load-balancing, orchestrating, pre-provisioning, dynamic batching, GPU-inference, micro-services working together via the Lightning Apps framework.
Python · TypeScript★ 39121 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0