Todas las etiquetas
Etiqueta del catálogo

#model-serving

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

10 registros
Con la etiqueta «model-serving»Ordenado por índice de salud
PyPI · crates.io
99Excepcionalíndice de salud
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2K4 ago 2026
Apache-2.04 ago 2026 · métricas 2.10.0
PyPI · crates.io
98Excepcionalíndice de salud
basetenlabs/truss
The simplest way to serve AI/ML models in production
Python · Rust★ 1195↓ 966.4K/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
Go · PyPI
97Excepcionalíndice de salud
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 583728 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
96Excepcionalíndice de salud
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8782↓ 256.2K/mes12 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI
95Excepcionalíndice de salud
mlrun/mlrun
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.
Python★ 16902 ago 2026
Apache-2.02 ago 2026 · métricas 2.10.0
Go · npm
87Excelenteíndice de salud
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
PyPI
59Moderadoíndice de salud
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
PyPI · npm · Maven
59Moderadoíndice de salud
hpnkv/a11
A streaming action runtime for AI agents, model serving, and multimodal APIs
C++ · Python · TypeScript★ 1↓ 17K/mes2 sept 2026
Sin licencia2 sept 2026 · métricas 2.10.0
PyPI
57Moderadoíndice de salud
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
PyPI
19Críticoíndice de salud
Lightning-Universe/stable-diffusion-deploy
Learn to serve Stable Diffusion models on cloud infrastructure at scale. This Lightning App shows load-balancing, orchestrating, pre-provisioning, dynamic batching, GPU-inference, micro-services working together via the Lightning Apps framework.
Python · TypeScript★ 39121 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0