Alle Tags
Katalog-Tag

#model-serving

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

10 Einträge
Getaggt als „model-serving“Geordnet nach Gesundheitsindex
PyPI · crates.io
99AußergewöhnlichGesundheitsindex
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2K4. Aug. 2026
Apache-2.04. Aug. 2026 · Metriken 2.10.0
PyPI · crates.io
98AußergewöhnlichGesundheitsindex
basetenlabs/truss
The simplest way to serve AI/ML models in production
Python · Rust★ 1.195↓ 966.4K/Monat27. Aug. 2026
MIT27. Aug. 2026 · Metriken 2.10.0
Go · PyPI
97AußergewöhnlichGesundheitsindex
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5.83728. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
PyPI
96AußergewöhnlichGesundheitsindex
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8.782↓ 256.2K/Monat12. Aug. 2026
Apache-2.012. Aug. 2026 · Metriken 2.10.0
PyPI
95AußergewöhnlichGesundheitsindex
mlrun/mlrun
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.
Python★ 1.6902. Aug. 2026
Apache-2.02. Aug. 2026 · Metriken 2.10.0
Go · npm
87ExzellentGesundheitsindex
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121. Juli 2026
Apache-2.021. Juli 2026 · Metriken 2.10.0
PyPI
59MittelGesundheitsindex
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118. Juli 2026
Apache-2.018. Juli 2026 · Metriken 2.10.0
PyPI · npm · Maven
59MittelGesundheitsindex
hpnkv/a11
A streaming action runtime for AI agents, model serving, and multimodal APIs
C++ · Python · TypeScript★ 1↓ 17K/Monat2. Sept. 2026
Keine Lizenz2. Sept. 2026 · Metriken 2.10.0
PyPI
57MittelGesundheitsindex
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118. Juli 2026
Apache-2.018. Juli 2026 · Metriken 2.10.0
PyPI
19KritischGesundheitsindex
Lightning-Universe/stable-diffusion-deploy
Learn to serve Stable Diffusion models on cloud infrastructure at scale. This Lightning App shows load-balancing, orchestrating, pre-provisioning, dynamic batching, GPU-inference, micro-services working together via the Lightning Apps framework.
Python · TypeScript★ 39121. Juli 2026
Apache-2.021. Juli 2026 · Metriken 2.10.0