Alle Tags
Katalog-Tag

#model-serving

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

7 Einträge
Getaggt als „model-serving“Geordnet nach Gesundheitsindex
crates.io · PyPI
88ExzellentGesundheitsindex
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 86.1K↓ 0/Monat13. Juli 2026
Apache-2.013. Juli 2026 · Metriken 1.13.0
Go · PyPI
87ExzellentGesundheitsindex
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5.70517. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
Go · PyPI
85ExzellentGesundheitsindex
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 5.70517. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
Go · npm
73GutGesundheitsindex
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121. Juli 2026
Apache-2.021. Juli 2026 · Metriken 1.13.0
PyPI
55MittelGesundheitsindex
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118. Juli 2026
Apache-2.018. Juli 2026 · Metriken 1.13.0
PyPI
55MittelGesundheitsindex
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118. Juli 2026
Apache-2.018. Juli 2026 · Metriken 1.13.0
PyPI
22KritischGesundheitsindex
Lightning-Universe/stable-diffusion-deploy
Learn to serve Stable Diffusion models on cloud infrastructure at scale. This Lightning App shows load-balancing, orchestrating, pre-provisioning, dynamic batching, GPU-inference, micro-services working together via the Lightning Apps framework.
Python · TypeScript★ 39121. Juli 2026
Apache-2.021. Juli 2026 · Metriken 1.13.0