Todas las etiquetas
Etiqueta del catálogo

#llm-serving

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

10 registros
Con la etiqueta «llm-serving»Ordenado por índice de salud
Maven · PyPI
100Excepcionalíndice de salud
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.4K5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI · crates.io
99Excepcionalíndice de salud
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2K4 ago 2026
Apache-2.04 ago 2026 · métricas 2.10.0
PyPI
96Excepcionalíndice de salud
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8782↓ 256.2K/mes12 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI
95Excepcionalíndice de salud
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5422↓ 2132/mes2 ago 2026
Apache-2.02 ago 2026 · métricas 2.10.0
PyPI · Go
95Excepcionalíndice de salud
skypilot-org/skypilot
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Python★ 10.5K↓ 2M/mes26 ago 2026
Apache-2.026 ago 2026 · métricas 2.10.0
Go
86Excelenteíndice de salud
NexusGPU/tensor-fusion
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 15817 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
Go
84Excelenteíndice de salud
nexusgpu/tensor-fusion-operator
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 15815 jul 2026
Apache-2.015 jul 2026 · métricas 2.10.0
Go
65Buenoíndice de salud
aimd54/palan
Pull, push, pack, and serve GGUF models as OCI ModelPack artifacts — daemonless, air-gap-first, one binary
Go★ 017 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
Go · PyPI
65Buenoíndice de salud
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2218/mes15 ago 2026
Apache-2.015 ago 2026 · métricas 2.10.0