Todas las etiquetas
Etiqueta del catálogo

#llm-inference

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

19 registros
Con la etiqueta «llm-inference»Ordenado por índice de salud
Maven · PyPI
90Excelenteíndice de salud
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.3K17 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
crates.io · PyPI
89Excelenteíndice de salud
dynemo-ai/dynemo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 752619 jul 2026
Licencia propia19 jul 2026 · métricas 1.13.0
PyPI · crates.io
88Excelenteíndice de salud
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 751417 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
Go · PyPI
87Excelenteíndice de salud
kubeflow/kfserving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 570517 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI
85Excelenteíndice de salud
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda · C++★ 5988↓ 3M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
Go · PyPI
85Excelenteíndice de salud
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 570517 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI
79Buenoíndice de salud
lemonade-sdk/lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
C++ · Python · TypeScript★ 496417 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
Packagist
78Buenoíndice de salud
neuron-core/neuron-ai
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that can interact with your data.
PHP★ 2012↓ 168.1K/mes15 jul 2026
MIT15 jul 2026 · métricas 1.13.0
Go · npm
75Buenoíndice de salud
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
crates.io · npm
75Buenoíndice de salud
ruvnet/ruvector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust★ 4351↓ 0/mes13 jul 2026
MIT13 jul 2026 · métricas 1.13.0
Go · npm
73Buenoíndice de salud
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI
70Buenoíndice de salud
Nayjest/lm-proxy
OpenAI-compatible HTTP LLM proxy / gateway for multi-provider inference (Google, Anthropic, OpenAI, PyTorch). Lightweight, extensible Python/FastAPI—use as library or standalone service.
Python★ 14115 jul 2026
MIT15 jul 2026 · métricas 1.13.0
npm
65Moderadoíndice de salud
requestyai/ai-sdk-requesty
The Requesty integration package for the Vercel AI SDK.
TypeScript★ 8↓ 15.5K/mes22 jul 2026
Apache-2.022 jul 2026 · métricas 1.13.0
npm
62Moderadoíndice de salud
webgptorg/promptbook
Turn your company's scattered knowledge into AI ready Books ✨
TypeScript★ 164↓ 11.5K/mes15 jul 2026
Licencia propia15 jul 2026 · métricas 1.13.0
PyPI
60Moderadoíndice de salud
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516 jul 2026
Licencia propia16 jul 2026 · métricas 1.13.0
Go · PyPI
59Moderadoíndice de salud
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI
55Moderadoíndice de salud
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
55Moderadoíndice de salud
google/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
npm
48En riesgoíndice de salud
empirical-run/empirical
Test and evaluate LLMs and model configurations, across all the scenarios that matter for your application
TypeScript★ 16715 jul 2026
MIT15 jul 2026 · métricas 1.13.0