Todas las etiquetas
Etiqueta del catálogo

#llm-inference

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

35 registros
Con la etiqueta «llm-inference»Ordenado por índice de salud
Maven · PyPI
100Excepcionalíndice de salud
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.4K5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
PyPI · crates.io
99Excepcionalíndice de salud
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7886↓ 59.4K/mes28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI
98Excepcionalíndice de salud
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 625612 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Python · Cuda★ 6263↓ 7.1M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
Go · PyPI
97Excepcionalíndice de salud
kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Go · Python★ 583728 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
96Excepcionalíndice de salud
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Python★ 8782↓ 256.2K/mes12 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI
95Excepcionalíndice de salud
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5422↓ 2132/mes2 ago 2026
Apache-2.02 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
lemonade-sdk/lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
C++ · Python · TypeScript★ 551028 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
monocle2ai/monocle
Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written in Python.
Python★ 337↓ 50.8K/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
93Excepcionalíndice de salud
Lightning-AI/litgpt
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Python★ 13.6K↓ 15.5K/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
npm
93Excepcionalíndice de salud
Nano-Collective/nanocoder
An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe nothing to anyone.
TypeScript★ 2376↓ 9541/mes25 ago 2026
Licencia propia25 ago 2026 · métricas 2.10.0
npm · crates.io
93Excepcionalíndice de salud
ruvnet/RuVector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
Rust · TypeScript★ 4441↓ 432.2K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
Packagist
92Excelenteíndice de salud
neuron-core/neuron-ai
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that can interact with your data.
PHP★ 2012↓ 168.1K/mes15 jul 2026
MIT15 jul 2026 · métricas 2.10.0
Go · npm
89Excelenteíndice de salud
matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
Go · TypeScript★ 25617 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
Go · npm
87Excelenteíndice de salud
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 48121 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
PyPI
86Excelenteíndice de salud
Nayjest/lm-proxy
OpenAI-compatible HTTP LLM proxy / gateway for multi-provider inference (Google, Anthropic, OpenAI, PyTorch). Lightweight, extensible Python/FastAPI—use as library or standalone service.
Python★ 150↓ 1822/mes26 ago 2026
MIT26 ago 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2506/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm
78Buenoíndice de salud
webgptorg/promptbook
Turn your company's scattered knowledge into AI ready Books ✨
TypeScript★ 167↓ 6355/mes6 sept 2026
Licencia propia6 sept 2026 · métricas 2.10.0
npm
75Buenoíndice de salud
requestyai/ai-sdk-requesty
The Requesty integration package for the Vercel AI SDK.
TypeScript★ 8↓ 15.5K/mes22 jul 2026
Apache-2.022 jul 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
asher/mlx-kquant
Native K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3360/mes23 ago 2026
MIT23 ago 2026 · métricas 2.10.0
Go · PyPI
65Buenoíndice de salud
anthony-chaudhary/fak
fak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 1517 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
RobTand/gridbook
Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
Python · Cuda★ 10↓ 2218/mes15 ago 2026
Apache-2.015 ago 2026 · métricas 2.10.0
Go
63Moderadoíndice de salud
flexigpt/inference-go
A single interface in Go to get inference from multiple llm/ai providers using their official SDKs
Go★ 23 ago 2026
MIT3 ago 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
Python★ 2516 jul 2026
Licencia propia16 jul 2026 · métricas 2.10.0
PyPI · crates.io
62Moderadoíndice de salud
theoddden/Terradev
An imperative command-line-interface for AI workload orchestration
Python★ 25↓ 3845/mes23 ago 2026
Apache-2.023 ago 2026 · métricas 2.10.0
60Moderadoíndice de salud
eastriverlee/LLM.swift
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
Swift★ 86628 jul 2026
MIT28 jul 2026 · métricas 2.10.0
60Moderadoíndice de salud
lean-dojo/LeanCopilot
LLMs as Copilots for Theorem Proving in Lean
C++★ 13074 ago 2026
MIT4 ago 2026 · métricas 2.10.0
PyPI
59Moderadoíndice de salud
ai-hypercomputer/jetstream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Python★ 45118 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
Go
59Moderadoíndice de salud
flexigpt/llmtools-go
LLM Tool implementations for Golang
Go★ 220 ago 2026
MIT20 ago 2026 · métricas 2.10.0