Todas las etiquetas
Etiqueta del catálogo

#inference

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

106 registros
Con la etiqueta «inference»Ordenado por índice de salud
PyPI · crates.io
99Excepcionalíndice de salud
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7886↓ 59.4K/mes28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI
99Excepcionalíndice de salud
huggingface/transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Python★ 163.3K↓ 179M/mes4 ago 2026
Apache-2.04 ago 2026 · métricas 2.10.0
PyPI · crates.io
99Excepcionalíndice de salud
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2K4 ago 2026
Apache-2.04 ago 2026 · métricas 2.10.0
PyPI · Go
98Excepcionalíndice de salud
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/mes15 ago 2026
Apache-2.015 ago 2026 · métricas 2.10.0
PyPI · crates.io
98Excepcionalíndice de salud
apache/tvm-ffi
Open ABI and FFI for Machine Learning Systems
C++ · Python · Rust★ 452↓ 7.9M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
PyPI
98Excepcionalíndice de salud
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 625612 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI · crates.io
98Excepcionalíndice de salud
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 31.3K↓ 144M/mes4 ago 2026
Apache-2.04 ago 2026 · métricas 2.10.0
PyPI
98Excepcionalíndice de salud
triton-inference-server/server
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Python · Shell · C++★ 10.9K27 ago 2026
BSD-3-Clause27 ago 2026 · métricas 2.10.0
PyPI · RubyGems
97Excepcionalíndice de salud
deepspeedai/DeepSpeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Python · C++★ 42.9K↓ 1.1M/mes13 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 290921 jul 2026
Licencia propia21 jul 2026 · métricas 2.10.0
PyPI
96Excepcionalíndice de salud
mozilla-ai/any-llm
Communicate with an LLM provider using a single interface
Python★ 2133↓ 130.6K/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
npm
96Excepcionalíndice de salud
vercel/ai
The AI Toolkit for TypeScript. From the creators of Next.js, the AI SDK is a free open-source library for building AI-powered applications and agents
TypeScript · MDX★ 26K↓ 87M/mes5 ago 2026
Licencia propia5 ago 2026 · métricas 2.10.0
PyPI · crates.io
95Excepcionalíndice de salud
ai-dynamo/aiconfigurator
Offline optimization of your disaggregated Dynamo graph
Python · Rust★ 37426 jul 2026
Apache-2.026 jul 2026 · métricas 2.10.0
crates.io
95Excepcionalíndice de salud
ai-dynamo/modelexpress
Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
Python · Rust★ 103↓ 246.6K/mes1 ago 2026
Apache-2.01 ago 2026 · métricas 2.10.0
PyPI
95Excepcionalíndice de salud
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5422↓ 2132/mes2 ago 2026
Apache-2.02 ago 2026 · métricas 2.10.0
npm
95Excepcionalíndice de salud
huggingface/huggingface.js
Use Hugging Face with JavaScript
TypeScript★ 2503↓ 22.4M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
huggingface/optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 344821 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
crates.io · npm · PyPI
94Excepcionalíndice de salud
pykeio/ort
Fast ML inference & training for ONNX models in Rust
Rust★ 2477↓ 4.3M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
npm · RubyGems
94Excepcionalíndice de salud
vega/vega
A visualization grammar.
JavaScript · TypeScript★ 12K↓ 28.4M/mes27 ago 2026
BSD-3-Clause27 ago 2026 · métricas 2.10.0
PyPI · npm · RubyGems
93Excepcionalíndice de salud
mlc-ai/xgrammar
Fast, Flexible and Portable Structured Generation
C++ · Python★ 1845↓ 8.2M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
PyPI
93Excepcionalíndice de salud
pytorch/TensorRT
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
Python · Jupyter Notebook · C++★ 298430 jul 2026
BSD-3-Clause30 jul 2026 · métricas 2.10.0
PyPI
93Excepcionalíndice de salud
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Python★ 33922 ago 2026
Apache-2.02 ago 2026 · métricas 2.10.0
Go
93Excepcionalíndice de salud
run-ai/karta
Translation layer that maps any Kubernetes framework's Custom Resource Definitions (CRDs) into a standardized, generic structure.
Go★ 605 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
crates.io · PyPI
92Excelenteíndice de salud
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/mes16 jul 2026
Apache-2.016 jul 2026 · métricas 2.10.0
PyPI · npm
92Excelenteíndice de salud
roboflow/inference
Turn any computer or edge device into a command center for your computer vision projects.
Python★ 24385 sept 2026
Licencia propia5 sept 2026 · métricas 2.10.0
PyPI
91Excelenteíndice de salud
triton-inference-server/model_analyzer
Triton Model Analyzer is a CLI tool to help with better understanding of the compute and memory requirements of the Triton Inference Server models.
Python★ 523↓ 10.6K/mes29 jul 2026
Apache-2.029 jul 2026 · métricas 2.10.0
Go
91Excelenteíndice de salud
utkuozdemir/nvidia_gpu_exporter
Nvidia GPU exporter for prometheus using nvidia-smi binary
Go★ 151022 jul 2026
MIT22 jul 2026 · métricas 2.10.0
90Excelenteíndice de salud
NVIDIA/nvcf
Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
Go · Rust★ 18017 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI
90Excelenteíndice de salud
Tencent/ncnn
ncnn is a high-performance neural network inference framework optimized for the mobile platform
C++ · C★ 23.7K12 ago 2026
Licencia propia12 ago 2026 · métricas 2.10.0
npm
90Excelenteíndice de salud
colinhacks/zod
TypeScript-first schema validation with static type inference
TypeScript★ 43.4K↓ 992M/mes4 ago 2026
MIT4 ago 2026 · métricas 2.10.0