Todas las etiquetas
Etiqueta del catálogo

#inference

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

54 registros
Con la etiqueta «inference»Ordenado por índice de salud
PyPI · crates.io
88Excelenteíndice de salud
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 751417 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
PyPI
88Excelenteíndice de salud
huggingface/transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Python★ 162.7K20 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0
crates.io · PyPI
88Excelenteíndice de salud
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 86.1K↓ 0/mes13 jul 2026
Apache-2.013 jul 2026 · métricas 1.13.0
PyPI · crates.io
86Excelenteíndice de salud
apache/tvm-ffi
Open ABI and FFI for Machine Learning Systems
C++ · Python★ 435↓ 7M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI · RubyGems
85Excelenteíndice de salud
deepspeedai/DeepSpeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Python · C++★ 42.7K↓ 1.3M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI
84Buenoíndice de salud
mozilla-ai/any-llm
Communicate with an LLM provider using a single interface
Python★ 2133↓ 130.6K/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI
84Buenoíndice de salud
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 290921 jul 2026
Licencia propia21 jul 2026 · métricas 1.13.0
PyPI
84Buenoíndice de salud
triton-inference-server/server
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Python · Shell · C++★ 10.9K19 jul 2026
BSD-3-Clause19 jul 2026 · métricas 1.13.0
PyPI
82Buenoíndice de salud
huggingface/optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 344821 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI · npm · RubyGems
82Buenoíndice de salud
mlc-ai/xgrammar
Fast, Flexible and Portable Structured Generation
C++ · Python★ 1792↓ 6.6M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
crates.io · PyPI
82Buenoíndice de salud
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 30.3K↓ 274M/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
npm
80Buenoíndice de salud
colinhacks/zod
TypeScript-first schema validation with static type inference
TypeScript★ 43.3K↓ 939M/mes22 jul 2026
MIT22 jul 2026 · métricas 1.13.0
PyPI · npm
80Buenoíndice de salud
google/mediapipe
Cross-platform, customizable ML solutions for live and streaming media.
C++★ 36.1K↓ 2.8M/mes15 jul 2026
Apache-2.015 jul 2026 · métricas 1.13.0
crates.io · PyPI
78Buenoíndice de salud
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/mes16 jul 2026
Apache-2.016 jul 2026 · métricas 1.13.0
Go
77Buenoíndice de salud
utkuozdemir/nvidia_gpu_exporter
Nvidia GPU exporter for prometheus using nvidia-smi binary
Go★ 151022 jul 2026
MIT22 jul 2026 · métricas 1.13.0
Go
76Buenoíndice de salud
run-ai/karta
Translation layer that maps any Kubernetes framework's Custom Resource Definitions (CRDs) into a standardized, generic structure.
Go★ 5518 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
74Buenoíndice de salud
NVIDIA/nvcf
Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
Go · Rust★ 18017 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
Go
74Buenoíndice de salud
defilantech/llmkube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
Go
71Buenoíndice de salud
NexusGPU/tensor-fusion
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 15817 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI
71Buenoíndice de salud
OpenNMT/CTranslate2
Fast inference engine for Transformer models
C++ · Python★ 4577↓ 9.4M/mes21 jul 2026
MIT21 jul 2026 · métricas 1.13.0
PyPI · crates.io · npm
71Buenoíndice de salud
alexsjones/llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Rust★ 29.6K↓ 15.8K/mes17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
PyPI · npm
71Buenoíndice de salud
roboflow/inference
Turn any computer or edge device into a command center for your computer vision projects.
Python★ 2374↓ 87.6K/mes14 jul 2026
Licencia propia14 jul 2026 · métricas 1.13.0
PyPI
70Buenoíndice de salud
NVIDIA/TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
C++★ 13.2K↓ 930.7K/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
Go
70Buenoíndice de salud
nexusgpu/tensor-fusion-operator
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 15815 jul 2026
Apache-2.015 jul 2026 · métricas 1.13.0
crates.io · PyPI
69Moderadoíndice de salud
StarlightSearch/EmbedAnything
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
Rust★ 1281↓ 13.2K/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
Packagist
68Moderadoíndice de salud
symfony/ai-platform
PHP library for interacting with AI platform provider.
PHP★ 52↓ 197.4K/mes15 jul 2026
MIT15 jul 2026 · métricas 1.13.0
PyPI
67Moderadoíndice de salud
mahimailabs/voicegateway
The open-source profiler for voice agents
Python★ 9↓ 2196/mes19 jul 2026
MIT19 jul 2026 · métricas 1.13.0
PyPI
63Moderadoíndice de salud
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/mes22 jul 2026
MIT22 jul 2026 · métricas 1.13.0
PyPI
62Moderadoíndice de salud
NevermindNilas/Nelux
NeLux: High-performance video processing for Python, powered by FFmpeg and PyTorch. Ultra-fast video decoding directly to tensors.
C++ · Python · Cuda★ 25↓ 3141/mes19 jul 2026
AGPL-3.019 jul 2026 · métricas 1.13.0
npm · crates.io
62Moderadoíndice de salud
sipemu/anofox-regression
Regression analysis in Rust.
Rust★ 5↓ 6781/mes17 jul 2026
MIT17 jul 2026 · métricas 1.13.0