All tags
Catalogue tag

#inference

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

54 records
Tagged “inference”Ranked by health index
PyPI · crates.io
88Excellenthealth index
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,514Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
PyPI
88Excellenthealth index
huggingface/transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Python★ 162.7KJul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 1.13.0
crates.io · PyPI
88Excellenthealth index
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 86.1K↓ 0/moJul 13, 2026
Apache-2.0Jul 13, 2026 · metrics 1.13.0
PyPI · crates.io
86Excellenthealth index
apache/tvm-ffi
Open ABI and FFI for Machine Learning Systems
C++ · Python★ 435↓ 7M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI · RubyGems
85Excellenthealth index
deepspeedai/DeepSpeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Python · C++★ 42.7K↓ 1.3M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI
84Goodhealth index
mozilla-ai/any-llm
Communicate with an LLM provider using a single interface
Python★ 2,133↓ 130.6K/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI
84Goodhealth index
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 2,909Jul 21, 2026
Custom licenseJul 21, 2026 · metrics 1.13.0
PyPI
84Goodhealth index
triton-inference-server/server
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Python · Shell · C++★ 10.9KJul 19, 2026
BSD-3-ClauseJul 19, 2026 · metrics 1.13.0
PyPI
82Goodhealth index
huggingface/optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 3,448Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI · npm · RubyGems
82Goodhealth index
mlc-ai/xgrammar
Fast, Flexible and Portable Structured Generation
C++ · Python★ 1,792↓ 6.6M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
crates.io · PyPI
82Goodhealth index
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 30.3K↓ 274M/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
npm
80Goodhealth index
colinhacks/zod
TypeScript-first schema validation with static type inference
TypeScript★ 43.3K↓ 939M/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
PyPI · npm
80Goodhealth index
google/mediapipe
Cross-platform, customizable ML solutions for live and streaming media.
C++★ 36.1K↓ 2.8M/moJul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 1.13.0
crates.io · PyPI
78Goodhealth index
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/moJul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 1.13.0
Go
77Goodhealth index
utkuozdemir/nvidia_gpu_exporter
Nvidia GPU exporter for prometheus using nvidia-smi binary
Go★ 1,510Jul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
Go
76Goodhealth index
run-ai/karta
Translation layer that maps any Kubernetes framework's Custom Resource Definitions (CRDs) into a standardized, generic structure.
Go★ 55Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
74Goodhealth index
NVIDIA/nvcf
Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
Go · Rust★ 180Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
Go
74Goodhealth index
defilantech/llmkube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
Go
71Goodhealth index
NexusGPU/tensor-fusion
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 158Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
PyPI
71Goodhealth index
OpenNMT/CTranslate2
Fast inference engine for Transformer models
C++ · Python★ 4,577↓ 9.4M/moJul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
PyPI · crates.io · npm
71Goodhealth index
alexsjones/llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Rust★ 29.6K↓ 15.8K/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI · npm
71Goodhealth index
roboflow/inference
Turn any computer or edge device into a command center for your computer vision projects.
Python★ 2,374↓ 87.6K/moJul 14, 2026
Custom licenseJul 14, 2026 · metrics 1.13.0
PyPI
70Goodhealth index
NVIDIA/TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
C++★ 13.2K↓ 930.7K/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
Go
70Goodhealth index
nexusgpu/tensor-fusion-operator
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 158Jul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 1.13.0
crates.io · PyPI
69Moderatehealth index
StarlightSearch/EmbedAnything
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
Rust★ 1,281↓ 13.2K/moJul 14, 2026
Apache-2.0Jul 14, 2026 · metrics 1.13.0
Packagist
68Moderatehealth index
symfony/ai-platform
PHP library for interacting with AI platform provider.
PHP★ 52↓ 197.4K/moJul 15, 2026
MITJul 15, 2026 · metrics 1.13.0
PyPI
67Moderatehealth index
mahimailabs/voicegateway
The open-source profiler for voice agents
Python★ 9↓ 2,196/moJul 19, 2026
MITJul 19, 2026 · metrics 1.13.0
PyPI
63Moderatehealth index
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
PyPI
62Moderatehealth index
NevermindNilas/Nelux
NeLux: High-performance video processing for Python, powered by FFmpeg and PyTorch. Ultra-fast video decoding directly to tensors.
C++ · Python · Cuda★ 25↓ 3,141/moJul 19, 2026
AGPL-3.0Jul 19, 2026 · metrics 1.13.0
npm · crates.io
62Moderatehealth index
sipemu/anofox-regression
Regression analysis in Rust.
Rust★ 5↓ 6,781/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0