All tags
Catalogue tag

#inference

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

106 records
Tagged “inference”Ranked by health index
PyPI · crates.io
99Exceptionalhealth index
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,886↓ 59.4K/moAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
PyPI
99Exceptionalhealth index
huggingface/transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Python★ 163.3K↓ 179M/moAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI · crates.io
99Exceptionalhealth index
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 88.2KAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI · Go
98Exceptionalhealth index
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Python★ 11.2K↓ 98.8K/moAug 15, 2026
Apache-2.0Aug 15, 2026 · metrics 2.10.0
PyPI · crates.io
98Exceptionalhealth index
apache/tvm-ffi
Open ABI and FFI for Machine Learning Systems
C++ · Python · Rust★ 452↓ 7.9M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI
98Exceptionalhealth index
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6,256Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI · crates.io
98Exceptionalhealth index
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 31.3K↓ 144M/moAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI
98Exceptionalhealth index
triton-inference-server/server
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Python · Shell · C++★ 10.9KAug 27, 2026
BSD-3-ClauseAug 27, 2026 · metrics 2.10.0
PyPI · RubyGems
97Exceptionalhealth index
deepspeedai/DeepSpeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Python · C++★ 42.9K↓ 1.1M/moAug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
97Exceptionalhealth index
pytorch/ao
PyTorch native quantization and sparsity for training and inference
Python · C++★ 2,909Jul 21, 2026
Custom licenseJul 21, 2026 · metrics 2.10.0
PyPI
96Exceptionalhealth index
mozilla-ai/any-llm
Communicate with an LLM provider using a single interface
Python★ 2,133↓ 130.6K/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
npm
96Exceptionalhealth index
vercel/ai
The AI Toolkit for TypeScript. From the creators of Next.js, the AI SDK is a free open-source library for building AI-powered applications and agents
TypeScript · MDX★ 26K↓ 87M/moAug 5, 2026
Custom licenseAug 5, 2026 · metrics 2.10.0
PyPI · crates.io
95Exceptionalhealth index
ai-dynamo/aiconfigurator
Offline optimization of your disaggregated Dynamo graph
Python · Rust★ 374Jul 26, 2026
Apache-2.0Jul 26, 2026 · metrics 2.10.0
crates.io
95Exceptionalhealth index
ai-dynamo/modelexpress
Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
Python · Rust★ 103↓ 246.6K/moAug 1, 2026
Apache-2.0Aug 1, 2026 · metrics 2.10.0
PyPI
95Exceptionalhealth index
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Python★ 5,422↓ 2,132/moAug 2, 2026
Apache-2.0Aug 2, 2026 · metrics 2.10.0
npm
95Exceptionalhealth index
huggingface/huggingface.js
Use Hugging Face with JavaScript
TypeScript★ 2,503↓ 22.4M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
huggingface/optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 3,448Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
crates.io · npm · PyPI
94Exceptionalhealth index
pykeio/ort
Fast ML inference & training for ONNX models in Rust
Rust★ 2,477↓ 4.3M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
npm · RubyGems
94Exceptionalhealth index
vega/vega
A visualization grammar.
JavaScript · TypeScript★ 12K↓ 28.4M/moAug 27, 2026
BSD-3-ClauseAug 27, 2026 · metrics 2.10.0
PyPI · npm · RubyGems
93Exceptionalhealth index
mlc-ai/xgrammar
Fast, Flexible and Portable Structured Generation
C++ · Python★ 1,845↓ 8.2M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI
93Exceptionalhealth index
pytorch/TensorRT
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
Python · Jupyter Notebook · C++★ 2,984Jul 30, 2026
BSD-3-ClauseJul 30, 2026 · metrics 2.10.0
PyPI
93Exceptionalhealth index
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Python★ 3,392Aug 2, 2026
Apache-2.0Aug 2, 2026 · metrics 2.10.0
Go
93Exceptionalhealth index
run-ai/karta
Translation layer that maps any Kubernetes framework's Custom Resource Definitions (CRDs) into a standardized, generic structure.
Go★ 60Aug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
crates.io · PyPI
92Excellenthealth index
lightseekorg/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/moJul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 2.10.0
PyPI · npm
92Excellenthealth index
roboflow/inference
Turn any computer or edge device into a command center for your computer vision projects.
Python★ 2,438Sep 5, 2026
Custom licenseSep 5, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
triton-inference-server/model_analyzer
Triton Model Analyzer is a CLI tool to help with better understanding of the compute and memory requirements of the Triton Inference Server models.
Python★ 523↓ 10.6K/moJul 29, 2026
Apache-2.0Jul 29, 2026 · metrics 2.10.0
Go
91Excellenthealth index
utkuozdemir/nvidia_gpu_exporter
Nvidia GPU exporter for prometheus using nvidia-smi binary
Go★ 1,510Jul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
90Excellenthealth index
NVIDIA/nvcf
Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
Go · Rust★ 180Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI
90Excellenthealth index
Tencent/ncnn
ncnn is a high-performance neural network inference framework optimized for the mobile platform
C++ · C★ 23.7KAug 12, 2026
Custom licenseAug 12, 2026 · metrics 2.10.0
npm
90Excellenthealth index
colinhacks/zod
TypeScript-first schema validation with static type inference
TypeScript★ 43.4K↓ 992M/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0