PyPI · crates.io88Excellenthealth index
ai-dynamo/dynamoA Datacenter Scale Distributed Inference Serving Framework
Rust · Python · Go★ 7,514Jul 17, 2026
PyPI88Excellenthealth index
huggingface/transformers🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Python★ 162.7KJul 20, 2026
crates.io · PyPI88Excellenthealth index
vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 86.1K↓ 0/moJul 13, 2026
PyPI · crates.io86Excellenthealth index
apache/tvm-ffiOpen ABI and FFI for Machine Learning Systems
C++ · Python★ 435↓ 7M/moJul 21, 2026
PyPI · RubyGems85Excellenthealth index
deepspeedai/DeepSpeedDeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Python · C++★ 42.7K↓ 1.3M/moJul 21, 2026
mozilla-ai/any-llmCommunicate with an LLM provider using a single interface
Python★ 2,133↓ 130.6K/moJul 21, 2026
pytorch/aoPyTorch native quantization and sparsity for training and inference
Python · C++★ 2,909Jul 21, 2026
triton-inference-server/serverThe Triton Inference Server provides an optimized cloud and edge inferencing solution.
Python · Shell · C++★ 10.9KJul 19, 2026
huggingface/optimum🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Python★ 3,448Jul 21, 2026
PyPI · npm · RubyGems82Goodhealth index
mlc-ai/xgrammarFast, Flexible and Portable Structured Generation
C++ · Python★ 1,792↓ 6.6M/moJul 21, 2026
crates.io · PyPI82Goodhealth index
sgl-project/sglangSGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 30.3K↓ 274M/moJul 14, 2026
colinhacks/zodTypeScript-first schema validation with static type inference
TypeScript★ 43.3K↓ 939M/moJul 22, 2026
PyPI · npm80Goodhealth index
google/mediapipeCross-platform, customizable ML solutions for live and streaming media.
C++★ 36.1K↓ 2.8M/moJul 15, 2026
crates.io · PyPI78Goodhealth index
lightseekorg/smgEngine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Rust · Python★ 398↓ 39/moJul 16, 2026
utkuozdemir/nvidia_gpu_exporterNvidia GPU exporter for prometheus using nvidia-smi binary
Go★ 1,510Jul 22, 2026
run-ai/kartaTranslation layer that maps any Kubernetes framework's Custom Resource Definitions (CRDs) into a standardized, generic structure.
Go★ 55Jul 18, 2026
NVIDIA/nvcfPlatform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
Go · Rust★ 180Jul 17, 2026
defilantech/llmkubeKubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 167↓ 0/moJul 14, 2026
NexusGPU/tensor-fusionTensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 158Jul 17, 2026
OpenNMT/CTranslate2Fast inference engine for Transformer models
C++ · Python★ 4,577↓ 9.4M/moJul 21, 2026
PyPI · crates.io · npm71Goodhealth index
alexsjones/llmfitHundreds of models & providers. One command to find what runs on your hardware.
Rust★ 29.6K↓ 15.8K/moJul 17, 2026
PyPI · npm71Goodhealth index
roboflow/inferenceTurn any computer or edge device into a command center for your computer vision projects.
Python★ 2,374↓ 87.6K/moJul 14, 2026
NVIDIA/TensorRTNVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
C++★ 13.2K↓ 930.7K/moJul 21, 2026
nexusgpu/tensor-fusion-operatorTensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 158Jul 15, 2026
crates.io · PyPI69Moderatehealth index
StarlightSearch/EmbedAnythingHighly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
Rust★ 1,281↓ 13.2K/moJul 14, 2026
Packagist68Moderatehealth index
symfony/ai-platformPHP library for interacting with AI platform provider.
PHP★ 52↓ 197.4K/moJul 15, 2026
PyPI67Moderatehealth index
mahimailabs/voicegatewayThe open-source profiler for voice agents
Python★ 9↓ 2,196/moJul 19, 2026
PyPI63Moderatehealth index
varjoranta/turboquant-vllmTurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fallback.
Python · C++ · Cuda★ 76↓ 960/moJul 22, 2026
PyPI62Moderatehealth index
NevermindNilas/NeluxNeLux: High-performance video processing for Python, powered by FFmpeg and PyTorch. Ultra-fast video decoding directly to tensors.
C++ · Python · Cuda★ 25↓ 3,141/moJul 19, 2026
npm · crates.io62Moderatehealth index
sipemu/anofox-regressionRegression analysis in Rust.
Rust★ 5↓ 6,781/moJul 17, 2026