All tags
Catalogue tag

#inference

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

106 records
Tagged “inference”Ranked by health index
Go
89Excellenthealth index
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 207Sep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
PyPI
88Excellenthealth index
Andyyyy64/whichllm
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Python★ 6,463Aug 24, 2026
MITAug 24, 2026 · metrics 2.10.0
PyPI
88Excellenthealth index
NVIDIA/TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
C++★ 13.3K↓ 1M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
npm · PyPI
88Excellenthealth index
google-ai-edge/mediapipe
Cross-platform, customizable ML solutions for live and streaming media.
C++★ 36.5KAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
crates.io
88Excellenthealth index
praxis-proxy/praxis
AI and cloud-native proxy server and framework
Rust★ 71↓ 10.7K/moAug 15, 2026
MITAug 15, 2026 · metrics 2.10.0
crates.io · PyPI · npm
87Excellenthealth index
AlexsJones/llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Rust★ 31.1K↓ 2,251/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
crates.io · npm · PyPI
87Excellenthealth index
EricLBuehler/mistral.rs
Fast, flexible LLM inference
Rust · Cuda★ 7,586↓ 259.3K/moAug 12, 2026
MITAug 12, 2026 · metrics 2.10.0
npm
87Excellenthealth index
anolilab/lunora
Type-safe, real-time backend framework on your own Cloudflare account — Workers, Durable Objects, D1, R2, Queues. Convex-style DX, Vite-first.
TypeScript★ 261↓ 58.8K/moSep 5, 2026
Custom licenseSep 5, 2026 · metrics 2.10.0
crates.io
87Excellenthealth index
paiml/aprender
Next Generation Machine Learning, Statistics and Deep Learning in PURE Rust
Rust · HTML★ 110↓ 3,595/moJul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
crates.io · PyPI
87Excellenthealth index
typedb/typeql
TypeQL: Built for systems, not records
Rust★ 254↓ 3,889/moAug 16, 2026
MPL-2.0Aug 16, 2026 · metrics 2.10.0
Go
86Excellenthealth index
NexusGPU/tensor-fusion
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 158Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
OpenNMT/CTranslate2
Fast inference engine for Transformer models
C++ · Python★ 4,645↓ 12.4M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI · npm
86Excellenthealth index
n24q02m/mcp-core
Shared foundation for building MCP servers -- Streamable HTTP transport, OAuth 2.1, browser-based credential setup, and a shared embedding daemon.
Python · TypeScript★ 1↓ 17.7K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
Packagist · PyPI
84Excellenthealth index
cognesy/instructor-php
Unified LLM API, structured data outputs with LLMs, and agent SDK - in PHP
PHP★ 325↓ 5,250/moJul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
crates.io · Maven
84Excellenthealth index
eugenehp/llama-cpp-rs
A wrapper around the llama-cpp library for rust, including new Sampler API from llama-cpp.
Rust★ 46↓ 7,448/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
Go
84Excellenthealth index
nexusgpu/tensor-fusion-operator
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 158Jul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 2.10.0
PyPI · npm
84Excellenthealth index
rajveer43/VeloxQuant-MLX
Fast KV-cache quantization for Apple Silicon (MLX) — 43 research-adapted compression methods with Metal kernels
Python★ 15↓ 7,549/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
crates.io · PyPI
81Excellenthealth index
StarlightSearch/EmbedAnything
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
Rust★ 1,304↓ 13.2K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
mahimailabs/voicegateway
The open-source profiler for voice agents
Python★ 9↓ 2,196/moJul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
mobilint/mblt-model-zoo
Mobilint Model Zoo Project
Python★ 23Aug 1, 2026
BSD-3-ClauseAug 1, 2026 · metrics 2.10.0
Packagist
80Excellenthealth index
symfony/ai-platform
PHP library for interacting with AI platform provider.
PHP★ 54↓ 212.3K/moSep 6, 2026
MITSep 6, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2,506/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
crates.io · npm · PyPI
78Goodhealth index
ohdearquant/lattice
Run, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust · Python★ 40↓ 19.1K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
PyPI · crates.io
77Goodhealth index
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI
75Goodhealth index
cozy-creator/python-gen-worker
A collection of worker runtimes (with Dockerfiles) + worker functions, to be used by the gen-orchestrator.
Python★ 0↓ 18.3K/moSep 6, 2026
MITSep 6, 2026 · metrics 2.10.0
npm · PyPI
75Goodhealth index
nicolasmelo1/logion
Agent-native course marketplace and skill registry for executable AI-agent curricula
Python★ 33↓ 5,752/moJul 24, 2026
MITJul 24, 2026 · metrics 2.10.0
npm · crates.io
73Goodhealth index
sipemu/anofox-regression
Regression analysis in Rust.
Rust★ 5↓ 6,781/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
npm · crates.io · PyPI
71Goodhealth index
Rust★ 13↓ 7,573/moSep 6, 2026
No licenseSep 6, 2026 · metrics 2.10.0
crates.io
71Goodhealth index
ai-dynamo/frontend-crates
No repository description published.
Rust · Python★ 16↓ 271.8K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
npm · Maven
71Goodhealth index
hung-yueh/react-native-litert-lm
High-performance on-device LLM inference for React Native, powered by LiteRT-LM and Nitro Modules
C++ · TypeScript · C★ 49↓ 2,239/moJul 24, 2026
MITJul 24, 2026 · metrics 2.10.0