
defilantech/LLMKubeKubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 2072026年9月5日

Andyyyy64/whichllmFind the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Python★ 6,4632026年8月24日

NVIDIA/TensorRTNVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
C++★ 13.3K↓ 1M/月2026年8月27日
C++★ 36.5K2026年8月5日
Rust★ 71↓ 10.7K/月2026年8月15日
crates.io · PyPI · npm87优秀健康指数

AlexsJones/llmfitHundreds of models & providers. One command to find what runs on your hardware.
Rust★ 31.1K↓ 2,251/月2026年8月5日
crates.io · npm · PyPI87优秀健康指数
Rust · Cuda★ 7,586↓ 259.3K/月2026年8月12日

anolilab/lunoraType-safe, real-time backend framework on your own Cloudflare account — Workers, Durable Objects, D1, R2, Queues. Convex-style DX, Vite-first.
TypeScript★ 261↓ 58.8K/月2026年9月5日
paiml/aprenderNext Generation Machine Learning, Statistics and Deep Learning in PURE Rust
Rust · HTML★ 110↓ 3,595/月2026年7月28日
Rust★ 254↓ 3,889/月2026年8月16日
NexusGPU/tensor-fusionTensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 1582026年7月17日
C++ · Python★ 4,645↓ 12.4M/月2026年8月27日

n24q02m/mcp-coreShared foundation for building MCP servers -- Streamable HTTP transport, OAuth 2.1, browser-based credential setup, and a shared embedding daemon.
Python · TypeScript★ 1↓ 17.7K/月2026年8月22日
PHP★ 325↓ 5,250/月2026年7月20日
crates.io · Maven84优秀健康指数

eugenehp/llama-cpp-rsA wrapper around the llama-cpp library for rust, including new Sampler API from llama-cpp.
Rust★ 46↓ 7,448/月2026年8月22日
nexusgpu/tensor-fusion-operatorTensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
Go★ 1582026年7月15日

rajveer43/VeloxQuant-MLXFast KV-cache quantization for Apple Silicon (MLX) — 43 research-adapted compression methods with Metal kernels
Python★ 15↓ 7,549/月2026年9月5日
Rust★ 1,304↓ 13.2K/月2026年8月22日
Python★ 9↓ 2,196/月2026年7月19日
Python★ 232026年8月1日
PHP★ 54↓ 212.3K/月2026年9月6日
Python★ 2↓ 2,506/月2026年8月22日
crates.io · npm · PyPI78良好健康指数

ohdearquant/latticeRun, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust · Python★ 40↓ 19.1K/月2026年8月22日

ahb-sjsu/turboquant-proConsumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the downstream consumer actually uses. PCA-Matryoshka + TurboQuant (27x @ 99.8% recall@10), asymmetric K/V, CUDA/Triton kernels, vLLM plugin, replayable CI-gated claims. MIT.
Python★ 25↓ 277/月2026年9月5日
Python★ 0↓ 18.3K/月2026年9月6日
nicolasmelo1/logionAgent-native course marketplace and skill registry for executable AI-agent curricula
Python★ 33↓ 5,752/月2026年7月24日
Rust★ 5↓ 6,781/月2026年7月17日
npm · crates.io · PyPI71良好健康指数
Rust★ 13↓ 7,573/月2026年9月6日
Rust · Python★ 16↓ 271.8K/月2026年8月22日
C++ · TypeScript · C★ 49↓ 2,239/月2026年7月24日