PyPI99Exceptionalhealth index
C++ · C★ 122.7KAug 4, 2026
PyPI97Exceptionalhealth index
intel/auto-roundA SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
Python · C++★ 1,520Jul 16, 2026
npm95Exceptionalhealth index
TypeScript★ 2,503↓ 22.4M/moAug 27, 2026
PyPI · npm94Exceptionalhealth index

modelscope/FunASROpen-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Python · C · HTML★ 19.6KAug 5, 2026
Go89Excellenthealth index

defilantech/LLMKubeKubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 207Sep 5, 2026
crates.io · PyPI · npm87Excellenthealth index

AlexsJones/llmfitHundreds of models & providers. One command to find what runs on your hardware.
Rust★ 31.1K↓ 2,251/moAug 5, 2026
npm87Excellenthealth index

withcatai/node-llama-cppRun AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 2,162↓ 3.6M/moAug 27, 2026
PyPI86Excellenthealth index
MakazhanAlpamys/SoupSoup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.
Python★ 74Jul 15, 2026
PyPI · npm86Excellenthealth index

n24q02m/mcp-coreShared foundation for building MCP servers -- Streamable HTTP transport, OAuth 2.1, browser-based credential setup, and a shared embedding daemon.
Python · TypeScript★ 1↓ 17.7K/moAug 22, 2026
crates.io · Maven84Excellenthealth index

eugenehp/llama-cpp-rsA wrapper around the llama-cpp library for rust, including new Sampler API from llama-cpp.
Rust★ 46↓ 7,448/moAug 22, 2026
Go81Excellenthealth index
Go★ 295Sep 5, 2026
Go · npm80Excellenthealth index
kdeps/kdepsRun AI workflows locally. Or deploy them anywhere. AI agent framework in YAML — workflow pipelines + autonomous agent loop. NVIDIA Inception member. Build, deploy, export as Docker/K8s/ISO.
Go★ 35Jul 24, 2026
Python★ 2↓ 2,506/moAug 22, 2026
crates.io77Goodhealth index
Rust★ 6↓ 3,633/moAug 5, 2026
hybridgroup/yzmaGo with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
Go★ 520Jul 17, 2026

asher/mlx-kquantNative K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/moAug 23, 2026
PyPI · crates.io67Goodhealth index
C++ · C★ 197↓ 7,559/moJul 17, 2026

wundercorp/openmodelUse any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.
TypeScript · JavaScript · CSS★ 18↓ 3,140/moAug 7, 2026
aimd54/palanPull, push, pack, and serve GGUF models as OCI ModelPack artifacts — daemonless, air-gap-first, one binary
Go★ 0Jul 17, 2026
Go · PyPI65Goodhealth index
anthony-chaudhary/fakfak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
PyPI · crates.io63Moderatehealth index

FedericoTs/quantprobeRun a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/moAug 20, 2026
npm62Moderatehealth index
therealtimex/node-llama-cppRun AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 1↓ 41.2K/moJul 17, 2026
eastriverlee/LLM.swiftLLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
Swift★ 866Jul 28, 2026
Maven · npm60Moderatehealth index
Java★ 3Sep 4, 2026
npm56Moderatehealth index
mohitsoni48/TurboLLMRun any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
TypeScript★ 173↓ 6,179/moJul 16, 2026
PyPI54Moderatehealth index
jjang-ai/jangqJANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 214Jul 22, 2026
PyPI · npm · crates.io53Moderatehealth index
JavaScript · Rust★ 2↓ 3,094/moSep 4, 2026
npm · crates.io · Maven51Moderatehealth index
Makefile · C · C++★ 7↓ 495/moJul 31, 2026
Python★ 60Aug 31, 2026