PyPI · Maven95Exceptionalhealth index
maziyarpanahi/openmedLocal-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud, no patient data leaving your network. Apache-2.0
Python★ 4,827↓ 4M/moAug 4, 2026
PyPI93Exceptionalhealth index
raullenchai/Rapid-MLXThe fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Python★ 3,392Aug 2, 2026
Go89Excellenthealth index

defilantech/LLMKubeKubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 207Sep 5, 2026
PyPI88Excellenthealth index

Andyyyy64/whichllmFind the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Python★ 6,463Aug 24, 2026
npm87Excellenthealth index
TypeScript★ 14↓ 3,502/moSep 6, 2026
PyPI86Excellenthealth index
MakazhanAlpamys/SoupSoup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.
Python★ 74Jul 15, 2026
npm83Excellenthealth index
HybridAIOne/hybridclawEnterprise-ready self-hosted AI assistant runtime with sandboxed execution, secure credentials, approvals, and memory
TypeScript★ 126↓ 2,721/moAug 4, 2026
PyPI83Excellenthealth index

skynetcmd/m3-memoryLocal-first Memory Framework for AI Agents · 99.2% LongMemEval-S retrieval @ k=10 · Supports Claude · Antigravity · LangChain · Hermes · Gemini · OpenCode · OpenClaw · MCP-native and plugins · Hybrid search (FTS5 + vector + MMR) · GDPR · FIPS 140-3 ready · 100% local (fully offline) or cloud capable
Python★ 19↓ 8,178/moAug 9, 2026
PyPI81Excellenthealth index
jonigl/mcp-client-for-ollamaHarness the power of local LLMs with this TUI MCP Client for Ollama. Featuring all core MCP primitives (tools, prompts, resources), agent mode, multi-server, model switching, streaming responses, human-in-the-loop, thinking mode, model params config, system prompts, and saved preferences.
Python★ 782↓ 15.1K/moJul 24, 2026
npm81Excellenthealth index
manojmallick/sigmap97% token reduction for AI coding sessions — zero deps, 33 languages, MCP server
JavaScript★ 598↓ 10.8K/moJul 18, 2026
npm · Maven81Excellenthealth index
C++ · C★ 1,000↓ 57.2K/moJul 17, 2026
Go · npm · PyPI81Excellenthealth index
orneryd/NornicDBNornicdb is a distributed low-latency, Graph+Vector, Temporal MVCC with all sub-ms HNSW search, graph traversal, and writes. Using Neo4j Bolt/Cypher and qdrant's gRPC means you can switch with no changes while adding intelligent features like schemas, managed embeddings, reranking+llm, GPU accel, Auto-TLP, Policy-based Memory Decay, and MCP server.
Go★ 833Jul 27, 2026
npm80Excellenthealth index
dcostenco/prism-coderPersistent memory + local AI for coding agents. 2B–27B open-weight LLM fleet, cross-session Mind Palace, cognitive routing, L3 grounding verifier, multi-agent Hivemind. Works with Claude Code, Cursor, VS Code. Offline-first, HIPAA-ready. Free tier included.
TypeScript★ 155↓ 5,205/moJul 18, 2026
Go80Excellenthealth index
ionalpha/flynnA secure, self-improving agent operating system in a single Go binary. Bring any model, manage local models, point it at a goal, and grant it real authority: every action is sandboxed, governed, and sealed into a verifiable, tamper-evident record an independent party can check. Runs interactive or 24/7, or embed it in your own system.
Go★ 2Jul 18, 2026
crates.io · npm · PyPI78Goodhealth index

ohdearquant/latticeRun, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust · Python★ 40↓ 19.1K/moAug 22, 2026

KevRojo/DulusDulus Ai — Free Agentic AI, Making Gemini web cappable of running bash commands in your terminal! [Gui, Web, Cli, Telegram, 2,000 MCP, 100K Skills . LiteLLM (100+ providers), local models via Ollama, /lang in 34 languages, Mesa Redonda, I create the first utility coin that can be used 100% as AI quota or Fuel, is called $Dulus
Python★ 358↓ 7,957/moSep 5, 2026
TypeScript · JavaScript★ 6↓ 15.1K/moJul 26, 2026
npm · crates.io75Goodhealth index
Rust · TypeScript★ 152↓ 2,518/moAug 3, 2026
ASCIT31/Dark-MoonAutonomous AI pentesting engine, continuous offensive security across web, cloud, identity, CI/CD, IaC, databases, Active Directory, Kubernetes and IoT firmware. Agentic reasoning plus real exploit execution deliver proof-based vulnerabilities. Privacy gateway: the LLM never sees your real IPs, hosts or creds, nothing leaves your perimeter.
Python · TypeScript · Shell★ 802Aug 4, 2026
C++ · JavaScript · TypeScript★ 20↓ 1,947/moJul 25, 2026
TypeScript★ 75↓ 15.9K/moJul 15, 2026

asher/mlx-kquantNative K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/moAug 23, 2026
TypeScript · Python · JavaScript★ 1,830↓ 3,860/moJul 23, 2026
m62624/pi-code-plannerStructured planning, bounded memory, TDD, and Git guardrails for local coding models in Pi Code (I don't know TypeScript at all; this is mostly a local-model experiment, with occasional help from Claude Code)
TypeScript★ 4↓ 3,041/moJul 22, 2026
famclaw/famclawSelf-hosted family AI gateway with parental controls. Runs on Linux, macOS, and Android (Termux) — Raspberry Pi, mini PC, old laptop, homelab server, even a phone. Telegram, Discord, web. Privacy-first, works with any LLM (local or cloud), OPA content filtering, MCP skill scanning.
Go★ 2Jul 17, 2026
rtmx-ai/aegis-cliAir-gap-native agentic coding for closed environments — a hardened OpenCode TUI driven by a local model, with rtmx as the intent layer. Zero egress by construction.
Go★ 4Jul 20, 2026

wundercorp/openmodelUse any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.
TypeScript · JavaScript · CSS★ 18↓ 3,140/moAug 7, 2026
Go · PyPI65Goodhealth index
anthony-chaudhary/fakfak — the Fused Agent Kernel: one Go binary for AI agent loops. Wrap Claude Code/Codex/Cursor, keep long sessions cache-efficient, route work per call, run local GGUF models, and adjudicate tool calls.
Go · Python★ 15Jul 17, 2026
PyPI · crates.io63Moderatehealth index

FedericoTs/quantprobeRun a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/moAug 20, 2026

shotah/ai-gantryPersonal AI agent you can actually own: one static Go binary, one persona, any OpenAI-compat LLM (Ollama, Gemini, Grok), MCP tools, chat via Telegram/Discord/Slack. Outbound-only — no dashboard, no config UI, no open ports, ever. Hardened so small local models actually finish tool calls.
Go★ 0Aug 11, 2026