PyPI99Exceptionalhealth index
C++ · Python★ 27.8KAug 5, 2026
PyPI95Exceptionalhealth index

Blaizzy/mlx-audioA text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Python★ 7,796↓ 582.2K/moAug 28, 2026
PyPI95Exceptionalhealth index

Blaizzy/mlx-vlmMLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Python★ 5,431↓ 881.2K/moAug 28, 2026
PyPI · Maven95Exceptionalhealth index
maziyarpanahi/openmedLocal-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud, no patient data leaving your network. Apache-2.0
Python★ 4,827↓ 4M/moAug 4, 2026
PyPI95Exceptionalhealth index
Python★ 6,813↓ 1.2M/moAug 28, 2026
PyPI93Exceptionalhealth index
raullenchai/Rapid-MLXThe fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Python★ 3,392Aug 2, 2026
Go89Excellenthealth index

defilantech/LLMKubeKubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 207Sep 5, 2026
crates.io · PyPI · npm87Excellenthealth index

AlexsJones/llmfitHundreds of models & providers. One command to find what runs on your hardware.
Rust★ 31.1K↓ 2,251/moAug 5, 2026
PyPI86Excellenthealth index

arogozhnikov/einopsFlexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)
Python · Jupyter Notebook★ 9,582↓ 27.1M/moAug 28, 2026
PyPI · npm86Excellenthealth index
Python · C++★ 1,683↓ 817/moJul 18, 2026
PyPI · npm84Excellenthealth index

rajveer43/VeloxQuant-MLXFast KV-cache quantization for Apple Silicon (MLX) — 43 research-adapted compression methods with Metal kernels
Python★ 15↓ 7,549/moSep 5, 2026
PyPI · npm81Excellenthealth index
youssofal/MTPLX3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
Python · Swift★ 1,120↓ 2,694/moAug 2, 2026
PyPI80Excellenthealth index
Python★ 83↓ 1,116/moAug 28, 2026

jundot/omlxLLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Python · Swift★ 18.5KAug 5, 2026
Python★ 15↓ 2,711/moJul 18, 2026
jagoff/memoPersistent semantic memory for AI agents — 100% local on Apple Silicon (MLX) or Linux/Ubuntu (CPU). Markdown source of truth, sqlite-vec + BM25 hybrid search, a codegraph-backed knowledge graph, MCP server + CLI. No cloud, no keys.
Python★ 7↓ 3,119/moJul 22, 2026
npm · crates.io75Goodhealth index
Rust · TypeScript★ 152↓ 2,518/moAug 3, 2026
binlecode/actopApple Silicon (M1–M4) power, GPU, ANE & memory-bandwidth monitor — sudoless TUI + Python API for profiling local LLM / MLX / CoreML inference
Python★ 0↓ 2,580/moJul 29, 2026

asher/mlx-kquantNative K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/moAug 23, 2026
PyPI · npm69Goodhealth index
jjang-ai/vmlxvMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4,850/moJul 24, 2026
codam-coding-college/MLX42Codam's own fixed, functioning and open source alternative of the miniLibX. MLX42 is a simple cross-platform graphics library running on GLFW and OpenGL.
C★ 476Jul 23, 2026
PyPI63Moderatehealth index
druide67/asiaiMulti-engine LLM benchmark & monitoring CLI for Apple Silicon
Python★ 11↓ 4,662/moJul 18, 2026
npm63Moderatehealth index
TypeScript★ 14↓ 11.2K/moJul 26, 2026
crates.io · PyPI57Moderatehealth index

jbg/eredua Rust runtime for local language models
Rust★ 3↓ 16.5K/moSep 5, 2026
npm · PyPI · crates.io56Moderatehealth index
TaeSooPark-PTS/LatticeAILocal-first private AI memory layer / Digital Brain for conversations, documents, decisions, and model-agnostic knowledge.
Python · TypeScript★ 1↓ 5,135/moJul 19, 2026
PyPI54Moderatehealth index
Python★ 44Jul 31, 2026
PyPI54Moderatehealth index
jjang-ai/jangqJANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 214Jul 22, 2026
PyPI54Moderatehealth index
Python · Jupyter Notebook★ 8,876Aug 12, 2026
PyPI53Moderatehealth index

ARahim3/mlx-dsparkUp to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, Ornith-1.0, ternary Bonsai-27B.
Python · Swift★ 430↓ 5,564/moAug 19, 2026
PyPI · npm53Moderatehealth index
Python★ 0↓ 6,042/moJul 18, 2026