crates.io · npm96Exceptionalhealth index

hashintel/hash🚀 The open-source, multi-tenant platform for self-building knowledge graphs and simulation
TypeScript · Rust★ 1,641↓ 244.7K/moAug 16, 2026
PyPI95Exceptionalhealth index

Blaizzy/mlx-audioA text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Python★ 7,796↓ 582.2K/moAug 28, 2026
PyPI95Exceptionalhealth index

Blaizzy/mlx-vlmMLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Python★ 5,431↓ 881.2K/moAug 28, 2026
PyPI93Exceptionalhealth index
raullenchai/Rapid-MLXThe fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Python★ 3,392Aug 2, 2026
Go89Excellenthealth index

defilantech/LLMKubeKubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Go★ 207Sep 5, 2026
npm · crates.io86Excellenthealth index
drakulavich/kesha-voice-kitGive your tools a voice — speech to text and back, 25 languages, up to ~19× faster than Whisper. On your machine.
Rust · TypeScript★ 67↓ 1,029/moJul 28, 2026
PyPI · npm84Excellenthealth index

rajveer43/VeloxQuant-MLXFast KV-cache quantization for Apple Silicon (MLX) — 43 research-adapted compression methods with Metal kernels
Python★ 15↓ 7,549/moSep 5, 2026
npm83Excellenthealth index
achiya-automation/safari-mcpNative Safari browser automation for AI agents. 80 tools via AppleScript — zero overhead, keeps logins, runs silently in background. Drop-in alternative to Chrome DevTools MCP with 40-60% less CPU/heat on Apple Silicon.
JavaScript★ 151↓ 6,053/moJul 17, 2026
PyPI · npm81Excellenthealth index
youssofal/MTPLX3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
Python · Swift★ 1,120↓ 2,694/moAug 2, 2026
PyPI80Excellenthealth index
Python★ 83↓ 1,116/moAug 28, 2026
Arthur-Ficial/apfelThe free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API keys, no cloud, no downloads.
Swift · Python★ 6,240Aug 4, 2026

jundot/omlxLLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Python · Swift★ 18.5KAug 5, 2026
Python★ 15↓ 2,711/moJul 18, 2026
crates.io · npm · PyPI78Goodhealth index

ohdearquant/latticeRun, quantize, and fine-tune LLMs on Apple Silicon. Pure Rust, no Python, no CUDA, no ONNX
Rust · Python★ 40↓ 19.1K/moAug 22, 2026
michaelellis003/smcxState-space inference in JAX: Kalman and particle filters, tempered SMC, and SMC²
Python★ 3Jul 26, 2026
tillahoffmann/jax-mpsA JAX backend for Apple Metal Performance Shaders (MPS), enabling GPU-accelerated JAX computations on Apple Silicon.
C++ · Python★ 192↓ 3,492/moJul 18, 2026
arcships/light-ocrFast, offline OCR for Node.js & C++. PP-OCRv6 with Core ML / WebGPU hardware acceleration — recognize text in images with confidence scores & coordinates. npm: @arcships/light-ocr
C++ · JavaScript★ 455Jul 31, 2026
jagoff/memoPersistent semantic memory for AI agents — 100% local on Apple Silicon (MLX) or Linux/Ubuntu (CPU). Markdown source of truth, sqlite-vec + BM25 hybrid search, a codegraph-backed knowledge graph, MCP server + CLI. No cloud, no keys.
Python★ 7↓ 3,119/moJul 22, 2026
binlecode/actopApple Silicon (M1–M4) power, GPU, ANE & memory-bandwidth monitor — sudoless TUI + Python API for profiling local LLM / MLX / CoreML inference
Python★ 0↓ 2,580/moJul 29, 2026

asher/mlx-kquantNative K-quant support for MLX, with a quantization and fine-tuning toolchain for Apple Silicon
C++ · Python★ 6↓ 3,360/moAug 23, 2026

ooples/AiDotNet.TensorsThe fastest .NET tensor library. Beats MathNet (6x), NumSharp (3200x), matches TorchSharp CPU - pure managed C# with hand-tuned AVX2/FMA SIMD kernels. Optional CUDA/OpenCL GPU acceleration.
C#★ 11Sep 6, 2026
PyPI · npm69Goodhealth index
jjang-ai/vmlxvMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Python · TypeScript★ 777↓ 4,850/moJul 24, 2026
PyPI63Moderatehealth index
druide67/asiaiMulti-engine LLM benchmark & monitoring CLI for Apple Silicon
Python★ 11↓ 4,662/moJul 18, 2026
PyPI · npm · RubyGems63Moderatehealth index
Ruby★ 1,360↓ 368/moAug 4, 2026
DLTcollab/sse2neonA translator from Intel SSE intrinsics to Arm/Aarch64 NEON implementation
C++ · C · Python★ 1,519Jul 20, 2026
saiyam1814/kiacLocal Kubernetes on Apple's container framework - every node is its own lightweight VM. Metrics, storage, and LoadBalancer included.
Go★ 261Jul 18, 2026
lynicis/applecontainer-goA testcontainers-go-style Go library for spinning up Apple Container CLI Linux containers as test dependencies on macOS.
Go★ 3Jul 15, 2026
PyPI54Moderatehealth index
jjang-ai/jangqJANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon
Python · Swift★ 214Jul 22, 2026
PyPI53Moderatehealth index

ARahim3/mlx-dsparkUp to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, Ornith-1.0, ternary Bonsai-27B.
Python · Swift★ 430↓ 5,564/moAug 19, 2026
PyPI · npm53Moderatehealth index
Python★ 0↓ 6,042/moJul 18, 2026