PyPI88Excellenthealth index
huggingface/transformers🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Python★ 162.7KJul 20, 2026
crates.io · PyPI88Excellenthealth index
vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
Python★ 86.1K↓ 0/moJul 13, 2026
PyPI87Excellenthealth index
huggingface/pytorch-pretrained-BERT🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Python★ 162.7KJul 16, 2026
crates.io · PyPI82Goodhealth index
sgl-project/sglangSGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 30.3K↓ 274M/moJul 14, 2026
PyPI · npm82Goodhealth index
unslothai/unslothUnsloth is a local UI for training and running Gemma 4, Qwen3.6, DeepSeek, Kimi, GLM and other models.
Python · TypeScript★ 68.5K↓ 2.3M/moJul 20, 2026
ollama/ollamaGet up and running with Kimi-K2.6, GLM-5.1, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Go · C★ 176.2KJul 15, 2026
diegosouzapw/omnirouteNever stop coding. Free AI gateway: one endpoint, 231+ providers (50+ free), connect Claude Code, Codex, Cursor, Cline & Copilot to FREE Claude/GPT/Gemini. RTK+Caveman stacked compression saves 15-95% tokens, smart auto-fallback, MCP/A2A, multimodal APIs, Desktop/PWA.
TypeScript★ 17.7K↓ 70K/moJul 16, 2026
jmorganca/ollamaGet up and running with Kimi-K2.6, GLM-5.1, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Go · C★ 176.3KJul 17, 2026
crates.io · npm · PyPI80Goodhealth index
openinterpreter/openinterpreterA coding agent for low-cost models
Rust★ 65.4KJul 15, 2026
lemonade-sdk/lemonadeLemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
C++ · Python · TypeScript★ 4,964Jul 17, 2026
decolua/9routerUnlimited FREE AI coding. Connect Claude Code, Codex, Cursor, Cline, Copilot, Antigravity to FREE Claude/GPT/Gemini via 40+ providers. Auto-fallback, RTK -40% tokens, never hit limits.
JavaScript★ 22.1K↓ 166.8K/moJul 14, 2026
labring/FastGPTFastGPT is a knowledge-based platform built on the LLMs, offers a comprehensive suite of out-of-the-box capabilities such as data processing, RAG retrieval, and visual AI workflow orchestration, letting you easily develop and deploy complex question-answering systems without the need for extensive setup or configuration.
TypeScript★ 29.1K↓ 2,932/moJul 22, 2026
Go · npm73Goodhealth index
ome-projects/omeOpen Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Go★ 481Jul 21, 2026
dyad-sh/dyadLocal, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!
TypeScript★ 20.9KJul 16, 2026
crates.io · npm72Goodhealth index
tailcallhq/forgecodeAI enabled pair programmer for Claude, GPT, O Series, Grok, Deepseek, Gemini and 300+ models
Rust★ 7,460Jul 19, 2026
PyPI · npm71Goodhealth index
lightseekorg/tokenspeedTokenSpeed is a speed-of-light LLM inference engine.
Python★ 1,639↓ 1.7M/moJul 21, 2026
jnMetaCode/superpowers-zh🦸 AI 编程超能力 · 中文增强版 — superpowers(116k+ ⭐)完整汉化 + 6 个中国原创 skills,让 Claude Code / Copilot CLI / Hermes Agent / Cursor / Windsurf / Kiro / Gemini CLI 等 16 款 AI 编程工具真正会干活
JavaScript · Shell★ 7,136↓ 11.3K/moJul 22, 2026
npm67Moderatehealth index
ChatLunaLab/chatluna多平台模型接入,可扩展,多种输出格式,提供大语言模型聊天服务的插件 | A bot plugin for LLM chat with multi-model integration, extensibility, and various output formats
TypeScript★ 426↓ 0/moJul 14, 2026
crates.io · npm67Moderatehealth index
Piebald-AI/splitrailFast, cross-platform, real-time token usage tracker and cost monitor for Gemini CLI / Claude Code / Codex CLI / Qwen Code / Cline / Roo Code / Kilo Code / GitHub Copilot / OpenCode / Pi Agent / Piebald.
Rust★ 209↓ 0/moJul 13, 2026
PyPI67Moderatehealth index
lpalbou/mlx-genGenerative image model runtimes for MLX.
Python★ 15↓ 2,711/moJul 18, 2026
npm · crates.io66Moderatehealth index
snomiao/agent-yesRun AI coding agents (Claude, Codex, Gemini …) unattended — auto-answer prompts, auto-retry on rate limits, and list/tail/steer every agent locally or from agent-yes.com
TypeScript · Rust★ 29↓ 24.7K/moJul 18, 2026
PyPI65Moderatehealth index
athola/claude-night-market23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context optimization, research, and multi-LLM delegation. 186 skills, 128 commands, 54 agents.
Python★ 323Jul 18, 2026
PyPI65Moderatehealth index
jagoff/memoPersistent semantic memory for AI agents — 100% local on Apple Silicon (MLX) or Linux/Ubuntu (CPU). Markdown source of truth, sqlite-vec + BM25 hybrid search, a codegraph-backed knowledge graph, MCP server + CLI. No cloud, no keys.
Python★ 7↓ 3,119/moJul 22, 2026
PyPI · npm64Moderatehealth index
KevRojo/DulusUse Web AI as an agent without an API key. $0 | litellm , local models via Ollama, /lang in 34 languages, Mesa Redonda, voice, OCR, MemPalace, embedded sandbox OS. (The Hermes Killer)
Python★ 391Jul 15, 2026
npm · crates.io64Moderatehealth index
SeemSeam/claude_codex_bridgeVisible multi-agent CLI workspace for mixing Codex, Claude, Gemini, Kimi, Qwen, Cursor, Copilot, Pi, OpenCode, and other AI coding agents
Python · Dart★ 3,301↓ 8,635/moJul 22, 2026
npm · Go64Moderatehealth index
ethanhq/cc-fleet🚢 Run Claude Code's ⚙️ Dynamic Workflows, 👥 Agent Teams & ⚡ Subagents on any third-party model — DeepSeek · GLM · Kimi · Qwen … or your Codex subscription. No Anthropic subscription needed. | 🚢 让 Claude Code 的 ⚙️ Dynamic Workflow、👥 Agent Team、⚡ Subagent 用上任意第三方模型 — DeepSeek · GLM · Kimi · Qwen…… 或你的 Codex 订阅,无需 Claude 订阅
Go★ 190↓ 2,150/moJul 15, 2026
invergent-ai/surogateTraining/Fine-tuning at the speed of light
C++ · Python · Cuda★ 806↓ 0/moJul 21, 2026
npm59Moderatehealth index
therealtimex/node-llama-cppRun AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
TypeScript★ 1↓ 41.2K/moJul 17, 2026
npm · PyPI · Go58Moderatehealth index
startvibecoding/mothxUltra cost-effective terminal AI coding assistant with excellent token cache hit rate. Built in ~10K lines of Go, it uses DeepSeek by default with multiple modes and sandbox.
Go★ 12↓ 10.1K/moJul 15, 2026
mlhher/late-cliStop degrading your model's reasoning. A minimal, zero-config AI coding agent. Enforced ephemeral subagents keep context pure. From tiny local models up to Sol, Fable and Kimi K3.
Go★ 382Jul 20, 2026