npm · crates.io · PyPI94Exceptionalhealth index

headroomlabs-ai/headroomCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Python · Rust★ 64.8K↓ 115.9K/moAug 4, 2026
crates.io · npm93Exceptionalhealth index

yvgude/lean-ctxControl what your AI can see. LeanCTX (Lean Context) is the context intelligence layer for AI agents — one local Rust binary that decides what they read, remembers what they learn, guards what they touch, and proves what they save. 60–90% fewer tokens as the receipt. 76 MCP tools, 30+ agents, local-first.
Rust★ 3,631↓ 3,190/moAug 22, 2026
crates.io · npm90Excellenthealth index

rtk-ai/rtkCLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
Rust★ 74.7KAug 4, 2026
PyPI86Excellenthealth index

jgravelle/jdocmunch-mcpThe leading, most token-efficient MCP server for documentation exploration and retrieval via structured section indexing
Python★ 203↓ 19.1K/moAug 22, 2026
npm86Excellenthealth index
ooples/token-optimizer-mcpIntelligent token optimization for Claude Code - achieving 95%+ token reduction through caching, compression, and smart tool intelligence
TypeScript★ 454↓ 2,292/moJul 29, 2026
npm84Excellenthealth index
TypeScript★ 7,291↓ 9,307/moAug 28, 2026
PyPI · npm83Excellenthealth index
jgravelle/jcodemunch-mcpCut AI token costs 95%+ on code exploration. The leading MCP server for precise, symbol-level GitHub code retrieval via tree-sitter AST. Works with Claude Code, Cursor & any MCP client. 313B+ tokens saved.
Python★ 2,024↓ 88.4K/moJul 16, 2026
Go81Excellenthealth index
edouard-claude/snipCLI proxy that reduces LLM token usage by 60-90%. Declarative YAML filters for Claude Code, Cursor, Copilot, Gemini. rtk alternative in Go.
Go★ 374Jul 19, 2026
PyPI · npm · crates.io80Excellenthealth index
juyterman1000/entrolyAuditable context engineering for AI agents: context optimization, recoverable context compression, receipts, answer verification, and MCP for Claude Code, Codex, OpenClaw.
Python · Rust★ 428↓ 14.6K/moJul 21, 2026
npm · crates.io78Goodhealth index

dPeluChe/trsToken-Reducing Shell — terminal output compression for AI coding agents
Rust★ 11↓ 857/moSep 5, 2026
PyPI · npm77Goodhealth index
SonAIengine/graph-tool-callGraph-based tool retrieval for LLM agents — 248 tools → 82% accuracy, 79% fewer tokens. Zero dependencies. OpenAPI / MCP / LangChain.
Python★ 7↓ 2,636/moAug 1, 2026
crates.io · npm75Goodhealth index
ratel-ai/ratelContext engineering for AI agents. ~80% fewer tokens. Fix tool overload. Skills and memory with in-process BM25 and semantic retrieval. Progressive Disclosure. No vector DB.
Rust · Python · TypeScript★ 233Jul 20, 2026
TypeScript★ 2,129↓ 10.7K/moJul 17, 2026
PyPI · crates.io73Goodhealth index
fkiene/llmtrimLocal proxy that compresses your LLM API requests so you pay less, with no change to the answers. Trims wasted tokens from prompts, history, tool output, and code before they're sent: -31% input / -74% output, measured live. Any provider, no extra model calls. Also an MCP server and embeddable library (Rust, Python, Ruby, Kotlin, Swift, JS/TS).
Rust★ 179↓ 5,878/moJul 26, 2026
maheshmakvana/graphsiftToken Saver for Claude, GPT-5 & Gemini. 80-150x code context reduction, F1 0.85. AST dependency graph, ranked context selection, 19 CLI compressors, MCP server, agent memory. Save LLM tokens — zero telemetry.
Python★ 4↓ 2,797/moAug 4, 2026
sphragis-oss/isthmosLocal context-compression layer for agent tool outputs. Claude Code PostToolUse hook or generic filter, single Go binary.
Go★ 0Jul 31, 2026

wundercorp/openmodelUse any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.
TypeScript · JavaScript · CSS★ 18↓ 3,140/moAug 7, 2026
npm · crates.io · PyPI62Moderatehealth index
Tura-AI/turaAcross 348 long-horizon benchmark sessions, Tura used up to 83.1% fewer turns on the rewrite benchmark and improved the DeepSWE pass rate by up to 16.7 percentage points compared with Codex CLI.
Rust · TypeScript · JavaScript★ 56↓ 3,378/moJul 16, 2026
PyPI60Moderatehealth index
Chuzom/ChuzomLightweight signal-driven LLM router for Claude Code, Cursor, Codex, Gemini CLI, and Codex CLI
Python★ 19↓ 2,660/moAug 1, 2026
npm59Moderatehealth index
TaewoooPark/Agent-BlackboxLocal-first flight recorder for coding agents : replay every run as a live session map, score the context bill, and write the fix back into AGENTS.md — no API key, one npx command.
TypeScript★ 57↓ 3,988/moJul 21, 2026
Packagist · npm59Moderatehealth index
PHP★ 144↓ 11K/moJul 23, 2026
npm59Moderatehealth index
psjostrom/frontloadLocal-first context gateway that helps AI coding agents read less, spend less, and stay grounded in your repo.
TypeScript★ 0↓ 2,191/moJul 24, 2026
npm59Moderatehealth index
TypeScript★ 10↓ 2,951/moJul 23, 2026
npm57Moderatehealth index
TypeScript★ 9↓ 7,997/moAug 29, 2026
npm · Go · PyPI57Moderatehealth index

firstops-dev/whittleCarves your agent's tool outputs down to what matters. Never cuts what doesn't come back.
Go · Python★ 61↓ 37/moSep 5, 2026
npm56Moderatehealth index
HTML · Shell★ 1↓ 6,010/moJul 18, 2026
npm56Moderatehealth index
kitepon-rgb/aiterm-mcpOne persistent MCP terminal your AI drives — and launches other coding agents (Codex/Grok/Composer) into. SSH, containers, and REPLs nest as text you send in. tmux-backed, token-reduced reads, headless over MCP.
JavaScript · TypeScript · Python★ 1↓ 2,175/moJul 15, 2026

blackwell-systems/gcf-goGCF Go implementation. 100% LLM comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ round-trips verified. Zero dependencies.
Go★ 4Aug 22, 2026
PyPI54Moderatehealth index
wesleysimplicio/simplicio-loop🔁 Finishes your entire backlog while you sleep. The AI orchestrator that DOES the work end-to-end on ANY LLM — discover → implement → verify → merge → 24/7 — behind safety gates, at up to 90% fewer tokens. 48 extension points. Not a chatbot. A worker.
Python★ 10↓ 3,430/moJul 31, 2026
npm51Moderatehealth index
bassprofressor-lab/openwolf-enhancedEnhanced fork of OpenWolf — a token-conscious second brain for Claude Code, with bounded storage, self-maintenance (openwolf doctor), .wolfignore scoping, and tunable retention. AGPL-3.0.
TypeScript · JavaScript★ 6↓ 5,230/moJul 30, 2026