Python · TypeScript★ 11.2K2026年8月28日
Python★ 4,442↓ 14.6M/月2026年8月27日

mastra-ai/mastraMastra is the modern TypeScript framework for AI-powered applications and agents.
TypeScript★ 26.9K↓ 61.2K/月2026年8月5日
Python★ 27.3K↓ 210.2K/月2026年8月5日
Python★ 4,6942026年8月27日
MCPJam/inspectorTesting and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2,069↓ 55.8K/月2026年7月18日

truera/trulensEvaluation and Tracking for LLM Experiments and AI Agents
Python★ 3,4872026年8月6日

UiPath/coder_evalTest that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/月2026年8月19日
TypeScript · Astro★ 18↓ 25.1K/月2026年7月26日

AgentOps-AI/agentopsPython SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5,801↓ 248.9K/月2026年8月28日
Python★ 317↓ 6,133/月2026年8月9日
o-stepper/graphorinProject Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/月2026年8月1日
Go · Shell★ 22026年7月20日

attenlabs/hotatoFind what broke in your agent calls. Pin it so it never ships again. Local voice-agent call forensics and regression guards.
Python★ 1↓ 1,680/月2026年8月22日
TypeScript · Go★ 15↓ 7,904/月2026年7月30日

superlinear-ai/raglite🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL
Python★ 1,2002026年8月10日
zernie/vigilesLike Lighthouse for your agent harness - verify your CLAUDE.md/AGENTS.md, skills & hooks are real, then test and measure they actually work. Claude Code + Codex.
TypeScript · JavaScript★ 12↓ 4,369/月2026年7月19日
Python★ 2↓ 2,171/月2026年7月23日
PHP★ 4↓ 7,386/月2026年8月4日
aryaminus/controlkeelAgent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Elixir★ 102026年7月17日
shulmansj/teamiControl plane for orchestrating, evaluating, and improving agent work across a company
JavaScript★ 0↓ 2,136/月2026年7月30日
spences10/my-piComposable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
TypeScript★ 88↓ 5,088/月2026年7月18日
farazhassan/gantryA tiny testable, Go-native agent runtime for teams that want control, conformance, and no framework lock-ins.
Go★ 12026年7月18日
inferock/inferock-benchLocal LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.
TypeScript★ 123↓ 6,119/月2026年7月29日
mykim-aus/hey-llm-you-okayHey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2,332/月2026年7月30日
TypeScript★ 0↓ 48.8K/月2026年7月29日
LilMGenius/paperthinLow-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
Shell · JavaScript★ 119↓ 4,136/月2026年7月31日
eigenpal/cliCreate, evaluate, and deploy workflows from your terminal. Agent-ready.
TypeScript★ 2↓ 4,949/月2026年7月18日
Python★ 0↓ 3,588/月2026年7月16日