Python · TypeScript★ 1,039↓ 148M/月2026年8月27日

Agenta-AI/agentaAgenta is a workspace where you and your team build agents and automations.
TypeScript · Python★ 4,573↓ 17.6K/月2026年8月28日
TypeScript★ 3,487↓ 1,393/月2026年8月13日

comet-ml/opikDebug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/月2026年8月5日

promptfoo/promptfooTest your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/月2026年8月5日

Tencent/WeKnoraOpen-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 19.4K2026年8月5日
Python · Jupyter Notebook★ 3,3622026年7月19日
trpc-group/trpc-agent-goA Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 1,5612026年7月19日
Python★ 27.3K↓ 210.2K/月2026年8月5日

langfuse/langfuse🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5K2026年8月5日
modelscope/evalscopeA streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Python · TypeScript★ 3,172↓ 67.9K/月2026年8月1日
Python · MDX★ 1,055↓ 406.4K/月2026年7月18日

joshuaswarren/remnicOpen-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.
TypeScript★ 176↓ 203.7K/月2026年8月22日
MCPJam/inspectorTesting and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2,069↓ 55.8K/月2026年7月18日
Marker-Inc-Korea/AutoRAGAutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4,963↓ 309/月2026年8月2日

UiPath/coder_evalTest that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/月2026年8月19日
Python★ 35↓ 153.9K/月2026年8月22日
Python · Jupyter Notebook★ 15.2K↓ 1.6M/月2026年8月8日
hidai25/eval-viewRegression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Python★ 124↓ 2,207/月2026年7月26日
Python · Jupyter Notebook★ 332026年7月31日
Python★ 4,309↓ 210.7K/月2026年8月28日

asc-community/AngouriMathOpen-source cross-platform symbolic algebra library for C# and F#. Can be used for both production and research purposes.
C#★ 8272026年8月23日

robocurve/inspect-robotsOpen source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.
Python★ 269↓ 5,131/月2026年9月6日
npm · crates.io · PyPI84优秀健康指数

PSU3D0/formualizerEmbeddable spreadsheet engine - parse, evaluate & mutate Excel workbooks from Rust, Python, or the browser. Arrow-powered, 400+ functions.
Rust★ 177↓ 10.9K/月2026年9月5日
npm · Go · crates.io +281优秀健康指数
TypeScript · C#★ 5↓ 3,643/月2026年9月3日
aikdna/kdnaKDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/月2026年7月22日
o-stepper/graphorinProject Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/月2026年8月1日
huggingface/evaluate🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 2,4652026年7月18日

mjpost/sacrebleuReference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1,259↓ 4.3M/月2026年8月27日
TypeScript★ 5↓ 92.5K/月2026年8月1日