All tags
Catalogue tag

#agent-evaluation

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

7 records
Tagged “agent-evaluation”Ranked by health index
PyPI
97Exceptionalhealth index
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python★ 5,775Aug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
90Excellenthealth index
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3,487Aug 6, 2026
MITAug 6, 2026 · metrics 2.10.0
PyPI · npm
89Excellenthealth index
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
PyPI · npm
88Excellenthealth index
hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Python★ 124↓ 2,207/moJul 26, 2026
Apache-2.0Jul 26, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
benchflow-ai/benchflow
Research infra for creating RL environments, post-training, and evals.
Python★ 317↓ 6,133/moAug 9, 2026
Apache-2.0Aug 9, 2026 · metrics 2.10.0
PyPI
77Goodhealth index
alizahidraja/isnad
Grade every agent, scraper and model in a claim's chain — provenance, trust scoring and audit evidence for LLM pipelines
Python★ 37↓ 4,204/moAug 29, 2026
Apache-2.0Aug 29, 2026 · metrics 2.10.0
npm · PyPI
45Weakhealth index
SynthiaResearch/synthia-sdk
Synthia SDKs (npm + PyPI: synthiaresearch) and the synthia CLI — eval your AI agent against simulated users
TypeScript · Python★ 0↓ 3,724/moJul 23, 2026
MITJul 23, 2026 · metrics 2.10.0