All tags
Catalogue tag

#agent-testing

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

5 records
Tagged “agent-testing”Ranked by health index
npm · PyPI
94Exceptionalhealth index
langwatch/scenario
Agentic testing for agentic codebases
Python · TypeScript★ 942↓ 50K/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · npm
89Excellenthealth index
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
npm
80Excellenthealth index
reticlehq/reticle
AI agents can generate code, but they still struggle to understand what they build. Reticle gives them runtime perception of web applications.
TypeScript · JavaScript★ 191↓ 14.3K/moJul 23, 2026
Custom licenseJul 23, 2026 · metrics 2.10.0
PyPI · npm · crates.io
65Goodhealth index
automators-com/flowproof
Agents are starting to run real business processes. flowproof tests them like anything else: record one real run, then replay it on every commit with zero LLM calls, asserting which tools were called, with which arguments, in which order, and which were not.
Rust★ 5↓ 5,937/moAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI
59Moderatehealth index
mrwersa/agentverity
Your agent test passed. Would it pass again? Checks whether AI agent test results are repeatable and varied enough to trust as a regression baseline. Reads Promptfoo and DeepEval runs you already have.
Python★ 1Aug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0