Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Agents are starting to run real business processes. flowproof tests them like anything else: record one real run, then replay it on every commit with zero LLM calls, asserting which tools were called, with which arguments, in which order, and which were not.
Your agent test passed. Would it pass again? Checks whether AI agent test results are repeatable and varied enough to trust as a regression baseline. Reads Promptfoo and DeepEval runs you already have.