All tags
Catalogue tag

#ai-evaluation

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

4 records
Tagged “ai-evaluation”Ranked by health index
PyPI · npm
89Excellenthealth index
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
b7n0de/proofbundle
Offline cryptographic receipts for AI evaluation results — Ed25519 + RFC 6962 Merkle + optional SD-JWT. Integrity, not truth
Python★ 2↓ 6,574/moJul 23, 2026
MITJul 23, 2026 · metrics 2.10.0
PyPI
59Moderatehealth index
mrwersa/agentverity
Your agent test passed. Would it pass again? Checks whether AI agent test results are repeatable and varied enough to trust as a regression baseline. Reads Promptfoo and DeepEval runs you already have.
Python★ 1Aug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI
50Moderatehealth index
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2,908/moJul 22, 2026
MITJul 22, 2026 · metrics 2.10.0