Alle Tags
Katalog-Tag

#ai-evaluation

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

4 Einträge
Getaggt als „ai-evaluation“Geordnet nach Gesundheitsindex
PyPI · npm
89ExzellentGesundheitsindex
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/Monat19. Aug. 2026
Apache-2.019. Aug. 2026 · Metriken 2.10.0
PyPI
73GutGesundheitsindex
b7n0de/proofbundle
Offline cryptographic receipts for AI evaluation results — Ed25519 + RFC 6962 Merkle + optional SD-JWT. Integrity, not truth
Python★ 2↓ 6.574/Monat23. Juli 2026
MIT23. Juli 2026 · Metriken 2.10.0
PyPI
59MittelGesundheitsindex
mrwersa/agentverity
Your agent test passed. Would it pass again? Checks whether AI agent test results are repeatable and varied enough to trust as a regression baseline. Reads Promptfoo and DeepEval runs you already have.
Python★ 15. Aug. 2026
Apache-2.05. Aug. 2026 · Metriken 2.10.0
PyPI
50MittelGesundheitsindex
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2.908/Monat22. Juli 2026
MIT22. Juli 2026 · Metriken 2.10.0