全部标签
目录标签

#ai-evaluation

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

4 条记录
标签为“ai-evaluation”按健康指数排序
PyPI · npm
89优秀健康指数
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/月2026年8月19日
Apache-2.02026年8月19日 · 指标 2.10.0
PyPI
73良好健康指数
b7n0de/proofbundle
Offline cryptographic receipts for AI evaluation results — Ed25519 + RFC 6962 Merkle + optional SD-JWT. Integrity, not truth
Python★ 2↓ 6,574/月2026年7月23日
MIT2026年7月23日 · 指标 2.10.0
PyPI
59中等健康指数
mrwersa/agentverity
Your agent test passed. Would it pass again? Checks whether AI agent test results are repeatable and varied enough to trust as a regression baseline. Reads Promptfoo and DeepEval runs you already have.
Python★ 12026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI
50中等健康指数
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2,908/月2026年7月22日
MIT2026年7月22日 · 指标 2.10.0