全部标签
目录标签

#agent-evaluation

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

7 条记录
标签为“agent-evaluation”按健康指数排序
PyPI
97卓越健康指数
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python★ 5,7752026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
PyPI
90优秀健康指数
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3,4872026年8月6日
MIT2026年8月6日 · 指标 2.10.0
PyPI · npm
89优秀健康指数
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/月2026年8月19日
Apache-2.02026年8月19日 · 指标 2.10.0
PyPI · npm
88优秀健康指数
hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Python★ 124↓ 2,207/月2026年7月26日
Apache-2.02026年7月26日 · 指标 2.10.0
PyPI
80优秀健康指数
benchflow-ai/benchflow
Research infra for creating RL environments, post-training, and evals.
Python★ 317↓ 6,133/月2026年8月9日
Apache-2.02026年8月9日 · 指标 2.10.0
PyPI
77良好健康指数
alizahidraja/isnad
Grade every agent, scraper and model in a claim's chain — provenance, trust scoring and audit evidence for LLM pipelines
Python★ 37↓ 4,204/月2026年8月29日
Apache-2.02026年8月29日 · 指标 2.10.0
npm · PyPI
45薄弱健康指数
SynthiaResearch/synthia-sdk
Synthia SDKs (npm + PyPI: synthiaresearch) and the synthia CLI — eval your AI agent against simulated users
TypeScript · Python★ 0↓ 3,724/月2026年7月23日
MIT2026年7月23日 · 指标 2.10.0