PyPI97卓越健康指数Giskard-AI/giskard-oss🐢 Open-Source Evaluation & Testing library for LLM AgentsPython★ 5,7752026年8月28日Apache-2.02026年8月28日 · 指标 2.10.0
PyPI90优秀健康指数truera/trulensEvaluation and Tracking for LLM Experiments and AI AgentsPython★ 3,4872026年8月6日MIT2026年8月6日 · 指标 2.10.0
PyPI · npm89优秀健康指数UiPath/coder_evalTest that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.Python · TypeScript★ 116↓ 13.9K/月2026年8月19日Apache-2.02026年8月19日 · 指标 2.10.0
PyPI · npm88优秀健康指数hidai25/eval-viewRegression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.Python★ 124↓ 2,207/月2026年7月26日Apache-2.02026年7月26日 · 指标 2.10.0
PyPI80优秀健康指数benchflow-ai/benchflowResearch infra for creating RL environments, post-training, and evals.Python★ 317↓ 6,133/月2026年8月9日Apache-2.02026年8月9日 · 指标 2.10.0
PyPI77良好健康指数alizahidraja/isnadGrade every agent, scraper and model in a claim's chain — provenance, trust scoring and audit evidence for LLM pipelinesPython★ 37↓ 4,204/月2026年8月29日Apache-2.02026年8月29日 · 指标 2.10.0
npm · PyPI45薄弱健康指数SynthiaResearch/synthia-sdkSynthia SDKs (npm + PyPI: synthiaresearch) and the synthia CLI — eval your AI agent against simulated usersTypeScript · Python★ 0↓ 3,724/月2026年7月23日MIT2026年7月23日 · 指标 2.10.0