Alle Tags
Katalog-Tag

#agent-evaluation

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

7 Einträge
Getaggt als „agent-evaluation“Geordnet nach Gesundheitsindex
PyPI
97AußergewöhnlichGesundheitsindex
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python★ 5.77528. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
PyPI
90ExzellentGesundheitsindex
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3.4876. Aug. 2026
MIT6. Aug. 2026 · Metriken 2.10.0
PyPI · npm
89ExzellentGesundheitsindex
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/Monat19. Aug. 2026
Apache-2.019. Aug. 2026 · Metriken 2.10.0
PyPI · npm
88ExzellentGesundheitsindex
hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Python★ 124↓ 2.207/Monat26. Juli 2026
Apache-2.026. Juli 2026 · Metriken 2.10.0
PyPI
80ExzellentGesundheitsindex
benchflow-ai/benchflow
Research infra for creating RL environments, post-training, and evals.
Python★ 317↓ 6.133/Monat9. Aug. 2026
Apache-2.09. Aug. 2026 · Metriken 2.10.0
PyPI
77GutGesundheitsindex
alizahidraja/isnad
Grade every agent, scraper and model in a claim's chain — provenance, trust scoring and audit evidence for LLM pipelines
Python★ 37↓ 4.204/Monat29. Aug. 2026
Apache-2.029. Aug. 2026 · Metriken 2.10.0
npm · PyPI
45SchwachGesundheitsindex
SynthiaResearch/synthia-sdk
Synthia SDKs (npm + PyPI: synthiaresearch) and the synthia CLI — eval your AI agent against simulated users
TypeScript · Python★ 0↓ 3.724/Monat23. Juli 2026
MIT23. Juli 2026 · Metriken 2.10.0