Alle Tags
Katalog-Tag

#evaluation-metrics

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

4 Einträge
Getaggt als „evaluation-metrics“Geordnet nach Gesundheitsindex
PyPI · npm
94AußergewöhnlichGesundheitsindex
confident-ai/deepeval
The LLM Evaluation Framework
Python · TypeScript★ 17.4K↓ 6.3M/Monat5. Aug. 2026
Apache-2.05. Aug. 2026 · Metriken 2.10.0
PyPI · npm
80ExzellentGesundheitsindex
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5.801↓ 248.9K/Monat28. Aug. 2026
MIT28. Aug. 2026 · Metriken 2.10.0
RubyGems
67GutGesundheitsindex
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 117. Juli 2026
Eigene Lizenz17. Juli 2026 · Metriken 2.10.0
Go
41SchwachGesundheitsindex
CircleCI-Research/evalbench
Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.
HTML★ 025. Juli 2026
MPL-2.025. Juli 2026 · Metriken 2.10.0