All tags
Catalogue tag

#evaluation-metrics

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

4 records
Tagged “evaluation-metrics”Ranked by health index
PyPI · npm
94Exceptionalhealth index
confident-ai/deepeval
The LLM Evaluation Framework
Python · TypeScript★ 17.4K↓ 6.3M/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · npm
80Excellenthealth index
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5,801↓ 248.9K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
RubyGems
67Goodhealth index
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 1Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
Go
41Weakhealth index
CircleCI-Research/evalbench
Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.
HTML★ 0Jul 25, 2026
MPL-2.0Jul 25, 2026 · metrics 2.10.0