All tags
Catalogue tag

#evaluation-metrics

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

3 records
Tagged “evaluation-metrics”Ranked by health index
npm · PyPI
78Goodhealth index
confident-ai/deepeval
The LLM Evaluation Framework
Python · TypeScript★ 17K↓ 17.7K/moJul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 1.13.0
PyPI · npm
71Goodhealth index
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5,716Jul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
RubyGems
61Moderatehealth index
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 1Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0