全部标签
目录标签

#evaluation-metrics

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

4 条记录
标签为“evaluation-metrics”按健康指数排序
PyPI · npm
94卓越健康指数
confident-ai/deepeval
The LLM Evaluation Framework
Python · TypeScript★ 17.4K↓ 6.3M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
PyPI · npm
80优秀健康指数
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5,801↓ 248.9K/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
RubyGems
67良好健康指数
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 12026年7月17日
自定义许可证2026年7月17日 · 指标 2.10.0
Go
41薄弱健康指数
CircleCI-Research/evalbench
Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.
HTML★ 02026年7月25日
MPL-2.02026年7月25日 · 指标 2.10.0