全部标签
目录标签

#evaluation-framework

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

5 条记录
标签为“evaluation-framework”按健康指数排序
PyPI
83良好健康指数
EleutherAI/lm-evaluation-harness
A framework for few-shot evaluation of language models.
Python★ 13.3K↓ 1.5M/月2026年7月21日
MIT2026年7月21日 · 指标 1.13.0
npm
80良好健康指数
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.4K↓ 1.7M/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
npm · PyPI
78良好健康指数
confident-ai/deepeval
The LLM Evaluation Framework
Python · TypeScript★ 17K↓ 17.7K/月2026年7月20日
Apache-2.02026年7月20日 · 指标 1.13.0
RubyGems
61中等健康指数
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 12026年7月17日
自定义许可证2026年7月17日 · 指标 1.13.0
npm
48存在风险健康指数
empirical-run/empirical
Test and evaluate LLMs and model configurations, across all the scenarios that matter for your application
TypeScript★ 1672026年7月15日
MIT2026年7月15日 · 指标 1.13.0