全部标签
目录标签

#model-comparison

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

3 条记录
标签为“model-comparison”按健康指数排序
78良好健康指数
stan-dev/loo
loo R package for approximate leave-one-out cross-validation (LOO-CV) and Pareto smoothed importance sampling (PSIS)
R★ 1572026年8月30日
自定义许可证2026年8月30日 · 指标 2.10.0
RubyGems
59中等健康指数
justi/ruby_llm-contract
Validate and retry LLM outputs for ruby_llm. Describe the JSON response you expect, fall back to a stronger model when the cheaper one fails the rules, and gate CI on regressions — all as one contract object per step.
Ruby★ 352026年8月22日
MIT2026年8月22日 · 指标 2.10.0
Go
41薄弱健康指数
CircleCI-Research/evalbench
Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.
HTML★ 02026年7月25日
MPL-2.02026年7月25日 · 指标 2.10.0