All tags
Catalogue tag

#model-comparison

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

3 records
Tagged “model-comparison”Ranked by health index
78Goodhealth index
stan-dev/loo
loo R package for approximate leave-one-out cross-validation (LOO-CV) and Pareto smoothed importance sampling (PSIS)
R★ 157Aug 30, 2026
Custom licenseAug 30, 2026 · metrics 2.10.0
RubyGems
59Moderatehealth index
justi/ruby_llm-contract
Validate and retry LLM outputs for ruby_llm. Describe the JSON response you expect, fall back to a stronger model when the cheaper one fails the rules, and gate CI on regressions — all as one contract object per step.
Ruby★ 35Aug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
Go
41Weakhealth index
CircleCI-Research/evalbench
Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.
HTML★ 0Jul 25, 2026
MPL-2.0Jul 25, 2026 · metrics 2.10.0