Alle Tags
Katalog-Tag

#model-comparison

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

3 Einträge
Getaggt als „model-comparison“Geordnet nach Gesundheitsindex
78GutGesundheitsindex
stan-dev/loo
loo R package for approximate leave-one-out cross-validation (LOO-CV) and Pareto smoothed importance sampling (PSIS)
R★ 15730. Aug. 2026
Eigene Lizenz30. Aug. 2026 · Metriken 2.10.0
RubyGems
59MittelGesundheitsindex
justi/ruby_llm-contract
Validate and retry LLM outputs for ruby_llm. Describe the JSON response you expect, fall back to a stronger model when the cheaper one fails the rules, and gate CI on regressions — all as one contract object per step.
Ruby★ 3522. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
Go
41SchwachGesundheitsindex
CircleCI-Research/evalbench
Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.
HTML★ 025. Juli 2026
MPL-2.025. Juli 2026 · Metriken 2.10.0