Öffentliches Register
Software-GesundheitsberichtSchema 0.27.0 · Metriken 1.13.0 · 2026-07-25 08:58 UTC

CircleCI-Research / evalbench

Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.

HTMLMPL-2.0★ 0 Sterne⑂ 0 Forksseit März 2026ForkAuf GitHub ansehen ↗

CircleCI-Research/evalbench erreicht einen Gesundheitsindex von 43 von 100 und liegt damit im Bereich Gefährdet. Am stärksten schneidet es bei AI Readiness (65/100) ab, am schwächsten bei Community & Adoption (12/100). Zuletzt vor 127 Tagen aktualisiert. Ein einzelner Mitwirkender trägt den Großteil der jüngsten Arbeit.

43
gesamt / 100
Gefährdet

Software-Gesundheitsindex

Metriken werden auf einer Skala von 1–100 in gewichtete Kategorien gruppiert. Der Gesamtwert beginnt als ihr Mittel; sobald öffentliche Evidenz die Richtlinie für Hochrisikojurisdiktionen auslöst, wird die Bewertung angepasst und erhält die Obergrenze 49 (Gefährdet). AI Readiness liegt außerhalb.

43
Exzellent85-100Vorbildlich; erfüllt im Wesentlichen alle geprüften Kriterien
Gut70-84Gesund; geringfügige Lücken
Mittel50-69Akzeptabel mit deutlichen Lücken; Überprüfung empfohlen
Gefährdet30-49Erhebliche Schwächen; eine Übernahme erfordert Vorsicht
Kritisch1-29Schwerwiegende Probleme (aufgegeben, nur ein Maintainer, keine Hygiene)
VitalitätCommunity &VerbreitungNachhaltigkeit &GovernanceEngineering-QualitätSicherheitAI Readiness

Bewertungsprofil

Jede Achse ist eine Kategorie. Die Form zählt mehr als der Durchschnitt — ein gesundes Projekt füllt die gesamte Fläche, während ein Profil aus Spitzen und Kratern bedeutet, dass Stärke in einer Dimension Risiken in einer anderen verdeckt.

Eigentümerschaft

CircleCI ResearchOrganisation
3 Follower4 öffentliche Reposseit Okt. 2022

Dieses Repository wird von einer Organisation getragen — geteilte, rechenschaftspflichtige Trägerschaft, die jeden einzelnen Maintainer überdauern kann.

Paket-Ökosysteme

RegistryPaketVersionDownloads / MonatVersionenZuletzt veröffentlicht
Gogithub.com/CircleCI-Research/evalbenchv0.17.0-33vor 128 Tagen

Metriken nach Kategorie

Vitalität

Lebt das Projekt — wird Code geschrieben und werden Releases ausgeliefert?

56Mittel · 22 % des Gesamtindex
Wie die Bewertung erfolgt
9.9/36Push-Aktualität — letzter Push vor 127 Tagen
14.5/36Commit-Rhythmus — 21/52 Wochen mit Commits
17.4/18Commit-Volumen — 86 Commits im letzten Jahr
0/10OpenSSF Scorecard: Maintained — 0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0
Verwendete Eingangsdaten
commits_last_year86
human_commit_share1
days_since_last_push127
active_weeks_last_year21
Wie die Bewertung erfolgt
16.2/27Liefert Releases aus — 33 Versions-Tags (keine GitHub-Releases)
27/36Release-Aktualität — letztes Release vor 128 Tagen
27/27Release-Rhythmus — ein Release etwa alle 13,2 Tage
0/10OpenSSF Scorecard: Signed-Releases — keine Daten
Verwendete Eingangsdaten
releases_count33
latest_release_tagv0.17.0
releases_from_tagsja
days_since_latest_release128
mean_days_between_releases13,2
Von der Bewertung ausgeschlossen (keine Daten oder nicht anwendbar): OpenSSF Scorecard: Signed-Releases. Die verbleibenden Gewichte wurden renormalisiert.

Community & Verbreitung

Hat das Projekt Nutzer, Downloads, Aufmerksamkeit und ein einladendes Umfeld für Beitragende?

12Kritisch · 18 % des Gesamtindex
Wie die Bewertung erfolgt
0/60Stars — 0 Stars
0/25Forks — 0 Forks
0/15Watcher — 0 Watcher
Verwendete Eingangsdaten
forks0
stars0
watchers0
growth_stateunverified
growth_factor_pct100
growth_unverified_reasonno_history
Wie die Bewertung erfolgt
0/22.5README
22.5/22.5Lizenz — anerkannte Lizenz (MPL-2.0)
0/18CONTRIBUTING-Leitfaden
0/13.5Verhaltenskodex
0/7.2Issue-Vorlage
0/6.3PR-Vorlage
Verwendete Eingangsdaten
has_readmenein
has_licensenein
has_contributingnein
has_issue_templatenein
has_code_of_conductnein
has_pull_request_templatenein

Nachhaltigkeit & Governance

Überdauert das Projekt die Menschen, die es tragen — Bus-Faktor, Reaktionsfähigkeit, Trägerschaft und Paketpflege?

57Mittel · 24 % des Gesamtindex
Wie die Bewertung erfolgt
9/54Bus-Faktor — 1 Beitragende decken die Hälfte aller Commits ab
7.2/22.5Commit-Verteilung — wichtigste beitragende Person verfasste 68 % der Commits
2.7/13.5Breite der Beitragenden — 2 Beitragende
6/10OpenSSF Scorecard: Contributors — project has 2 contributing companies or organizations -- score normalized to 6
Verwendete Eingangsdaten
bus_factor1
contributors_sampled2
top_contributor_share0,681
Wie die Bewertung erfolgt
0/46.8Issue-Lösungsquote — keine Issues oder keine Daten
38.2/38.3PR-Annahme — 1/1 entschiedene PRs gemergt
0/15OpenSSF Scorecard: Code-Review — Found 0/26 approved changesets -- score normalized to 0
Verwendete Eingangsdaten
merged_prs1
open_issues0
closed_issues0
issue_closed_ratio
closed_unmerged_prs0
Von der Bewertung ausgeschlossen (keine Daten oder nicht anwendbar): Issue-Lösungsquote. Die verbleibenden Gewichte wurden renormalisiert.
Wie die Bewertung erfolgt
30/30Organisatorische Trägerschaft — im Besitz einer Organisation
0/20Verifizierte Domain
4.3/25Reichweite des Inhabers — 3 Follower von CircleCI-Research
12.7/25Kontohistorie — 4 öffentliche Repos, Kontoalter ca. 3 Jahre
Verwendete Eingangsdaten
followers3
owner_typeOrganization
is_verified
owner_loginCircleCI-Research
public_repos4
account_age_days1.387

Paketpflege

100Exzellent
Wie die Bewertung erfolgt
25/25Veröffentlicht & auflösbar — 1 Paket(e) auf go
35/35Veröffentlichungsaktualität — letzte Veröffentlichung vor 128 Tagen
20/20Versionshistorie — 33 veröffentlichte Versionen
20/20Nicht veraltet — aktiv, nicht veraltet oder zurückgezogen
Verwendete Eingangsdaten
packagesgithub.com/CircleCI-Research/evalbench
ecosystemsgo
any_deprecatednein
min_days_since_publish128

Engineering-Qualität

Sind grundlegende Engineering- und Dokumentationspraktiken vorhanden?

51Mittel · 20 % des Gesamtindex
Wie die Bewertung erfolgt
24/24CI-Workflows — 1 Workflow(s)
24/24Tests vorhanden
0/16Linter-Konfiguration
0/9.6Pre-Commit-Hooks
0/6.4.editorconfig
20/20OpenSSF Scorecard: CI-Tests — 1 out of 1 merged PRs checked by a CI test -- score normalized to 10
Verwendete Eingangsdaten
has_cija
has_testsja
has_editorconfignein
has_linter_confignein
has_precommit_confignein

Dokumentation

25Kritisch
Wie die Bewertung erfolgt
0/30README
0/25Dokumentationsverzeichnis
15/15Dokumentations-/Homepage-Site — https://loop.circleci.com
0/10Repository-Beschreibung
10/10Topics — 10 Topics
0/10Wiki
Verwendete Eingangsdaten
topicsai, ai-agents, benchmark-framework, ci-cd, circleci, devops, evaluation-framework, evaluation-metrics, llm, model-comparison
has_wikinein
homepagehttps://loop.circleci.com
has_readmenein
has_docs_dirnein
has_descriptionnein

Sicherheit

Sind die sichtbaren Sicherheits- und Lieferkettenpraktiken belastbar, ohne ungeklärte Exposition gegenüber Hochrisikojurisdiktionen?

30Gefährdet · 16 % des Gesamtindex

Sicherheitslage

30Gefährdet
Wie die Bewertung erfolgt
7.5/7.5Binary-Artifacts — no binaries found in the repo
3.8/7.5Branch-Protection — branch protection is not maximal on development and all release branches
2.5/2.5CI-Tests — 1 out of 1 merged PRs checked by a CI test -- score normalized to 10
0/2.5CII-Best-Practices — no effort to earn an OpenSSF best practices badge detected
0/7.5Code-Review — Found 0/26 approved changesets -- score normalized to 0
1.5/2.5Contributors — project has 2 contributing companies or organizations -- score normalized to 6
10/10Dangerous-Workflow — no dangerous workflow patterns detected
0/7.5Dependency-Update-Tool — no update tool detected
0/5Fuzzing — project is not fuzzed
2.5/2.5Lizenz — license file detected
0/7.5Maintained — 0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0
0/5Packaging — keine Daten
0/5Pinned-Dependencies — dependency not pinned by hash detected -- score normalized to 0
0/5SAST — SAST tool is not run on all commits -- score normalized to 0
0/5Security-Policy — security policy file not detected
0/7.5Signed-Releases — keine Daten
0/7.5Token-Permissions — detected GitHub workflow tokens with excessive permissions
0/7.5Vulnerabilities — 44 existing vulnerabilities detected
Verwendete Eingangsdaten
sourceopenssf_scorecard
checks_evaluated16
scorecard_versionv5.5.0
checks_inconclusive2
scorecard_aggregate3
Von der Bewertung ausgeschlossen (keine Daten oder nicht anwendbar): packaging, signed_releases. Die verbleibenden Gewichte wurden renormalisiert.

AI Readiness

Wie gut ist das Repository dafür ausgestattet, mit KI-Coding-Agenten entwickelt und gepflegt zu werden? Ein unabhängiges, experimentelles Badge — Gewicht 0,0, es wird eigenständig ausgewiesen und verändert den Gesamt-Gesundheitswert nicht.

65Mittel · 0 % des Gesamtindex
Wie die Bewertung erfolgt
45/45Agentenanweisungen — AGENTS.md
0/15Maschinenlesbare Doku (llms.txt)
40/40Lesbare Commit-Historie — 96 von 100 menschlichen Commits benennen ihre Absicht (strukturierter Betreff oder erläuternder Text)
Verwendete Eingangsdaten
has_llms_txtnein
legible_history_share0,96
agent_instruction_filesAGENTS.md
agent_instruction_max_bytes3.085
Wie die Bewertung erfolgt
12.6/18Bootstrap mit einem Befehl — go.mod (Toolchain-Konvention, kein Task-Runner)
22/22Automatisierte Tests
0/11Lint-/Format-Konfiguration
0/11Statische Typprüfung
10/10Reproduzierbare Umgebung — lockfile
10/10Belegte Agentenpraxis — 11 der letzten 100 Commits von Agenten verfasst oder ihnen zugeschrieben
0/8Automatisierte Wartung — keine automatisierten Abhängigkeits-Updates beobachtet
0/10OpenSSF Scorecard: Pinned-Dependencies — dependency not pinned by hash detected -- score normalized to 0
Verwendete Eingangsdaten
has_nixnein
has_testsja
lockfilesgo.sum
has_dockerfilenein
typed_languagenein
bootstrap_files
has_devcontainernein
has_linter_confignein
typecheck_configs
agent_commit_share0,11
toolchain_manifestsgo.mod
dependency_bot_commit_share0
Wie die Bewertung erfolgt
0/45Typprüfbarer Code — HTML ohne Typprüfungs-Konfiguration
54.4/55Handhabbare Dateigrößen — 5/442 Quelldateien über 60 KB
Verwendete Eingangsdaten
primary_languageHTML
largest_source_bytes73.199
source_files_sampled442
oversized_source_files5

Eckdaten

0GitHub-Sterne
2Mitwirkende
86Commits, letzte 12 Monate
127Tage seit letztem Push
33Releases
1Bus-Faktor
0offene Issues
GoPaket-Ökosysteme

Warnungen zur Datenerhebung

  • Community profile unavailable
  • GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository

Weitere Details

OpenSSF Scorecard 3.0 / 10
3.0Gesamtwert

Unabhängige, werkzeugneutrale Sicherheitsbewertung durch das quelloffene OpenSSF Scorecard. Jede Prüfung honoriert eine Sicherheits-Praxis, nicht das Werkzeug eines bestimmten Anbieters. Prüfungen, die Scorecard nicht ermitteln konnte, sind mit k. A. markiert und vom Sicherheitswert ausgeschlossen (nie als null gezählt).Scorecard v5.5.0 · 2026-07-25 08:58 UTC

10Binary-Artifactsno binaries found in the repo
5Branch-Protectionbranch protection is not maximal on development and all release branches
10CI-Tests1 out of 1 merged PRs checked by a CI test -- score normalized to 10
0CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
0Code-ReviewFound 0/26 approved changesets -- score normalized to 0
6Contributorsproject has 2 contributing companies or organizations -- score normalized to 6
10Dangerous-Workflowno dangerous workflow patterns detected
0Dependency-Update-Toolno update tool detected
0Fuzzingproject is not fuzzed
10Licenselicense file detected
0Maintained0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0
k. A.Packagingpackaging workflow not detected
0Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
0SASTSAST tool is not run on all commits -- score normalized to 0
0Security-Policysecurity policy file not detected
k. A.Signed-Releasesno releases found
0Token-Permissionsdetected GitHub workflow tokens with excessive permissions
0Vulnerabilities44 existing vulnerabilities detected
Direkte Abhängigkeiten 24
RegistryPaketVersionsvorgabeManifest
Gogithub.com/charmbracelet/x/termv0.2.1go.mod
Gogithub.com/containerd/errdefsv1.0.0go.mod
Gogithub.com/docker/dockerv28.5.2+incompatiblego.mod
Gogithub.com/google/uuidv1.6.0go.mod
Gogithub.com/invopop/jsonschemav0.13.0go.mod
Gogithub.com/kaptinlin/jsonrepairv0.2.8go.mod
Gogithub.com/oklog/ulid/v2v2.1.1go.mod
Gogithub.com/openai/openai-go/v3v3.26.0go.mod
Gogithub.com/santhosh-tekuri/jsonschema/v6v6.0.2go.mod
Gogithub.com/sethvargo/go-retryv0.3.0go.mod
Gogithub.com/stretchr/testifyv1.11.1go.mod
Gogolang.org/x/timev0.14.0go.mod
Gogoogle.golang.org/genaiv1.48.0go.mod
Gogopkg.in/validator.v2v2.0.1go.mod
Gogopkg.in/yaml.v3v3.0.1go.mod
Gogithub.com/anthropics/anthropic-sdk-gov1.26.0go.mod
Gogithub.com/charmbracelet/bubbles/v2v2.0.0-beta.1go.mod
Gogithub.com/charmbracelet/bubbletea/v2v2.0.0-beta.1go.mod
Gogithub.com/charmbracelet/lipgloss/v2v2.0.0-beta.1go.mod
Gogithub.com/cohesion-org/deepseek-gov1.3.3go.mod
Gogithub.com/go-playground/validator/v10v10.30.1go.mod
Gogithub.com/rs/zerologv1.34.0go.mod
Gogithub.com/sergi/go-diffv1.4.0go.mod
Gogolang.org/x/expv0.0.0-20260218203240-3dfff04db8fago.mod
Alle Abhängigkeiten nicht erhoben

Der aufgelöste Abhängigkeitssatz konnte für diesen Bericht nicht erhoben werden: GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository

JSON-Rohbericht maschinenlesbar
{
  "data": {
    "repo": {
      "topics": [
        "ai",
        "ai-agents",
        "benchmark-framework",
        "ci-cd",
        "circleci",
        "devops",
        "evaluation-framework",
        "evaluation-metrics",
        "llm",
        "model-comparison"
      ],
      "is_fork": true,
      "size_kb": 7617,
      "has_wiki": false,
      "homepage": "https://loop.circleci.com",
      "languages": {
        "Go": 876390,
        "HTML": 20005421,
        "Shell": 1009,
        "Python": 5074,
        "Go Template": 78025
      },
      "pushed_at": "2026-03-19T15:10:09Z",
      "created_at": "2026-03-19T14:20:11Z",
      "owner_type": "Organization",
      "updated_at": "2026-03-19T16:25:31Z",
      "description": "Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.",
      "is_archived": false,
      "is_disabled": false,
      "license_spdx": "MPL-2.0",
      "default_branch": "main",
      "license_spdx_raw": "MPL-2.0",
      "primary_language": "HTML",
      "significant_languages": [
        "HTML"
      ]
    },
    "owner": {
      "blog": "https://circleci.com",
      "name": "CircleCI Research",
      "type": "Organization",
      "login": "CircleCI-Research",
      "company": null,
      "location": "United States of America",
      "followers": 3,
      "avatar_url": "https://avatars.githubusercontent.com/u/115158100?v=4",
      "created_at": "2022-10-06T12:04:53Z",
      "is_verified": null,
      "public_repos": 4,
      "account_age_days": 1387
    },
    "license": {
      "state": "standard",
      "spdx_id": "MPL-2.0",
      "raw_spdx": "MPL-2.0",
      "file_present": true,
      "scorecard_found": true,
      "profile_has_license": false
    },
    "activity": {
      "releases": [
        {
          "tag": "v0.17.0",
          "kind": "minor",
          "published_at": "2026-03-18T22:43:28Z"
        },
        {
          "tag": "v0.16.0",
          "kind": "minor",
          "published_at": "2026-03-12T19:58:14Z"
        },
        {
          "tag": "v0.15.0",
          "kind": "minor",
          "published_at": "2026-02-27T22:17:14Z"
        },
        {
          "tag": "v0.14.1",
          "kind": "patch",
          "published_at": "2026-02-07T15:46:27Z"
        },
        {
          "tag": "v0.14.0",
          "kind": "minor",
          "published_at": "2026-02-01T03:21:58Z"
        },
        {
          "tag": "v0.13.4",
          "kind": "patch",
          "published_at": "2026-01-17T19:22:30Z"
        },
        {
          "tag": "v0.13.3",
          "kind": "patch",
          "published_at": "2025-12-23T19:35:43Z"
        },
        {
          "tag": "v0.13.2",
          "kind": "patch",
          "published_at": "2025-12-12T18:17:51Z"
        },
        {
          "tag": "v0.13.1",
          "kind": "patch",
          "published_at": "2025-12-08T00:26:37Z"
        },
        {
          "tag": "v0.13.0",
          "kind": "minor",
          "published_at": "2025-11-20T05:56:34Z"
        },
        {
          "tag": "v0.12.2",
          "kind": "patch",
          "published_at": "2025-11-12T21:39:53Z"
        },
        {
          "tag": "v0.12.1",
          "kind": "patch",
          "published_at": "2025-10-11T01:09:29Z"
        },
        {
          "tag": "v0.12.0",
          "kind": "minor",
          "published_at": "2025-10-09T18:05:48Z"
        },
        {
          "tag": "v0.11.1",
          "kind": "patch",
          "published_at": "2025-10-03T00:05:30Z"
        },
        {
          "tag": "v0.11.0",
          "kind": "minor",
          "published_at": "2025-10-02T15:35:36Z"
        },
        {
          "tag": "v0.10.1",
          "kind": "patch",
          "published_at": "2025-09-23T21:16:27Z"
        },
        {
          "tag": "v0.10.0",
          "kind": "minor",
          "published_at": "2025-09-19T16:11:27Z"
        },
        {
          "tag": "v0.9.0",
          "kind": "minor",
          "published_at": "2025-09-16T18:47:58Z"
        },
        {
          "tag": "v0.8.0",
          "kind": "minor",
          "published_at": "2025-09-03T14:00:21Z"
        },
        {
          "tag": "v0.7.2",
          "kind": "patch",
          "published_at": "2025-08-28T14:57:25Z"
        },
        {
          "tag": "v0.7.1",
          "kind": "patch",
          "published_at": "2025-08-26T16:25:26Z"
        },
        {
          "tag": "v0.7.0",
          "kind": "minor",
          "published_at": "2025-08-21T20:16:13Z"
        },
        {
          "tag": "v0.6.1",
          "kind": "patch",
          "published_at": "2025-08-16T17:01:31Z"
        },
        {
          "tag": "v0.6.0",
          "kind": "minor",
          "published_at": "2025-08-11T18:11:56Z"
        },
        {
          "tag": "v0.5.0",
          "kind": "minor",
          "published_at": "2025-07-30T03:28:50Z"
        },
        {
          "tag": "v0.4.2",
          "kind": "patch",
          "published_at": "2025-07-08T19:46:19Z"
        },
        {
          "tag": "v0.4.1",
          "kind": "patch",
          "published_at": "2025-07-07T17:30:18Z"
        },
        {
          "tag": "v0.3.2",
          "kind": "patch",
          "published_at": "2025-06-27T16:30:55Z"
        },
        {
          "tag": "v0.3.1",
          "kind": "patch",
          "published_at": "2025-06-18T04:38:04Z"
        },
        {
          "tag": "v0.3.0",
          "kind": "minor",
          "published_at": "2025-05-31T13:41:35Z"
        },
        {
          "tag": "v0.2.1",
          "kind": "patch",
          "published_at": "2025-05-24T02:21:21Z"
        },
        {
          "tag": "v0.2.0",
          "kind": "minor",
          "published_at": "2025-05-19T18:49:36Z"
        },
        {
          "tag": "v0.1.0",
          "kind": "minor",
          "published_at": "2025-04-26T18:50:00Z"
        }
      ],
      "recent_commits": [
        {
          "oid": "64b9f020db13d4629897badb64f94f85a8003544",
          "body": "fix: restore GHA workflow, fix failing tests, add AGENTS.md",
          "is_bot": false,
          "headline": "Merge pull request #1 from CircleCI-Research/fix-gha",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T15:10:03Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "d0db7a898f6f628cb467accc7850547be5c042a4",
          "body": "…sitive\n\nReverts the sequential run change — runs within a provider intentionally\nexecute in parallel for throughput. Updates the doc comment to match.\n\nFixes TestRunnerRun by sorting results by (Run, Task, Got, Want) before\ncomparing, since parallel execution produces non-deterministic ordering.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "fix: restore parallel run execution and make runner tests order-insen…",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T15:06:08Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "9a263d0487b9867f614754612e963e7044385cbe",
          "body": "…golden files\n\nRuns within a single provider were incorrectly launched in parallel\ngoroutines, causing non-deterministic result ordering that violated the\ndocumented contract (\"individual runs on a single provider are executed\nsequentially\") and broke TestRunnerRun assertions.\n\nAlso updates log formatter golden files to include the Score column\nadded to the log output format.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "fix: make runs sequential within a provider and update log formatter …",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T15:01:09Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "625baa335648b4da4d8c41ccae05e3fe2ce5a214",
          "body": "Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "docs: add AGENTS.md with repo guidance and remote enforcement",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T14:54:31Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "ef7744e213fd6ed148d7378b3cf47cd6032ef80e",
          "body": "The workflow was removed during the CircleCI migration but the README\nbadge still references it, causing a broken badge on the repo.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "fix: restore GitHub Actions workflow for build badge",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T14:44:15Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "001659b4955de4026e29d49a8077a65059c19d54",
          "body": null,
          "is_bot": false,
          "headline": "refactor: tighten up README for ease of use",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T14:30:07Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "1dd80d60d0afea570bcdf8a4f1c7ad486b06bc68",
          "body": "- Rename binary from mindtrial to evalbench (cmd/mindtrial -> cmd/evalbench)\n- Migrate Go module path from github.com/petmal/mindtrial to\n  github.com/CircleCI-Research/evalbench\n- Replace MindTrial product name with EvalBench throughout codebase\n- Update README badges, links, and install commands for CircleCI-Research\n- Add attribution: \"Originally created by Petr Malik as MindTrial\"\n- Preserve MPL 2.0 license and all original copyright headers\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "chore: rebrand MindTrial to EvalBench and migrate to CircleCI-Research",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T14:08:38Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "ea8a6456e0884da66f75d40e99839f35e371a5e3",
          "body": "feat: Live voice race announcer for model comparisons 🎙️🤠",
          "is_bot": false,
          "headline": "Merge pull request #2 from ryan-circleci/add-live-voice-announcer",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T15:37:16Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "cd2b9aee8d945d9b4c76951cfa3b2478d605af0c",
          "body": "Made-with: Cursor",
          "is_bot": false,
          "headline": "fix: remove 30s announcer interval option (cron minimum is 1m)",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T15:35:52Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "4a37c3fbd8815c57aff691e4fa42577a73f76a30",
          "body": "…uto-loop note\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "docs: update announcer usage section with with-announcer syntax and a…",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T09:23:01Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "3fe62cea75d6750962b83dc1877fe8c4c1ce91c9",
          "body": "- Ask user for announcer check-in frequency when with-announcer is set but no interval specified\n- Support numeric arg immediately after `with-announcer` to set interval directly\n- Fix arg parsing so config file detection uses .yaml suffix instead of positional order\n- Wire ANNOUNCER_INTERVAL through to cron schedule and loop interval file\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: add interactive announcer interval prompt and flexible arg parsing",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T09:18:08Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "8738c92a52a7be5361cb7599c96ff3d663839142",
          "body": null,
          "is_bot": false,
          "headline": "Create open-dashboard.sh",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T09:05:44Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b4d9ed5b7afd9549e45b55ecee3ef3f8a887b271",
          "body": "…r flag\n\nBoth /run-model-comparison and /simulate-model-comparison now default to\nsilent mode and ask the user upfront whether they want live voice commentary.\nPass with-announcer as an argument to skip the question. Silent mode suggests\n/loop Nm /announce-model-comparison as an easy on-ramp.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: make announcer opt-in with interactive prompt and with-announce…",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T08:59:55Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "957b0b89a47d1251c3361453feb6bd5c1b665e86",
          "body": null,
          "is_bot": false,
          "headline": "Update README.md",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T07:11:02Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9e80a0530bcfc29722884363dc9d73e6b58b8ee5",
          "body": "Background agents lose Bash permissions mid-run; /loop runs in the main\nsession where permissions are already granted, so it's the only reliable\npath. Removes the Task-based auto-launch and the Step 0b mode picker.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "refactor: drop with-announcer Task mode, always use /loop for announcer",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T06:53:06Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "530854d5ab37fd4be4c05f32444a16f7e5bc4fdf",
          "body": null,
          "is_bot": false,
          "headline": "feat: add loop interval-specific color & realism to announcer commentary",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T06:25:57Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "740b9acc71da17e605ee356c8127b9de822319a0",
          "body": "- Use YYYY-MM-DD/HH-MM-SS ISO format everywhere (simulation + eval runs)\n- Unify eval results under results/eval/ prefix across slash commands and config YAMLs\n- Pass -output-dir and -output-basename to mindtrial so CSV/HTML land alongside eval.log\n- Rename /tmp/.race_* temp files to /tmp/.eval_* for consistency\n- Add interactive config picker to run-model-comparison\n- Ignore results/ and .claude/scheduled_tasks.lock in .gitignore\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: standardize results directory naming and ignore runtime artifacts",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T04:29:16Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "f067409a54ed9fe10d6aacecfc0bf4138f8309e6",
          "body": "Accidentally changed python3 to bash when refactoring the log path.\nThe .sh file is actually a Python script.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "fix: restore python3 invocation for simulate script",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T03:42:38Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "f3adbbc16f07a7f917d465128f6ead99e93498ef",
          "body": "Store eval.log at $RESULTS_DIR/eval.log instead of logs/eval.log so\neach run has a self-contained results folder. Path is written to\n/tmp/.race_log_file at init time and read by simulate, run, announce,\nand stop commands via a logs/eval.log fallback.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: co-locate eval.log inside results directory",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T03:40:07Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "e0bb5c1a78ba31f73ad85c9c90f956419b7dd02f",
          "body": "- Add model/provider count to opening announcement in run-model-comparison\n- Change results folder date format from YYYY-MM-DD to MM-DD-YYYY\n- Change results time subfolder format to 12h am/pm (e.g. 11-23pm)\n- Update stop-model-comparison TTS voice from af_heart to am_michael\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: improve voice announcer UX and results folder naming",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T03:33:11Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "75c0ec4bbc927f34462b9c020a537ff8aadc01cc",
          "body": null,
          "is_bot": false,
          "headline": "update voice",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T01:57:41Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "f33a80f86cd1650e4e7a0ec33ec5beb573d5ea50",
          "body": "Rename commands to run-model-comparison / simulate-model-comparison /\nstop-model-comparison for clarity. Add simulate-model-comparison.sh\nscript that generates realistic MindTrial log output without API calls,\nenabling end-to-end testing of the voice announcer.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "refactor: rename slash commands and add race simulation",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T01:48:40Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "66e378393345bda128d79254bd5b3a987bf62dd8",
          "body": null,
          "is_bot": false,
          "headline": "Update run-eval-suite.md to default to ci/cd evals",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T01:38:10Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9092c211d236b82f95c2e1a60242907e1350396a",
          "body": "Add Claude Code slash commands that launch model-vs-model evals with a\nlive Vin Scully-style sports commentator calling the play-by-play via\nKokoro TTS. The /loop prompt parses MindTrial's zerolog output to track\nper-model standings, speed comparisons, lead changes, and race completion.\n\n- .claude/c\n[…]\nnds/stop-eval-suite.md: kills eval, finalizes transcript\n- voice_announcer.md: original design spec and reference\n- .gitignore: exclude runtime artifacts (transcript, logs, summary)\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add live voice race announcer for MindTrial evals",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T01:17:27Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "efb66f8a9fe87817002478271a8f008ba2f3aeec",
          "body": "…ents\n\nfeat: Evaluation framework enhancements — CI/CD tasks, cost tracking, CircleCI migration",
          "is_bot": false,
          "headline": "Merge pull request #1 from ryan-circleci/feat/eval-framework-enhancem…",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T00:57:06Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "d5898f5e3070ce76e4232483a0721e08bd53aa80",
          "body": "Include all HTML, CSV, JSONL, and log outputs from:\n- General intelligence eval (71 tasks x 15 models, 2026-03-12)\n- CI/CD eval v1 (19 tasks x 15 models, 2026-03-13)\n- CI/CD eval v2 (100 tasks x 15 models, 2026-03-16)\n\nResults are committed directly since CI is not yet active.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "chore: add evaluation results for general-intelligence and CI/CD runs",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T00:49:39Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "aeb8b49bea996ed84cab23dc43cd8b10460d6ff2",
          "body": "Add 81 new CI/CD evaluation tasks covering CircleCI (25 total),\nGitHub Actions (12), GitLab CI/CD (10), Jenkins (8), Azure DevOps (5),\nshell scripting (11), Docker (9), Kubernetes (7), Git (4), and\ndeployment strategies (5). Mixed difficulty (easy/medium/hard) to\nbetter differentiate model capabilities.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: expand CI/CD benchmark from 19 to 100 tasks",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-16T17:07:46Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "3ac13e13594bf3cde3de60e31024e79483c7dfaa",
          "body": "Add config-eval-top3-cicd.yaml pointing at tasks-cicd.yaml with\nresults directed to results/cicd/ for clean separation from\ngeneral-intelligence eval runs.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add dedicated CI/CD eval config for top-3 providers",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-14T00:22:01Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9515c123b133f82e2113e064a41b087e39866e4e",
          "body": "Made-with: Cursor",
          "is_bot": false,
          "headline": "chore: gitignore compiled binary and evaluation results",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-13T15:45:51Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "cfaae1bd5f5228f074e635368b62c3788a223fca",
          "body": "Add OpenAI Responses API provider (openai_responses.go) for GPT-5.3+\nmodels, which use a different API surface than Chat Completions.\nUpdate the top-3 eval config to include GPT-5.4, GPT-5.4 Pro, and\nrefresh the model lineup across OpenAI, Google, and Anthropic.\nExpand the pricing catalog with latest model prices (GPT-5.4/Pro,\nGemini 3 Flash, Gemini 3.1 Flash-Lite, Claude Opus 4.5/4.6,\nClaude Haiku 4.5, Claude Sonnet 4.6). Bump openai-go SDK to v3.26.0.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add GPT-5.4 Responses API support and update top-3 eval config",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-13T15:44:32Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b4b626aea37958a16e119f547d8a682164c760e5",
          "body": "Focused evaluation config targeting the 5 latest models from each of\nthe top 3 providers (15 models total):\n\n- OpenAI: GPT-5.4 Pro, GPT-5.4 Thinking, GPT-5.3-Codex, GPT-5.3\n  Instant, GPT-5.2\n- Google: Gemini 3.1 Pro, 3.1 Flash-Lite, 3 Flash, 2.5 Pro, 2.5 Flash\n- Anthropic: Claude Opus 4.6, Sonnet 4\n[…]\n Haiku 4.5, Opus 4.5,\n  Sonnet 4.5\n\nUses env var auto-fallback for API keys (OPENAI_API_KEY, GOOGLE_API_KEY,\nANTHROPIC_API_KEY). Judge uses Anthropic Claude for semantic evaluation.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add top-3 provider eval config (OpenAI, Google, Anthropic)",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:57:18Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "1567787d5f21abb899cbb6dadd578a17026ecc03",
          "body": "Two complementary mechanisms for keeping API keys out of config YAML:\n\n1. ${ENV_VAR} expansion: any value in config.yaml can reference an env\n   var using ${VAR_NAME} syntax, expanded before YAML parsing. Unset\n   variables are left as-is.\n\n2. Automatic fallback: if a provider's api-key is empty aft\n[…]\n GOOGLE_API_KEY for\n   google, etc.). Works for both providers and judges.\n\nResolution order: ${...} expansion on raw YAML, then auto-fallback for\nstill-empty keys, then validation.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add environment variable support for API keys",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:56:35Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "afdeeeb7a692d5274fde63ed865f565d4b392326",
          "body": "Extend the CircleCI config with two evaluation workflows:\n\n1. scheduled-eval: Runs every Monday at 06:00 UTC against all configured\n   models, producing HTML, CSV, and JSONL reports stored as artifacts.\n\n2. new-model-eval: API-triggered pipeline for on-demand evaluation when\n   a new model drops. Tr\n[…]\nth use a run-eval job that builds MindTrial, runs the evaluation with\nconfigurable task files, appends JSONL results to a history file, and\nstores all outputs as CircleCI artifacts.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add scheduled and new-model-drop evaluation pipelines",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:23:45Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "8f9823090e36bf0edeb3d6a500fc358869cf4a27",
          "body": "Add a new JSONL (JSON Lines) formatter that outputs one JSON object per\nline, summarizing each provider/run combination with pass rate, duration,\ntoken counts, and estimated cost. Register the formatter with a -jsonl\nCLI flag.\n\nThis format is designed for append-friendly historical result files,\nenabling trend analysis across evaluation runs over time.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add JSONL formatter for historical result tracking",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:23:20Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "df9b365c7dd128466e4b8ac7c400abb6e7b4cae8",
          "body": "Introduce a pricing catalog (pricing/pricing.go) with per-million-token\ncosts for models across OpenAI, Anthropic, Google, DeepSeek, Mistral,\nxAI, Alibaba, Moonshot, and OpenRouter. Add a Model field to RunResult\nso pricing can be looked up at report time.\n\nUpdate all formatters (CSV, HTML, summary \n[…]\nimated Cost columns. Add helper functions in\nformatters/utils.go for token aggregation and cost formatting.\n\nUpdate runner tests and regenerate golden files to match the new output.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add per-task dollar cost calculation",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:23:15Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "169e6877777a81958192300f8beb7a9a5117b4b3",
          "body": "Replace the GitHub Actions workflow (.github/workflows/go.yml) with a\nCircleCI configuration. The build-and-test workflow runs on every\npush/PR, building the project and running the test suite with race\ndetection enabled on cimg/go:1.25.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: migrate CI from GitHub Actions to CircleCI",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:23:05Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b869d9acbdf6d4f13d03cd96f5ecebdfa3fd38fe",
          "body": "Add 19 new evaluation tasks targeting CI/CD and DevOps knowledge:\nshell arithmetic, pipeline stages, semantic versioning, environment\nvariable resolution, CircleCI config debugging, Docker port mapping,\nbuild parallelism, cache key resolution, Kubernetes resource calculation,\nGit branch extraction, \n[…]\nation, shell subshell bugs, and CircleCI orb concepts.\n\nThese complement the existing general intelligence tasks with\ndomain-specific challenges that test practical CI/CD reasoning.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add CI/CD-domain evaluation tasks",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:22:54Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "8446cc90d786989bc8ee083424b9bc0df9720547",
          "body": "Add configurable `max-turns` limit per task to prevent unbounded\nconversation loops (e.g., when a model repeatedly requests exhausted\ntools). The limit is set globally in `task-config` and can be\noverridden per task. A value of 0 means unlimited.\n\nHandle a known Google Gemini issue where the model k\n[…]\nogFinishReason` debug logging to all provider conversation\nloops for improved observability.\n\nExample config:\n\n```\ntask-config:\n  max-turns: 100\n  tasks:\n    - name: \"my task\"\n      max-turns: 200\n```",
          "is_bot": false,
          "headline": "feat: Add conversation turn limit and Gemini tool fix",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:17:14Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "c67c86e37e774d2abf208e5468ec03ec2e595431",
          "body": "Mistral text-only prompts now build Content3.String instead of always\nusing ArrayOfContentChunk, while keeping chunked content for file-based\n(multimodal) prompts. This fixes compatibility with some models,\nwhich rejects chunk arrays for text-only input.\n\nThe Mistral client has been re-generated from the latest Open API spec.\n\nAlso adds structured error logging support across provider errors and\nlogger implementations, with tests for structured field propagation.",
          "is_bot": false,
          "headline": "fix: Use string content for text-only Mistral prompts",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:17:13Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "de5d6d78d4f6b1ee0bc44be5c072df1e71729e46",
          "body": "Add support for the new Gemini 3.1 Pro models and their expanded\nthinking level capabilities in the Google provider.\n\nChanges include:\n- Add `minimal` and `medium` to supported thinking levels.\n- Add `gemini-3.1-pro-preview` and `gemini-3.1-pro-preview-customtools`\n  to the default configuration.\n- Update documentation to clarify that the `minimal` thinking level\n  does not guarantee thinking is completely disabled.",
          "is_bot": false,
          "headline": "feat: Add Gemini 3.1 support",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:17:06Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b02a965a56bb9d998d6fee0c8520e2a50be047eb",
          "body": "Standardize the conversation loop across all AI providers to safely\nhandle multi-turn tool calling and prevent infinite loops.\n\n- Introduce `isTerminalStopReason` to explicitly distinguish between\n  final responses and intermediate tool calls.\n- Add `ErrNoActionableContent` to safely break loops whe\n[…]\n chunks before unmarshaling.\n- Log skipped preamble text during non-terminal turns.\n- Make final answer unmarshal more robust by accepting JSON primitive\n  types stored directly as the answer content.",
          "is_bot": false,
          "headline": "refactor: Standardize provider conversation loops",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:14:20Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "12a188ac50774a86f470fcb343aa9f96dc57f774",
          "body": "Added `stream` configuration option to `AnthropicModelParams`.\n\nImplemented transparent response buffering using the Anthropic SDK's\nAccumulate method to support this mode while maintaining compatibility\nwith existing validation and tool logic.\n\nIntroduced retry support for streaming failures and tr\n[…]\n\nmodel-parameters:\n  max-tokens: 16384\n  effort: max\n  stream: true\n```\n\nRecommended for requests with large `max-tokens` values or extended\nthinking to prevent HTTP timeouts on long-running requests.",
          "is_bot": false,
          "headline": "feat: Add streaming support for Anthropic models",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:14:18Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "32d4d8ee4dcb6b3e0f28c8145215fc7264f3826c",
          "body": "Replace the tool-use workaround (`record_summary` tool with forced\n`tool_choice`) with the native `output_config.format` API for\nstructured JSON responses.\n\nThe old approach was incompatible with extended thinking (which\nrequires `tool_choice: auto`) and conflated the response schema\ntool with actua\n[…]\nt.\n\nAlso batch parallel tool results into a single user message per\nturn, and add an explicit error for responses with no actionable\ncontent (e.g., thinking budget exhausted without producing output).",
          "is_bot": false,
          "headline": "fix: Use native structured output for Anthropic",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:14:18Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "783103daf2e38a9ba3a424cdbed63e3b2c6bdd19",
          "body": "MoonshotAI's thinking models (kimi-k2-thinking, kimi-k2.5) send\na non-standard reasoning_content field that the Open AI SDK\nsilently drops.\n\nIntroduce `CompletionHandler` interface to let delegating providers\ncustomize how streaming chunks are accumulated and how response\nmessages are converted to r\n[…]\nequent conversation\nturns.\n\nSet `max-tokens` to 16000 in default thinking model configs per\nMoonshot AI documentation recommendation.\n\nAlso remove discontinued `kimi-latest` model from default config.",
          "is_bot": false,
          "headline": "fix: Preserve MoonshotAI reasoning content",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:12:00Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "ebd9512657c21f094fdf563244deeb65fe281536",
          "body": "Add runs for OpenRouter (INTELLECT-3, Mercury, Seed 1.6, GLM 4.6V,\nGLM 4.7, Step3), xAI (Grok 4.1 Fast), Alibaba (Qwen3-Max snapshot,\nQVQ-Max, QwQ-Plus), and Moonshot AI (Kimi K2.5).",
          "is_bot": false,
          "headline": "chore: Add new default model run configurations",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-07T15:46:27Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "18055408478e2493eb0b0ebca73921d8afad51cb",
          "body": "Add `effort` parameter to Anthropic provider for adaptive extended\nthinking (values: low, medium, high, max). When set, uses\n`thinking: {type: \"adaptive\"}` with `output_config.effort` instead\nof the deprecated fixed `budget_tokens` approach.\n\nExample config:\n\n```\nmodel-parameters:\n  max-tokens: 8192\n  effort: max\n```",
          "is_bot": false,
          "headline": "feat: Add adaptive thinking support for Claude 4.6",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-06T23:47:28Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "878da54e97e6be63943f087f34c8278d58c5a894",
          "body": "Added `stream` configuration option to `AlibabaModelParams`.\n\nImplemented transparent response buffering in the internal OpenAI\nprovider to support this mode while maintaining compatibility with\nexisting validation and tool logic.\n\nExample config:\n\n```yaml\nmodel-parameters:\n  stream: true\n```\n\nRequired for some Alibaba models, such as QwQ, QVQ, and Qwen-Omni,\nwhich mandate streaming responses.",
          "is_bot": false,
          "headline": "feat: Add streaming support for Alibaba models",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-04T02:57:40Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "990b6529f7e63079ceb5a63c71b91070f44a9fc3",
          "body": "Add `text-only` boolean option to run configurations\nthat skips tasks requiring any file attachments.\n\nExample usage:\n\n```\nruns:\n  - name: \"Text Model\"\n    model: \"model-id\"\n    text-only: true\n```\n\nUseful for text-only models like some OpenRouter-hosted LLMs.\nDefault behavior unchanged (runs all tasks).",
          "is_bot": false,
          "headline": "feat: Add text-only mode to skip tasks with files",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-01T03:21:58Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "8ae8916cb505d7e32689e502b9470bb563ee4fe2",
          "body": "Ensure prompts and usage are propagated with result objects even when\nall retry attempts fail, by capturing the last attempt's result value.",
          "is_bot": false,
          "headline": "fix: Preserve execution metadata on retry failures",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-01T03:17:23Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "ae2885fc12670db3e6483b7fcab57e5f94b06893",
          "body": "Replace the community-based client with the official Open AI SDK.\nThis improves maintainability, ensures compatibility with OpenAI's\nlatest API changes, and provides better long-term support.\n\nKey changes:\n\n- Update dependencies in go.mod and go.sum\n- Refactor OpenAI, Alibaba and MoonshotAI provider\n[…]\nng\n\nAll existing functionality is preserved, and all tests pass. The new\nimplementation handles response formatting, tool calls, and error\nhandling consistently across all OpenAI-compatible providers.",
          "is_bot": false,
          "headline": "feat: Replace OpenAI client with official SDK",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-01T03:16:48Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "7205ab691d1beb9c1734afdf617f55c3d4d8c508",
          "body": "Introduce `disable-structured-output` run-configuration flag to\ntreat model responses as plain text unstructured data,\nbypassing JSON parsing.\nWhen enabled, the entire model's response becomes the final answer,\nwith Title and Explanation populated with placeholders.\n\nKey changes:\n\n- Config: Add Disa\n[…]\nproviders:\n  - name: openai\n    runs:\n      - name: unstructured\n        model: gpt-4o\n        disable-structured-output: true\n```\n\nEnhances compatibility with models unable to generate JSON reliably.",
          "is_bot": false,
          "headline": "feat: Add flag for unstructured responses",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-01T03:12:22Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "885ef526df7c0c87bb77d1d7d720686775db20fc",
          "body": "Remove extra blank line after OpenAI parameters list to maintain\nconsistent style across all provider parameter sections.",
          "is_bot": false,
          "headline": "docs: fix formatting in README parameters section",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-01-17T23:41:08Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "6c5cd533ddf4bcc6e16a0d4b01d48d1b67234aeb",
          "body": "Introduce OpenRouter as a new provider for accessing models through\nthe OpenRouter API. This implementation uses the official OpenAI\nSDK v3, providing a modern and maintainable client foundation that\nmay later replace the community-based client.\n\nThe `OpenRouter` provider can be configured with:\n\n- \n[…]\nities.\n- Adds `Ptr[T]` utility function for creating pointers to values.\n- Includes comprehensive test coverage for parameter mapping.\n- Updates documentation in README.md with configuration examples.",
          "is_bot": false,
          "headline": "feat: Add OpenRouter provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-01-17T19:22:30Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "43c2d878b8164777f4d207e5c07a0010bb84a201",
          "body": "Expose rate metrics in summary outputs to make runs easier\nto compare at a glance.\n\nDefinitions:\n\n- Pass Rate = Passed/(Passed+Failed+Error)\n- Accuracy = Passed/(Passed+Failed)\n- Error Rate = Error/(Passed+Failed+Error)\n- Skipped tasks are excluded; rates are 0 when denominator is 0.\n\nChanges:\n\n- Add Pass Rate, Accuracy, and Error Rate to summary.log output.\n- Add the same columns to the HTML summary table (sortable).",
          "is_bot": false,
          "headline": "feat: Add pass/accuracy/error rates to summary",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-12-23T19:35:43Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "12a69da5ce2aef364d06222d28b3fd4801a5f95c",
          "body": "Extended OpenAI provider to support GPT-5.2 model with new\nreasoning_effort values (none, minimal, low, medium, high, xhigh)\nand verbosity parameter (low, medium, high) for output control.\n\nChanges:\n- Extended reasoning-effort validation to 6 values including new\n  xhigh option for GPT-5.2's maximum\n[…]\nhandling.\n- Added gpt-5.2 model configuration with xhigh reasoning effort\n  and medium verbosity.\n- Updated README documentation with complete parameter details\n  and legacy model compatibility notes.",
          "is_bot": false,
          "headline": "fix: Add support for GPT-5.2 models",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-12-12T18:17:51Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "214540c7855951e7aa8965423a6ecc450125bda8",
          "body": "Extend vision task support to latest Mistral AI models.\n\n- Support mistral-large, mistral-medium, mistral-small variants.\n- Add ministral, pixtral, magistral model families.",
          "is_bot": false,
          "headline": "fix: Update Mistral AI vision model support",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-12-08T00:26:37Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "5243da369e4d53a0fcb581dfeb9d836615d6487e",
          "body": "Add support for Google Gemini 3 models with new API features:\n\n- Add `thinking-level` parameter to control reasoning depth (low/high)\n- Add `media-resolution` parameter for image token allocation control\n- Add `text-response-format-with-tools` for backward compatibility\n  with pre-Gemini 3 models th\n[…]\ned text and tool\n  responses correctly\n\nPre-Gemini 3 models can force text response format with\ntools using `text-response-format-with-tools` parameter.\n\nAlso, fix several minor typos in task prompts.",
          "is_bot": false,
          "headline": "feat: Add Gemini 3 with thinking and media control",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-11-20T05:56:34Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "0755e3331f801643d7868ccae889d4f47870f62f",
          "body": "Introduce Moonshot AI (Kimi) as a new provider via OpenAI-compatible\nAPI. The current provider's implementation delegates request\nprocessing logic to the existing `OpenAI` provider.\n\nThe `Moonshot AI` provider can be configured with the following\nproperties:\n\n- `name`:     Must be set to \"moonshotai\n[…]\nes.\n  `LegacyJsonSchema`, adds format instruction to prompt while keeping\n  json_schema response format.\n  `LegacyJsonObject`, adds format instruction to prompt and uses\n  json_object response format.",
          "is_bot": false,
          "headline": "feat: Add Moonshot AI (Kimi) provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-11-12T21:39:53Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "e2678265b465416665fa3f8a429c84571081b045",
          "body": "Add support for persistent shared directories that enable data\nsharing across all tool invocations within a single task:\n\n- Add `shared-dir` field to `ToolConfig` for directory configuration.\n- Implement lazy directory creation on the first use.\n- Mount the shared directory to `shared-dir` path insi\n[…]\ncutor\n      shared-dir: /app/shared\n      auxiliary-dir: /app/data\n```\n\nFiles in `shared-dir` persist across all tool calls within a task,\nwhile `auxiliary-dir` contents are reset between invocations.",
          "is_bot": false,
          "headline": "feat: Add persistent shared directory for tool use",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-11T01:09:29Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "f7b296e00be0d6a52168d97e69b85c3956592833",
          "body": "Add validation to ensure Docker images for enabled tools\nare available locally before tasks begin:\n\n- Add ValidateTool method to DockerToolExecutor to check image\n  availability via Docker API.\n- Integrate validation in `defaultRunner` to validate\n  all enabled tools before execution starts.\n- Updat\n[…]\ndescription to clarify ephemeral container behavior.\n\nValidation prevents runtime failures by catching missing images\nearly, providing clear guidance on how to resolve issues before\nany tasks execute.",
          "is_bot": false,
          "headline": "feat: Validate tool images before task execution",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-11T01:04:43Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "232969edf87a3f63371174b9ed3c047ec89a714b",
          "body": "- Mount copies of all task files into `auxiliary-dir` when set.\n- Rename `file_mappings` property to `parameter-files` for consistency,\n  and to better distinguish it from auxiliary files.\n- Timeout now covers only the container run, not setup/cleanup.",
          "is_bot": false,
          "headline": "feat: Mount task files in tool container",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-09T18:05:48Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9b977533237bb132a48a498b2c0389799e1738a3",
          "body": "- Introduces unique TraceID (ULID) for each task execution result.\n  TraceIDs enable tracing and correlation across artifacts and logs.\n  CSV and text formatters include TraceID column; HTML shows\n  indicator icon with TraceID tooltip; logger prepends TraceID\n  to all result messages.\n\n- Adds tool usage indicators to HTML output with sorted tool\n  names and visual icon.",
          "is_bot": false,
          "headline": "feat: Add ID and tool usage indicators to results",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-03T00:05:30Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "2dd05e474b3138368cdb18564d6ec71a751ce61d",
          "body": "- Add Claude 4.5 Sonnet model with extended thinking\n- Update DeepSeek model names from V3.1 to V3.2",
          "is_bot": false,
          "headline": "fix: Add recent models to default config file",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-02T15:35:36Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "5ffea0b8745aaf55910a6eaf4accb13daa6a6035",
          "body": "- Add global tools config and per-task tool-selector.\n- Validate tool parameter schemas and task tool references.\n- Update all providers to support tool calling with conversation loops.\n- Implement DockerToolExecutor for secure tool execution in containers.\n- Track tool usage statistics (call count,\n[…]\napp/main.py\n```\n\nExample tool configuration:\n\n```yaml\ntool-selector:\n  tools:\n    - name: python-code-executor\n      max-calls: 10\n      timeout: 60s\n      max-memory-mb: 512\n      cpu-percent: 25\n```",
          "is_bot": false,
          "headline": "feat: Enable tool use in tasks",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-02T15:35:25Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "2f4a3a187dd9dea1e5804d547d936e1580464c87",
          "body": "Introduce Alibaba (Qwen) as a new provider via OpenAI-compatible API.\nThe current provider's implementation delegates the request processing\nlogic to the existing `OpenAI` provider and may not support all models\nand all options.\n\nThe `Alibaba` provider can be configured with the following properties\n[…]\nokens\n                             available to the model for generating\n                             a response.\n- Extend test coverage of `Google` model parameters.\n- Fix `DeepSeek` name references.",
          "is_bot": false,
          "headline": "feat: Add Alibaba (Qwen) provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-23T21:16:27Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "ea8a0ff0c574b4f95eaaec38f180b47f4f14eedd",
          "body": "- Add customizable judge prompt template with structured context.\n- Add verdict format and passing verdicts configuration.\n- Add default judge template with semantic equivalence logic.\n- Implement per-task validation rules resolution and caching.\n- Fix support for mapping objects in value sets.\n- En\n[…]\n      enum: [\"excellent\", \"good\", \"poor\"]\n      required: [\"quality_score\"]\n      additionalProperties: false\n    passing-verdicts:\n      - quality_score: \"excellent\"\n      - quality_score: \"good\"\n```",
          "is_bot": false,
          "headline": "feat: Judge validation with custom prompt template",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-19T16:11:27Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "d11edbc12141282ea9cafd6b2e86f98eaaa5168c",
          "body": "- `quiz - multiple choice questions - v1`:\n    accept answers without question numbers, as long as the choices are\n    correct and in sequence. (Anthropic, DeepSeek, Mistral AI)",
          "is_bot": false,
          "headline": "fix: Adjust task validation for model edge cases",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-16T18:47:58Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "ea24056e0a2f91ae2d635d16295acfdab7ae86fc",
          "body": "- Add support for structured schema-based result formats using JSON\n  schemas in task definitions alongside existing plain text formats.\n- Configure system prompt delivery per result format type with new\n  'enable-for' setting:\n  - `all`: Send for all tasks.\n  - `text`: Send for tasks with plain tex\n[…]\n Seed for deterministic generation.\n\nExample structured answer format:\n\n```\n  response-result-format:\n    type: object\n    properties:\n      answer: {type: string}\n      confidence: {type: number}\n```",
          "is_bot": false,
          "headline": "feat: Add structured JSON result format support",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-16T18:35:20Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "d9cb955aa4a36d45c7c0417391fde8705bdd3892",
          "body": "Refresh dependencies in go.mod.",
          "is_bot": false,
          "headline": "fix: Update dependencies",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-03T14:00:21Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b64109916b136ef09bc700157b09a99390eb5243",
          "body": "Fill upper triangle of run comparison matrix with disagreement\npercentages between runs. Lower triangle continues to show\nagreement data, creating a symmetrical view of run relationships.\n\nClicking disagreement cells opens a modal with tasks where the\ntwo runs produced different results. Updated legend reflects\nthe new matrix layout and color coding.",
          "is_bot": false,
          "headline": "feat: Add disagreement data to comparison matrix",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-03T13:59:10Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "42f26988b97f59a1b58459434b297ffe6347ad52",
          "body": "Add support for customizable system prompt templates in task\nconfiguration to control how response format instructions are\npresented to AI models:\n\n- Add SystemPrompt struct with Template field to config package\n- Support global system prompt template in task-config section\n- Allow per-task system p\n[…]\nhe final answer in exactly this format:\n        {{.ResponseResultFormat}}\n    tasks:\n      - name: \"my-task\"\n        system-prompt:\n          template: \"Answer format: {{.ResponseResultFormat}}\"\n  ```",
          "is_bot": false,
          "headline": "feat: Add configurable system prompt templates",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-03T13:58:52Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "dde5cfc4eef2682e01787c11449b52706317d33c",
          "body": "- Update the minimum required Go version to match the ```go.mod``` file.\n- Add missing blockquote.",
          "is_bot": false,
          "headline": "docs: Update README.md",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-28T16:43:00Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "4d463de2ab876596d93a0810d44e756256d42c2f",
          "body": "Add support for xAI (Grok) models as a new provider option.\n\nThe xAI provider can be configured with the following properties:\n\n  - `name`:    Must be set to \"xai\"\n  - `api-key`: API key for the xAI (Grok) models provider\n\nSupported model-specific parameters include:\n\n  - `temperature`: Controls ran\n[…]\nrok 4 is a reasoning model.\n  - ```presence-penalty``` and ```frequency-penalty``` parameters\n    are not supported by reasoning models.\n  - Grok 4 does not support a ```reasoning-effort``` parameter.",
          "is_bot": false,
          "headline": "feat: Add xAI (Grok) provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-28T14:57:25Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "d451d66def0472504548a75ca9f841c55fe7185a",
          "body": "Update response-result-format specifications in default\n```tasks.yaml``` to reduce model response formatting errors and\nimprove task validation accuracy.",
          "is_bot": false,
          "headline": "fix: Clarify answer format to reduce task failures",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-26T16:25:26Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "112b079de234d8a2a38b482ee6d7007e97f4c562",
          "body": "Align run naming with current DeepSeek models and support both\nthinking and non-thinking modes.\n\n- Rename \"DeepSeek-R1 - latest\" to \"DeepSeek-V3.1 - latest (thinking\n  mode)\" keeping deepseek-reasoner model.\n- Add \"DeepSeek-V3.1 - latest (non-thinking mode)\" using\n  deepseek-chat model.",
          "is_bot": false,
          "headline": "fix: Support DeepSeek V3.1 thinking modes",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-26T00:09:43Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "a585c5a94a470aa4b040621eeb0f34d8dec1f1dc",
          "body": "Enable clicking run comparison matrix cells to open a dialog\nlisting overlapping tasks grouped by status.",
          "is_bot": false,
          "headline": "feat: Add run details dialog to comparison matrix",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-25T20:30:15Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "8425f9e403ec8ed240ed2decf0f72b52cf649a3f",
          "body": "The summary table in the HTML report now includes checkboxes to select\nmultiple runs.\nA \"Compare Selected Runs\" button uses this selection to generate a\nvisual comparison matrix in a new window.\n\nThe matrix shows the percentage of agreement between any two runs\non their common, non-skipped tasks.\nCe\n[…]\nor-coded to show the distribution of matching results\n(passed, failed, error), with saturation indicating overlap strength.\nThe matrix is interactive, with hover effects to highlight rows and\ncolumns.",
          "is_bot": false,
          "headline": "feat: Add run comparison matrix to HTML report",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-21T20:16:13Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "71cf5a7bb8a680839589dfd633362e258fac7e94",
          "body": "Prevent duplicate names across configurations to avoid ambiguous\nresults and improve clarity in reports.\n\n- Enforce unique names within provider runs.\n- Enforce unique names within judge configurations.\n- Enforce unique names within task definitions.",
          "is_bot": false,
          "headline": "feat: Enforce unique name for runs, judges & tasks",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-21T20:16:13Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "64bde07c899ab893ff1a56f7bfd7a35994433c9e",
          "body": "Rename duplicate ```riddle - split words - v3``` to ```riddle - split words - v4```.",
          "is_bot": false,
          "headline": "fix: Fix duplicate task name",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-21T20:16:12Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "e281f8c760a646d193abdb58ad0a129f68f69ce6",
          "body": "- `quiz - multiple choice questions - v1`: tolerate trailing\n  whitespace in individual options from some models (Anthropic).\n- `quiz - analogies`: accept both verb and noun answers\n  (\"eat\" / \"food\") since \"sleep\" functions as both.\n- `riddle - first letter - v3`: fix format description\n  (4 groups = 4 letters), and add an alternative valid solution\n  with a new first letter for group #2.",
          "is_bot": false,
          "headline": "fix: Adjust task validation for model edge cases",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-16T17:01:31Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "4683f60b98e671a9840a7f57b83647165dfc1e69",
          "body": "Introduces `trim-lines` validation rule, token usage tracking,\nand enhanced HTML reports with regex filtering capabilities.\n\nNew Validation Rule:\n\n- `trim-lines` (boolean, default: false) - Trims leading/trailing\n  whitespace from each line while preserving internal spaces and\n  normalizing CRLF to \n[…]\nsections.\n- Improved filtering tooltip text with regex examples.\n\nToken Usage Tracking:\n\n- Added `TokenUsage` struct to run result details.\n- Displays usage statistics in HTML reports and CSV exports.",
          "is_bot": false,
          "headline": "feat: Add trim-lines, token usage, enhanced HTML",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-16T15:55:15Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "990411b06017959ea855ba5c9aecd85534f655f7",
          "body": "Retry on documented transient errors from OpenAI API requests.",
          "is_bot": false,
          "headline": "fix: Enhance OpenAI provider with transient error handling",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-11T18:11:56Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "1e9743f238598ef5616588e622225551ffe0c776",
          "body": "- Add OpenAI o4-mini, o3, GPT-5 mini, and GPT-5 with high reasoning\n  effort.\n- Add Google Gemini 2.5 Flash and Pro models.\n- Add Anthropic Claude 4.1 Opus model.\n- Increase rate limits for o1-mini and o3-mini to 20 requests/minute.\n- Increase request timeout for DeepSeek to 15 minutes.",
          "is_bot": false,
          "headline": "fix: Update default model list and rate limits",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-11T18:11:55Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "d29677a380b50363f3b58fa081534c42e55dfc12",
          "body": "- Add structured Details (Answer, Validation, Error) to RunResult.\n- CSV-only: Serialize Details as structured JSON.\n- HTML-only: Enrich HTML report with structured details section.\n- HTML-only: Add header filtering and click-to-filter functionality.\n- HTML-only: Add search by provider, run, and task names.\n- HTML-only: Enable column sorting.\n- HTML-only: Improve accessibility and add schema.org annotations.",
          "is_bot": false,
          "headline": "feat: Introduce structured Details and enhance HTML report",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-11T18:11:30Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "27fb07a208067e47ca8b3201e32575a244454314",
          "body": "- Regenerate Go models from Mistral AI OpenAPI spec.\n- Refresh dependencies in go.mod.",
          "is_bot": false,
          "headline": "fix: update Mistral AI client and dependencies",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-10T15:45:56Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "ae2080e7f1e3c4ab622540dbcf17f4e27825b460",
          "body": "This commit introduces a major refactoring to decouple task execution logic\nfrom runners and validators, and to standardize logging throughout the\napplication.\n\n- A new execution package centralizes provider execution, handling\n  retries and rate limiting through a unified Executor.\n- Default runner\n[…]\nnner is simplified by utilizing a logging.Logger\n  instance to handle logging and message emitting.\n- Providers and validators now receive a logging.Logger\n  instance, allowing for contextual logging.",
          "is_bot": false,
          "headline": "refactor: Decouple execution logic and standardize logging",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-30T03:28:50Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "28564435feba3d3ae8e01a77ca8a86ba1a17b07d",
          "body": "Add comprehensive validation framework supporting both exact value\nmatching and LLM-based semantic evaluation:\n\n- Decouple validation logic from providers and move it to a\n  new `validators` package.\n- Introduce value-match validator for traditional value matching.\n- Add judge validator using LLM pr\n[…]\nmodel: \"gpt-4o-mini\"\n\t       max-requests-per-minute: 10\n  ```\n\nJudge selector example:\n  ```yaml\n     validation-rules:\n       judge:\n         enabled: true\n\t name: \"my-judge\"\n\t variant: \"fast\"\n  ```",
          "is_bot": false,
          "headline": "feat: Add LLM judge validation feature",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-30T03:20:14Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "5354323c42d8c48bcf186a7453320e04bec6b7bf",
          "body": "Adds post-processing to ensure proper Go formatting with gofmt.\nNote: File paths must not contain spaces due to known generator issue:\nhttps://github.com/OpenAPITools/openapi-generator/issues/10839",
          "is_bot": false,
          "headline": "style: Format Go files generated by openapi-generator using gofmt",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-08T19:46:19Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "3c02134f7935c67e049047358aee051a2fc40c3c",
          "body": "Add support for image processing tasks in the Mistral AI provider.\nThe provider now supports uploading and processing image files with\nvision-capable models.\n\nSupported vision models:\n\n  - pixtral-12b-latest\n  - pixtral-large-latest\n  - mistral-medium-latest\n  - mistral-small-latest",
          "is_bot": false,
          "headline": "feat: Add vision support to Mistral AI provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-08T19:05:01Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "697bed78a77dc98adbbba1d544d8d2de5c7c91ee",
          "body": "Use v0.4.1 instead.",
          "is_bot": false,
          "headline": "fix: Retract v0.4.0 due to broken history",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-07T17:30:18Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9505c586464848c95c0f78aae675002d1b86a26b",
          "body": "Add support for automatic retry of failed requests due to rate limiting\nor other transient errors.\nA retry policy can be configured at the provider level to apply to all\nruns, or at the individual run level to override the provider setting.\n\nThe retry policy defines the following properties:\n\n  - `m\n[…]\n backoff starting with the initial delay.\n\nRate limiting is respected during retries - if a run configuration has\n`max-requests-per-minute` set, the rate limiter will be applied to each\nretry attempt.",
          "is_bot": false,
          "headline": "feat: Add configurable retry policy",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-07T16:48:18Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "debb2c06fb5e40437ab9a89fa9f88274900e5d6c",
          "body": "Add support for Mistral AI models as a new provider option.\n\nThe Mistral AI provider can be configured with the following properties:\n\n  - `name`:    Must be set to \"mistralai\"\n  - `api-key`: API key for the Mistral AI generative models provider\n\nSupported model-specific parameters include:\n\n  - `te\n[…]\n  - `prompt-mode`: When set to \"reasoning\", instructs model to reason\n                   if supported\n  - `safe-prompt`: Enables content filtering for compliance with\n                   usage policies",
          "is_bot": false,
          "headline": "feat: Add Mistral AI provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-07T16:47:13Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "b2ca7f427c612ede5b9e8a12dae305d351e79f99",
          "body": "- Add fifteen visual tasks.",
          "is_bot": false,
          "headline": "fix: Add new visual tasks",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-06-27T16:30:55Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "671a19a85316cd43fc0234ef0a80255e9b04ae82",
          "body": "- Ensure interactive configuration selector scrolls to keep\n  cursor visible when navigating beyond the viewport bounds.\n- Improve window resize logic for consistent TUI layout adjustments.",
          "is_bot": false,
          "headline": "fix: Add scrolling to config selector and improve resize layout",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-06-27T16:27:56Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "3168044731ea9ec82fec88e97b90b7228c814f4a",
          "body": "- Add new mostly visual tasks.\n- Ensure the images have appropriate dimensions\n  to minimize token usage (adopts common tile\n  sizes of 384px or 512px).\n- Move the prompt text after the file data for\n  improved context integrity.",
          "is_bot": false,
          "headline": "fix: Improve and extend visual task coverage",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-06-18T04:38:04Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "22da64fec18f2782ad41ac5d4ec2619ef6df7283",
          "body": "- Add alternative valid answer to `riddle - web words - v2`.\n- Handle common edge cases in result validation for specific models\n  (`o4-mini` dropping spaces in visual task, `Claude 4.0` including answer\n  text in quiz responses).",
          "is_bot": false,
          "headline": "fix: Improve task validation for model-specific output formats",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-31T13:41:35Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b989ac53e22ca8d10b18a7d078a07f08b3c5062f",
          "body": "Introduces `validation-rules` at both the global `task-config` level and\nper `task`.\nThese rules allow customization of how model-generated answers are\ncompared against expected results.\n\nSupported Rules:\n\n- `case-sensitive` (boolean, default: false)\n- `ignore-whitespace` (boolean, default: false)\n\nTask-specific rules override global settings.",
          "is_bot": false,
          "headline": "feat: Add validation rules for flexible answer checking",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-31T02:10:43Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "c8f4657c17010f2aaa65caa10b26f2190cc97bf0",
          "body": "- Add `Claude Opus 4` and `Claude Sonnet 4` models to default\n  configuration file.\n- Increase default `max-requests-per-minute` for Claude models.\n- Update Anthropic provider dependency.\n- Refactor Anthropic provider implementation to align with the updated API.",
          "is_bot": false,
          "headline": "fix: Add latest `Claude 4` models to default configuration",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-24T02:21:21Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "1d5ad621cd9625f5ff2af7f2d83818b292318a9f",
          "body": "Change `OnceWithContext` closures to receive *TaskFile as parameter\ninstead of capturing `basePath` from enclosing scope. This ensures\n`Content()` uses the current `basePath` value after `SetBasePath()` calls.",
          "is_bot": false,
          "headline": "fix: Resolve issue where attached file could not be opened",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-24T00:54:10Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "a2f7f5d8f963fcb418e4e88498418eaa887ee258",
          "body": "Replace the previous `time.Tick` based rate limiting implementation\nwith `golang.org/x/time/rate.Limiter`.\n\nThis change provides a more robust and flexible approach to rate\ncontrol. The new `rate.Limiter` allows for an initial burst of tasks up\nto the `MaxRequestsPerMinute` limit and then ensures the rate is\nmaintained. This improves throughput by executing tasks as quickly as\npossible within the defined limits.",
          "is_bot": false,
          "headline": "perf: Use `x/time/rate` for rate limiting",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-19T18:49:36Z",
          "body_truncated": false,
          "is_coding_agent": false
        }
      ],
      "releases_count": 33,
      "commits_last_year": 86,
      "latest_release_at": "2026-03-18T22:43:28Z",
      "latest_release_tag": "v0.17.0",
      "releases_from_tags": true,
      "days_since_last_push": 127,
      "active_weeks_last_year": 21,
      "days_since_latest_release": 128,
      "mean_days_between_releases": 13.2
    },
    "community": {
      "has_readme": false,
      "has_license": false,
      "has_description": false,
      "has_contributing": false,
      "health_percentage": null,
      "has_issue_template": false,
      "has_code_of_conduct": false,
      "has_pull_request_template": false
    },
    "ecosystem": {
      "packages": [
        {
          "name": "github.com/CircleCI-Research/evalbench",
          "exists": true,
          "license": null,
          "keywords": [],
          "ecosystem": "go",
          "matches_repo": true,
          "registry_url": "https://pkg.go.dev/github.com/CircleCI-Research/evalbench",
          "is_deprecated": false,
          "latest_version": "v0.17.0",
          "repository_url": "https://github.com/CircleCI-Research/evalbench",
          "versions_count": 33,
          "total_downloads": null,
          "dependents_count": null,
          "deprecation_note": null,
          "maintainers_count": null,
          "monthly_downloads": null,
          "first_published_at": null,
          "latest_published_at": "2026-03-18T22:43:28Z",
          "latest_version_yanked": null,
          "days_since_latest_publish": 128
        }
      ]
    },
    "popularity": {
      "forks": 0,
      "stars": 0,
      "watchers": 0,
      "fork_history": {
        "days": [],
        "complete": true,
        "collected": 0,
        "total_forks": 0
      },
      "star_history": {
        "days": [],
        "complete": true,
        "collected": 0,
        "total_stars": 0,
        "collected_at": null
      },
      "open_issues_and_prs": 0
    },
    "ai_readiness": {
      "has_nix": false,
      "example_dirs": [],
      "has_llms_txt": false,
      "has_dockerfile": false,
      "has_mcp_signal": false,
      "bootstrap_files": [],
      "api_schema_files": [],
      "has_devcontainer": false,
      "typecheck_configs": [],
      "toolchain_manifests": [
        "go.mod"
      ],
      "largest_source_bytes": 73199,
      "source_files_sampled": 442,
      "oversized_source_files": 5,
      "agent_instruction_files": [
        "AGENTS.md"
      ],
      "agent_instruction_max_bytes": 3085
    },
    "dependencies": {
      "manifests": [
        "go.mod"
      ],
      "advisories": {
        "error": null,
        "scope": null,
        "source": null,
        "findings": [],
        "collected": false,
        "malicious": [],
        "truncated": false,
        "by_severity": {},
        "advisory_count": 0,
        "affected_count": 0,
        "assessed_count": 0,
        "malicious_count": 0,
        "assessed_package": null,
        "unassessed_count": 0,
        "direct_affected_count": 0
      },
      "ecosystems": [
        "go"
      ],
      "dependencies": [
        {
          "name": "github.com/charmbracelet/x/term",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.2.1"
        },
        {
          "name": "github.com/containerd/errdefs",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.0.0"
        },
        {
          "name": "github.com/docker/docker",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v28.5.2+incompatible"
        },
        {
          "name": "github.com/google/uuid",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.6.0"
        },
        {
          "name": "github.com/invopop/jsonschema",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.13.0"
        },
        {
          "name": "github.com/kaptinlin/jsonrepair",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.2.8"
        },
        {
          "name": "github.com/oklog/ulid/v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.1.1"
        },
        {
          "name": "github.com/openai/openai-go/v3",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v3.26.0"
        },
        {
          "name": "github.com/santhosh-tekuri/jsonschema/v6",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v6.0.2"
        },
        {
          "name": "github.com/sethvargo/go-retry",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.3.0"
        },
        {
          "name": "github.com/stretchr/testify",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.11.1"
        },
        {
          "name": "golang.org/x/time",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.14.0"
        },
        {
          "name": "google.golang.org/genai",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.48.0"
        },
        {
          "name": "gopkg.in/validator.v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.0.1"
        },
        {
          "name": "gopkg.in/yaml.v3",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v3.0.1"
        },
        {
          "name": "github.com/anthropics/anthropic-sdk-go",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.26.0"
        },
        {
          "name": "github.com/charmbracelet/bubbles/v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.0.0-beta.1"
        },
        {
          "name": "github.com/charmbracelet/bubbletea/v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.0.0-beta.1"
        },
        {
          "name": "github.com/charmbracelet/lipgloss/v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.0.0-beta.1"
        },
        {
          "name": "github.com/cohesion-org/deepseek-go",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.3.3"
        },
        {
          "name": "github.com/go-playground/validator/v10",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v10.30.1"
        },
        {
          "name": "github.com/rs/zerolog",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.34.0"
        },
        {
          "name": "github.com/sergi/go-diff",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.4.0"
        },
        {
          "name": "golang.org/x/exp",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.0.0-20260218203240-3dfff04db8fa"
        }
      ],
      "all_dependencies": {
        "error": "GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository",
        "source": null,
        "packages": [],
        "collected": false,
        "truncated": false,
        "total_count": null,
        "direct_count": null,
        "indirect_count": null
      }
    },
    "maintainership": {
      "issues": {
        "open_prs": 0,
        "merged_prs": 1,
        "open_issues": 0,
        "closed_ratio": null,
        "closed_issues": 0,
        "closed_unmerged_prs": 0
      },
      "bus_factor": 1,
      "bot_contributors": 0,
      "top_contributors": [
        {
          "type": "User",
          "login": "petmal",
          "commits": 79,
          "avatar_url": "https://avatars.githubusercontent.com/u/4350408?v=4"
        },
        {
          "type": "User",
          "login": "ryan-circleci",
          "commits": 37,
          "avatar_url": "https://avatars.githubusercontent.com/u/104376313?v=4"
        }
      ],
      "contributors_sampled": 2,
      "top_contributor_share": 0.681
    },
    "quality_signals": {
      "has_ci": true,
      "has_tests": true,
      "ci_workflows": [
        "go.yml"
      ],
      "has_docs_dir": false,
      "linter_configs": [],
      "has_editorconfig": false,
      "has_linter_config": false,
      "has_precommit_config": false
    },
    "security_signals": {
      "lockfiles": [
        "go.sum"
      ],
      "scorecard": {
        "checks": [
          {
            "name": "Binary-Artifacts",
            "score": 10,
            "reason": "no binaries found in the repo",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#binary-artifacts"
          },
          {
            "name": "Branch-Protection",
            "score": 5,
            "reason": "branch protection is not maximal on development and all release branches",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#branch-protection"
          },
          {
            "name": "CI-Tests",
            "score": 10,
            "reason": "1 out of 1 merged PRs checked by a CI test -- score normalized to 10",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#ci-tests"
          },
          {
            "name": "CII-Best-Practices",
            "score": 0,
            "reason": "no effort to earn an OpenSSF best practices badge detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#cii-best-practices"
          },
          {
            "name": "Code-Review",
            "score": 0,
            "reason": "Found 0/26 approved changesets -- score normalized to 0",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#code-review"
          },
          {
            "name": "Contributors",
            "score": 6,
            "reason": "project has 2 contributing companies or organizations -- score normalized to 6",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#contributors"
          },
          {
            "name": "Dangerous-Workflow",
            "score": 10,
            "reason": "no dangerous workflow patterns detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#dangerous-workflow"
          },
          {
            "name": "Dependency-Update-Tool",
            "score": 0,
            "reason": "no update tool detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#dependency-update-tool"
          },
          {
            "name": "Fuzzing",
            "score": 0,
            "reason": "project is not fuzzed",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#fuzzing"
          },
          {
            "name": "License",
            "score": 10,
            "reason": "license file detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#license"
          },
          {
            "name": "Maintained",
            "score": 0,
            "reason": "0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#maintained"
          },
          {
            "name": "Packaging",
            "score": null,
            "reason": "packaging workflow not detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#packaging"
          },
          {
            "name": "Pinned-Dependencies",
            "score": 0,
            "reason": "dependency not pinned by hash detected -- score normalized to 0",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#pinned-dependencies"
          },
          {
            "name": "SAST",
            "score": 0,
            "reason": "SAST tool is not run on all commits -- score normalized to 0",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#sast"
          },
          {
            "name": "Security-Policy",
            "score": 0,
            "reason": "security policy file not detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#security-policy"
          },
          {
            "name": "Signed-Releases",
            "score": null,
            "reason": "no releases found",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#signed-releases"
          },
          {
            "name": "Token-Permissions",
            "score": 0,
            "reason": "detected GitHub workflow tokens with excessive permissions",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#token-permissions"
          },
          {
            "name": "Vulnerabilities",
            "score": 0,
            "reason": "44 existing vulnerabilities detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#vulnerabilities"
          }
        ],
        "commit": "64b9f020db13d4629897badb64f94f85a8003544",
        "ran_at": "2026-07-25T08:58:05Z",
        "aggregate_score": 3,
        "scorecard_version": "v5.5.0"
      },
      "has_codeql_workflow": false,
      "has_security_policy": false,
      "has_dependabot_config": false
    },
    "contribution_flow": {
      "collected": true,
      "ci_last_run_at": "2026-03-19T15:13:01Z",
      "oldest_open_prs": [],
      "last_merged_pr_at": "2026-03-19T15:10:06Z",
      "ci_last_conclusion": "SUCCESS",
      "oldest_open_issues": []
    }
  },
  "config": {
    "disabled_metrics": [],
    "disabled_categories": [],
    "disabled_components": {}
  },
  "source": {
    "url": "https://github.com/CircleCI-Research/evalbench",
    "host": "github.com",
    "name": "evalbench",
    "owner": "CircleCI-Research"
  },
  "metrics": {
    "overall": {
      "key": "overall",
      "band": "at_risk",
      "name": "Overall health",
      "note": null,
      "notes": [],
      "value": 43,
      "inputs": {
        "security": 30,
        "vitality": 56,
        "community": 12,
        "governance": 57,
        "engineering": 51
      },
      "components": []
    },
    "categories": [
      {
        "key": "vitality",
        "band": "moderate",
        "name": "Vitality",
        "value": 56,
        "weight": 0.22,
        "metrics": [
          {
            "key": "development_activity",
            "band": "at_risk",
            "name": "Development activity",
            "note": null,
            "notes": [],
            "value": 42,
            "inputs": {
              "commits_last_year": 86,
              "human_commit_share": 1,
              "days_since_last_push": 127,
              "active_weeks_last_year": 21
            },
            "components": [
              {
                "key": "push_recency",
                "name": "Push recency",
                "detail": "last push 127 days ago",
                "points": 9.9,
                "status": "partial",
                "details": [
                  {
                    "code": "push_recency",
                    "params": {
                      "days": 127
                    }
                  }
                ],
                "max_points": 36
              },
              {
                "key": "commit_cadence",
                "name": "Commit cadence",
                "detail": "21/52 weeks with commits",
                "points": 14.5,
                "status": "partial",
                "details": [
                  {
                    "code": "commit_cadence_weeks",
                    "params": {
                      "weeks": 21
                    }
                  }
                ],
                "max_points": 36
              },
              {
                "key": "commit_volume",
                "name": "Commit volume",
                "detail": "86 commits in the last year",
                "points": 17.4,
                "status": "partial",
                "details": [
                  {
                    "code": "commits_last_year",
                    "params": {
                      "count": 86
                    }
                  }
                ],
                "max_points": 18
              },
              {
                "key": "openssf_scorecard_maintained",
                "name": "OpenSSF Scorecard: Maintained",
                "detail": "0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 10
              }
            ]
          },
          {
            "key": "release_discipline",
            "band": "good",
            "name": "Release discipline",
            "note": "Excluded from scoring (no data or not applicable): OpenSSF Scorecard: Signed-Releases. Remaining weights renormalized.",
            "notes": [
              {
                "code": "excluded_no_data",
                "params": {
                  "components": [
                    "openssf_scorecard_signed_releases"
                  ]
                }
              },
              {
                "code": "weights_renormalized",
                "params": {}
              }
            ],
            "value": 78,
            "inputs": {
              "releases_count": 33,
              "latest_release_tag": "v0.17.0",
              "releases_from_tags": true,
              "days_since_latest_release": 128,
              "mean_days_between_releases": 13.2
            },
            "components": [
              {
                "key": "ships_releases",
                "name": "Ships releases",
                "detail": "33 version tags (no GitHub releases)",
                "points": 16.2,
                "status": "partial",
                "details": [
                  {
                    "code": "version_tags_no_releases",
                    "params": {
                      "count": 33
                    }
                  }
                ],
                "max_points": 27
              },
              {
                "key": "release_recency",
                "name": "Release recency",
                "detail": "latest release 128 days ago",
                "points": 27,
                "status": "partial",
                "details": [
                  {
                    "code": "release_recency",
                    "params": {
                      "days": 128
                    }
                  }
                ],
                "max_points": 36
              },
              {
                "key": "release_cadence",
                "name": "Release cadence",
                "detail": "a release every ~13.2 days",
                "points": 27,
                "status": "met",
                "details": [
                  {
                    "code": "release_cadence",
                    "params": {
                      "gap": 13.2
                    }
                  }
                ],
                "max_points": 27
              },
              {
                "key": "openssf_scorecard_signed_releases",
                "name": "OpenSSF Scorecard: Signed-Releases",
                "detail": "no releases found",
                "points": 0,
                "status": "excluded",
                "details": [
                  {
                    "code": "no_data",
                    "params": {}
                  }
                ],
                "max_points": 10
              }
            ]
          },
          {
            "key": "abandonment",
            "band": "excellent",
            "name": "Abandonment",
            "note": null,
            "notes": [],
            "value": 100,
            "inputs": {
              "cap": null,
              "state": "unverified",
              "guards": [],
              "signals": [],
              "red_flag": false,
              "multiplier_pct": 100,
              "declared_reason": null,
              "unverified_reason": "repository_too_young",
              "unanswered_open_prs": null,
              "unanswered_open_issues": null,
              "days_since_last_merged_pr": null,
              "days_since_last_human_commit": null,
              "days_since_last_human_commit_is_floor": false
            },
            "components": [
              {
                "key": "project_is_still_maintained",
                "name": "Project is still maintained",
                "detail": "maintenance record not established from the collected data",
                "points": 100,
                "status": "met",
                "details": [
                  {
                    "code": "abandonment_unverified",
                    "params": {}
                  }
                ],
                "max_points": 100
              }
            ]
          }
        ],
        "description": "Is the project alive — is code being written and are releases shipping?"
      },
      {
        "key": "community",
        "band": "critical",
        "name": "Community & Adoption",
        "value": 12,
        "weight": 0.18,
        "metrics": [
          {
            "key": "popularity",
            "band": "critical",
            "name": "Popularity & adoption",
            "note": null,
            "notes": [],
            "value": 1,
            "inputs": {
              "forks": 0,
              "stars": 0,
              "watchers": 0,
              "growth_state": "unverified",
              "growth_factor_pct": 100,
              "growth_unverified_reason": "no_history"
            },
            "components": [
              {
                "key": "stars",
                "name": "Stars",
                "detail": "0 stars",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "stars",
                    "params": {
                      "count": 0
                    }
                  }
                ],
                "max_points": 60
              },
              {
                "key": "forks",
                "name": "Forks",
                "detail": "0 forks",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "forks",
                    "params": {
                      "count": 0
                    }
                  }
                ],
                "max_points": 25
              },
              {
                "key": "watchers",
                "name": "Watchers",
                "detail": "0 watchers",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "watchers",
                    "params": {
                      "count": 0
                    }
                  }
                ],
                "max_points": 15
              }
            ]
          },
          {
            "key": "community_health",
            "band": "critical",
            "name": "Community health",
            "note": null,
            "notes": [],
            "value": 25,
            "inputs": {
              "has_readme": false,
              "has_license": false,
              "has_contributing": false,
              "has_issue_template": false,
              "has_code_of_conduct": false,
              "has_pull_request_template": false
            },
            "components": [
              {
                "key": "readme",
                "name": "README",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 22.5
              },
              {
                "key": "license",
                "name": "License",
                "detail": "recognized license (MPL-2.0)",
                "points": 22.5,
                "status": "met",
                "details": [
                  {
                    "code": "license_standard",
                    "params": {}
                  },
                  {
                    "code": "license_spdx",
                    "params": {
                      "spdx": "MPL-2.0"
                    }
                  }
                ],
                "max_points": 22.5
              },
              {
                "key": "contributing_guide",
                "name": "CONTRIBUTING guide",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 18
              },
              {
                "key": "code_of_conduct",
                "name": "Code of conduct",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 13.5
              },
              {
                "key": "issue_template",
                "name": "Issue template",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.2
              },
              {
                "key": "pr_template",
                "name": "PR template",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 6.3
              }
            ]
          }
        ],
        "description": "Does the project have users, downloads, attention, and a welcoming setup for contributors?"
      },
      {
        "key": "governance",
        "band": "moderate",
        "name": "Sustainability & Governance",
        "value": 57,
        "weight": 0.24,
        "metrics": [
          {
            "key": "maintainer_resilience",
            "band": "critical",
            "name": "Maintainer resilience (bus factor)",
            "note": null,
            "notes": [],
            "value": 25,
            "inputs": {
              "bus_factor": 1,
              "contributors_sampled": 2,
              "top_contributor_share": 0.681
            },
            "components": [
              {
                "key": "bus_factor",
                "name": "Bus factor",
                "detail": "1 contributor(s) cover half of all commits",
                "points": 9,
                "status": "partial",
                "details": [
                  {
                    "code": "bus_factor",
                    "params": {
                      "count": 1
                    }
                  }
                ],
                "max_points": 54
              },
              {
                "key": "commit_distribution",
                "name": "Commit distribution",
                "detail": "top contributor authored 68% of commits",
                "points": 7.2,
                "status": "partial",
                "details": [
                  {
                    "code": "top_contributor_share",
                    "params": {
                      "share": 68
                    }
                  }
                ],
                "max_points": 22.5
              },
              {
                "key": "contributor_breadth",
                "name": "Contributor breadth",
                "detail": "2 contributors",
                "points": 2.7,
                "status": "partial",
                "details": [
                  {
                    "code": "contributors_sampled",
                    "params": {
                      "count": 2
                    }
                  }
                ],
                "max_points": 13.5
              },
              {
                "key": "openssf_scorecard_contributors",
                "name": "OpenSSF Scorecard: Contributors",
                "detail": "project has 2 contributing companies or organizations -- score normalized to 6",
                "points": 6,
                "status": "partial",
                "details": [],
                "max_points": 10
              }
            ]
          },
          {
            "key": "responsiveness",
            "band": "good",
            "name": "Issue & PR responsiveness",
            "note": "Excluded from scoring (no data or not applicable): Issue resolution. Remaining weights renormalized.",
            "notes": [
              {
                "code": "excluded_no_data",
                "params": {
                  "components": [
                    "issue_resolution"
                  ]
                }
              },
              {
                "code": "weights_renormalized",
                "params": {}
              }
            ],
            "value": 72,
            "inputs": {
              "merged_prs": 1,
              "open_issues": 0,
              "closed_issues": 0,
              "issue_closed_ratio": null,
              "closed_unmerged_prs": 0
            },
            "components": [
              {
                "key": "issue_resolution",
                "name": "Issue resolution",
                "detail": "no issues or no data",
                "points": 0,
                "status": "excluded",
                "details": [
                  {
                    "code": "no_issues_or_data",
                    "params": {}
                  }
                ],
                "max_points": 46.75
              },
              {
                "key": "pr_acceptance",
                "name": "PR acceptance",
                "detail": "1/1 decided PRs merged",
                "points": 38.2,
                "status": "met",
                "details": [
                  {
                    "code": "decided_prs_merged",
                    "params": {
                      "merged": 1,
                      "decided": 1
                    }
                  }
                ],
                "max_points": 38.25
              },
              {
                "key": "openssf_scorecard_code_review",
                "name": "OpenSSF Scorecard: Code-Review",
                "detail": "Found 0/26 approved changesets -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 15
              }
            ]
          },
          {
            "key": "stewardship",
            "band": "at_risk",
            "name": "Ownership & stewardship",
            "note": null,
            "notes": [],
            "value": 47,
            "inputs": {
              "followers": 3,
              "owner_type": "Organization",
              "is_verified": null,
              "owner_login": "CircleCI-Research",
              "public_repos": 4,
              "account_age_days": 1387
            },
            "components": [
              {
                "key": "ownership_backing",
                "name": "Ownership backing",
                "detail": "organization-owned",
                "points": 30,
                "status": "met",
                "details": [
                  {
                    "code": "owner_organization",
                    "params": {}
                  }
                ],
                "max_points": 30
              },
              {
                "key": "verified_domain",
                "name": "Verified domain",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 20
              },
              {
                "key": "owner_reach",
                "name": "Owner reach",
                "detail": "3 followers of CircleCI-Research",
                "points": 4.3,
                "status": "partial",
                "details": [
                  {
                    "code": "owner_followers",
                    "params": {
                      "count": 3,
                      "login": "CircleCI-Research"
                    }
                  }
                ],
                "max_points": 25
              },
              {
                "key": "track_record",
                "name": "Track record",
                "detail": "4 public repos, account ~3 yr old",
                "points": 12.7,
                "status": "partial",
                "details": [
                  {
                    "code": "public_repos",
                    "params": {
                      "count": 4
                    }
                  },
                  {
                    "code": "account_age_years",
                    "params": {
                      "years": 3
                    }
                  }
                ],
                "max_points": 25
              }
            ]
          },
          {
            "key": "package_maintenance",
            "band": "excellent",
            "name": "Package maintenance",
            "note": null,
            "notes": [],
            "value": 100,
            "inputs": {
              "packages": [
                "github.com/CircleCI-Research/evalbench"
              ],
              "ecosystems": "go",
              "any_deprecated": false,
              "min_days_since_publish": 128
            },
            "components": [
              {
                "key": "published_resolvable",
                "name": "Published & resolvable",
                "detail": "1 package(s) on go",
                "points": 25,
                "status": "met",
                "details": [
                  {
                    "code": "packages_published",
                    "params": {
                      "count": 1,
                      "ecosystems": "go"
                    }
                  }
                ],
                "max_points": 25
              },
              {
                "key": "publish_recency",
                "name": "Publish recency",
                "detail": "latest publish 128 days ago",
                "points": 35,
                "status": "met",
                "details": [
                  {
                    "code": "publish_recency",
                    "params": {
                      "days": 128
                    }
                  }
                ],
                "max_points": 35
              },
              {
                "key": "version_history",
                "name": "Version history",
                "detail": "33 published versions",
                "points": 20,
                "status": "met",
                "details": [
                  {
                    "code": "published_versions",
                    "params": {
                      "count": 33
                    }
                  }
                ],
                "max_points": 20
              },
              {
                "key": "not_deprecated",
                "name": "Not deprecated",
                "detail": "active, not deprecated or yanked",
                "points": 20,
                "status": "met",
                "details": [
                  {
                    "code": "package_not_deprecated",
                    "params": {}
                  }
                ],
                "max_points": 20
              }
            ]
          }
        ],
        "description": "Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep?"
      },
      {
        "key": "engineering",
        "band": "moderate",
        "name": "Engineering Quality",
        "value": 51,
        "weight": 0.2,
        "metrics": [
          {
            "key": "engineering_practices",
            "band": "moderate",
            "name": "Engineering practices",
            "note": null,
            "notes": [],
            "value": 68,
            "inputs": {
              "has_ci": true,
              "has_tests": true,
              "has_editorconfig": false,
              "has_linter_config": false,
              "has_precommit_config": false
            },
            "components": [
              {
                "key": "ci_workflows",
                "name": "CI workflows",
                "detail": "1 workflow(s)",
                "points": 24,
                "status": "met",
                "details": [
                  {
                    "code": "ci_workflows",
                    "params": {
                      "count": 1
                    }
                  }
                ],
                "max_points": 24
              },
              {
                "key": "tests_present",
                "name": "Tests present",
                "detail": null,
                "points": 24,
                "status": "met",
                "details": [],
                "max_points": 24
              },
              {
                "key": "linter_config",
                "name": "Linter config",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 16
              },
              {
                "key": "pre_commit_hooks",
                "name": "Pre-commit hooks",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 9.6
              },
              {
                "key": "editorconfig",
                "name": ".editorconfig",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 6.4
              },
              {
                "key": "openssf_scorecard_ci_tests",
                "name": "OpenSSF Scorecard: CI-Tests",
                "detail": "1 out of 1 merged PRs checked by a CI test -- score normalized to 10",
                "points": 20,
                "status": "met",
                "details": [],
                "max_points": 20
              }
            ]
          },
          {
            "key": "documentation",
            "band": "critical",
            "name": "Documentation",
            "note": null,
            "notes": [],
            "value": 25,
            "inputs": {
              "topics": [
                "ai",
                "ai-agents",
                "benchmark-framework",
                "ci-cd",
                "circleci",
                "devops",
                "evaluation-framework",
                "evaluation-metrics",
                "llm",
                "model-comparison"
              ],
              "has_wiki": false,
              "homepage": "https://loop.circleci.com",
              "has_readme": false,
              "has_docs_dir": false,
              "has_description": false
            },
            "components": [
              {
                "key": "readme",
                "name": "README",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 30
              },
              {
                "key": "documentation_directory",
                "name": "Documentation directory",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 25
              },
              {
                "key": "documentation_homepage_site",
                "name": "Documentation / homepage site",
                "detail": "https://loop.circleci.com",
                "points": 15,
                "status": "met",
                "details": [],
                "max_points": 15
              },
              {
                "key": "repository_description",
                "name": "Repository description",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 10
              },
              {
                "key": "topics",
                "name": "Topics",
                "detail": "10 topics",
                "points": 10,
                "status": "met",
                "details": [
                  {
                    "code": "topics_count",
                    "params": {
                      "count": 10
                    }
                  }
                ],
                "max_points": 10
              },
              {
                "key": "wiki",
                "name": "Wiki",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 10
              }
            ]
          }
        ],
        "description": "Are baseline engineering and documentation practices in place?"
      },
      {
        "key": "security",
        "band": "at_risk",
        "name": "Security",
        "value": 30,
        "weight": 0.16,
        "metrics": [
          {
            "key": "security_posture",
            "band": "at_risk",
            "name": "Security posture",
            "note": "Excluded from scoring (no data or not applicable): Packaging, Signed-Releases. Remaining weights renormalized.",
            "notes": [
              {
                "code": "excluded_no_data",
                "params": {
                  "components": [
                    "packaging",
                    "signed_releases"
                  ]
                }
              },
              {
                "code": "weights_renormalized",
                "params": {}
              }
            ],
            "value": 30,
            "inputs": {
              "source": "openssf_scorecard",
              "checks_evaluated": 16,
              "scorecard_version": "v5.5.0",
              "checks_inconclusive": 2,
              "scorecard_aggregate": 3
            },
            "components": [
              {
                "key": "binary_artifacts",
                "name": "Binary-Artifacts",
                "detail": "no binaries found in the repo",
                "points": 7.5,
                "status": "met",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "branch_protection",
                "name": "Branch-Protection",
                "detail": "branch protection is not maximal on development and all release branches",
                "points": 3.8,
                "status": "partial",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "ci_tests",
                "name": "CI-Tests",
                "detail": "1 out of 1 merged PRs checked by a CI test -- score normalized to 10",
                "points": 2.5,
                "status": "met",
                "details": [],
                "max_points": 2.5
              },
              {
                "key": "cii_best_practices",
                "name": "CII-Best-Practices",
                "detail": "no effort to earn an OpenSSF best practices badge detected",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 2.5
              },
              {
                "key": "code_review",
                "name": "Code-Review",
                "detail": "Found 0/26 approved changesets -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "contributors",
                "name": "Contributors",
                "detail": "project has 2 contributing companies or organizations -- score normalized to 6",
                "points": 1.5,
                "status": "partial",
                "details": [],
                "max_points": 2.5
              },
              {
                "key": "dangerous_workflow",
                "name": "Dangerous-Workflow",
                "detail": "no dangerous workflow patterns detected",
                "points": 10,
                "status": "met",
                "details": [],
                "max_points": 10
              },
              {
                "key": "dependency_update_tool",
                "name": "Dependency-Update-Tool",
                "detail": "no update tool detected",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "fuzzing",
                "name": "Fuzzing",
                "detail": "project is not fuzzed",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 5
              },
              {
                "key": "license",
                "name": "License",
                "detail": "license file detected",
                "points": 2.5,
                "status": "met",
                "details": [],
                "max_points": 2.5
              },
              {
                "key": "maintained",
                "name": "Maintained",
                "detail": "0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "packaging",
                "name": "Packaging",
                "detail": "packaging workflow not detected",
                "points": 0,
                "status": "excluded",
                "details": [
                  {
                    "code": "no_data",
                    "params": {}
                  }
                ],
                "max_points": 5
              },
              {
                "key": "pinned_dependencies",
                "name": "Pinned-Dependencies",
                "detail": "dependency not pinned by hash detected -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 5
              },
              {
                "key": "sast",
                "name": "SAST",
                "detail": "SAST tool is not run on all commits -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 5
              },
              {
                "key": "security_policy",
                "name": "Security-Policy",
                "detail": "security policy file not detected",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 5
              },
              {
                "key": "signed_releases",
                "name": "Signed-Releases",
                "detail": "no releases found",
                "points": 0,
                "status": "excluded",
                "details": [
                  {
                    "code": "no_data",
                    "params": {}
                  }
                ],
                "max_points": 7.5
              },
              {
                "key": "token_permissions",
                "name": "Token-Permissions",
                "detail": "detected GitHub workflow tokens with excessive permissions",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "vulnerabilities",
                "name": "Vulnerabilities",
                "detail": "44 existing vulnerabilities detected",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              }
            ]
          },
          {
            "key": "high_risk_jurisdiction_exposure",
            "band": "excellent",
            "name": "High-Risk Jurisdiction Exposure",
            "note": "Only high-confidence self-published location evidence affects this multiplier. Ambiguous matches are review-only; country evidence is not proof of nationality, citizenship, legal registration, malicious intent, or sanctions status.",
            "notes": [
              {
                "code": "jurisdiction_evidence_limits",
                "params": {}
              }
            ],
            "value": 100,
            "inputs": {
              "meaning": "self-published location evidence; not nationality or citizenship",
              "red_flag": false,
              "exposures": [],
              "policy_countries": [
                "Russia",
                "Iran",
                "North Korea"
              ],
              "review_only_matches": 0,
              "assessed_self_published_locations": 4
            },
            "components": [
              {
                "key": "policy_exposure_multiplier",
                "name": "Policy exposure multiplier",
                "detail": "no confirmed policy-scope location match",
                "points": 100,
                "status": "met",
                "details": [
                  {
                    "code": "jurisdiction_no_match",
                    "params": {}
                  }
                ],
                "max_points": 100
              }
            ]
          }
        ],
        "description": "Are visible security and supply-chain practices strong, with no malicious dependency and no unresolved high-risk jurisdiction exposure?"
      },
      {
        "key": "ai_readiness",
        "band": "moderate",
        "name": "AI Readiness",
        "value": 65,
        "weight": 0,
        "metrics": [
          {
            "key": "ai_agent_context",
            "band": "excellent",
            "name": "Agent context & guidance",
            "note": null,
            "notes": [],
            "value": 85,
            "inputs": {
              "has_llms_txt": false,
              "legible_history_share": 0.96,
              "agent_instruction_files": [
                "AGENTS.md"
              ],
              "agent_instruction_max_bytes": 3085
            },
            "components": [
              {
                "key": "agent_instructions",
                "name": "Agent instructions",
                "detail": "AGENTS.md",
                "points": 45,
                "status": "met",
                "details": [
                  {
                    "code": "file_list",
                    "params": {
                      "files": "AGENTS.md"
                    }
                  }
                ],
                "max_points": 45
              },
              {
                "key": "machine_readable_docs_llms_txt",
                "name": "Machine-readable docs (llms.txt)",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 15
              },
              {
                "key": "legible_commit_history",
                "name": "Legible commit history",
                "detail": "96 of 100 human commits state their intent (structured subject or explanatory body)",
                "points": 40,
                "status": "met",
                "details": [
                  {
                    "code": "legible_history",
                    "params": {
                      "legible": 96,
                      "sampled": 100
                    }
                  }
                ],
                "max_points": 40
              }
            ]
          },
          {
            "key": "ai_verify_loop",
            "band": "moderate",
            "name": "Verify loop (build / test / typecheck)",
            "note": null,
            "notes": [],
            "value": 55,
            "inputs": {
              "has_nix": false,
              "has_tests": true,
              "lockfiles": [
                "go.sum"
              ],
              "has_dockerfile": false,
              "typed_language": false,
              "bootstrap_files": [],
              "has_devcontainer": false,
              "has_linter_config": false,
              "typecheck_configs": [],
              "agent_commit_share": 0.11,
              "toolchain_manifests": [
                "go.mod"
              ],
              "dependency_bot_commit_share": 0
            },
            "components": [
              {
                "key": "one_command_bootstrap",
                "name": "One-command bootstrap",
                "detail": "go.mod (toolchain convention, no task runner)",
                "points": 12.6,
                "status": "partial",
                "details": [
                  {
                    "code": "toolchain_convention",
                    "params": {
                      "files": "go.mod"
                    }
                  }
                ],
                "max_points": 18
              },
              {
                "key": "automated_tests",
                "name": "Automated tests",
                "detail": null,
                "points": 22,
                "status": "met",
                "details": [],
                "max_points": 22
              },
              {
                "key": "lint_format_config",
                "name": "Lint / format config",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 11
              },
              {
                "key": "static_type_checking",
                "name": "Static type checking",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 11
              },
              {
                "key": "reproducible_environment",
                "name": "Reproducible environment",
                "detail": "lockfile",
                "points": 10,
                "status": "met",
                "details": [
                  {
                    "code": "file_list",
                    "params": {
                      "files": "lockfile"
                    }
                  }
                ],
                "max_points": 10
              },
              {
                "key": "demonstrated_agent_practice",
                "name": "Demonstrated agent practice",
                "detail": "11 of the last 100 commits agent-authored or agent-credited",
                "points": 10,
                "status": "met",
                "details": [
                  {
                    "code": "agent_authored_commits",
                    "params": {
                      "count": 11,
                      "sampled": 100
                    }
                  }
                ],
                "max_points": 10
              },
              {
                "key": "automated_maintenance",
                "name": "Automated maintenance",
                "detail": "no automated dependency updates observed",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "no_dependency_automation",
                    "params": {}
                  }
                ],
                "max_points": 8
              },
              {
                "key": "openssf_scorecard_pinned_dependencies",
                "name": "OpenSSF Scorecard: Pinned-Dependencies",
                "detail": "dependency not pinned by hash detected -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 10
              }
            ]
          },
          {
            "key": "ai_code_legibility",
            "band": "moderate",
            "name": "Code legibility for models",
            "note": null,
            "notes": [],
            "value": 54,
            "inputs": {
              "primary_language": "HTML",
              "largest_source_bytes": 73199,
              "source_files_sampled": 442,
              "oversized_source_files": 5
            },
            "components": [
              {
                "key": "type_checkable_code",
                "name": "Type-checkable code",
                "detail": "HTML without a type-check config",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "no_typecheck_config_language",
                    "params": {
                      "language": "HTML"
                    }
                  }
                ],
                "max_points": 45
              },
              {
                "key": "manageable_file_sizes",
                "name": "Manageable file sizes",
                "detail": "5/442 source files over 60KB",
                "points": 54.4,
                "status": "partial",
                "details": [
                  {
                    "code": "oversized_source_files",
                    "params": {
                      "kb": 60,
                      "sampled": 442,
                      "oversized": 5
                    }
                  }
                ],
                "max_points": 55
              }
            ]
          }
        ],
        "description": "How well is the repo equipped to be developed and maintained with AI coding agents? An independent, experimental badge — weight 0.0, so it is surfaced on its own and does not affect the overall health score."
      }
    ],
    "metrics_version": "1.13.0"
  },
  "warnings": [
    "Community profile unavailable",
    "GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository"
  ],
  "report_type": "repository",
  "generated_at": "2026-07-25T08:58:13.715545Z",
  "schema_version": "0.27.0",
  "badge_url": "https://raw.githubusercontent.com/inspect-software/badges/main/v1/c/CircleCI-Research/evalbench.svg",
  "full_name": "CircleCI-Research/evalbench",
  "license_state": "standard",
  "license_spdx": "MPL-2.0"
}

Bewertungen sind Signale, keine Garantien. Sie spiegeln öffentlich sichtbare Praxis auf GitHub wider — kein Code-Audit und keine Sicherheitsgarantie.

Fehlende Daten werden ausgeschlossen und die Gewichte neu normiert, nie als null bewertet. Die Methodik ist versioniert und offen: Metriken v1.13.0, Schema v0.27.0 — vollständige Methodik · Metriken-Wiki.

Wie ein einzelnes Ergebnis im Gesamtregister steht: aggregierte StatistikenGo.