公开记录
软件健康报告模式 0.27.0 · 指标 1.13.0 · 2026-07-25 08:58 UTC

CircleCI-Research / evalbench

Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.

HTMLMPL-2.0★ 0 星标⑂ 0 复刻始于 2026年3月复刻在 GitHub 上查看 ↗

CircleCI-Research/evalbench 的健康指数为 100 分中的 43 分,处于「存在风险」区间。 其得分最高的类别是AI Readiness(65/100),最低的是Community & Adoption(12/100)。 最近一次更新在 127 天前。 近期的大部分工作由 1 位贡献者完成。

43
总分 / 100
存在风险

软件健康指数

指标归入加权类别,统一采用 1–100 量表。总体分先取类别加权平均;当公开证据触发高风险司法辖区政策时,评级会按政策调整,并设置 49(有风险)的上限。AI 就绪度不计入总体分。

43
优秀85-100堪称典范;基本满足所有检验标准
良好70-84健康;仅有轻微不足
中等50-69可接受,但存在明显不足;建议进行审查
存在风险30-49存在重大薄弱环节;采用时应保持审慎
危急1-29问题严重(项目被弃置、仅有单一维护者、缺乏基本工程规范)
活力社区与采用可持续性与治理工程质量安全AI 就绪度

评分画像

每条轴代表一个类别。形状比平均值更重要——健康的对象会填满整个图形,而“一峰一谷”式画像意味着某一维度的优势正掩盖另一维度的风险。

所有权

3 关注者4 个公开仓库始于 2022年10月

该仓库由组织支持——共同承担、可问责的托管责任,可延续于任何单一维护者之后。

软件包生态系统

注册表软件包版本月下载量版本数最近发布
Gogithub.com/CircleCI-Research/evalbenchv0.17.0-33128 天前

按类别列示的指标

活力

项目是否仍有生命——是否仍在编写代码,是否仍在发布版本?

56中等 · 占总体的 22%

开发活跃度

42存在风险
评分方式
9.9/36推送新近度 — 最近一次推送于 127 天前
14.5/36提交节奏 — 52 周中有 21 周有提交
17.4/18提交量 — 最近一年 86 次提交
0/10OpenSSF Scorecard:Maintained — 0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0
所用输入
commits_last_year86
human_commit_share1
days_since_last_push127
active_weeks_last_year21

发布纪律

78良好
评分方式
16.2/27有发布版本 — 33 个版本标签(无 GitHub 发布版本)
27/36发布时效 — 最近一次发布版本于 128 天前
27/27发布节奏 — 约每 13.2 天发布一次
0/10OpenSSF Scorecard:Signed-Releases — 无数据
所用输入
releases_count33
latest_release_tagv0.17.0
releases_from_tags
days_since_latest_release128
mean_days_between_releases13.2
已排除计分(无数据或不适用):OpenSSF Scorecard:Signed-Releases。 其余权重已重新归一化。

社区与采用

项目是否拥有用户、下载量与关注度,并具备欢迎贡献者参与的配置?

12危急 · 占总体的 18%
评分方式
0/60星标 — 0 个星标
0/25复刻 — 0 个复刻
0/15关注者 — 0 位关注者
所用输入
forks0
stars0
watchers0
growth_stateunverified
growth_factor_pct100
growth_unverified_reasonno_history

社区健康

25危急
评分方式
0/22.5README
22.5/22.5许可证 — 可识别的许可证(MPL-2.0)
0/18CONTRIBUTING 指南
0/13.5行为准则
0/7.2议题模板
0/6.3PR 模板
所用输入
has_readme
has_license
has_contributing
has_issue_template
has_code_of_conduct
has_pull_request_template

可持续性与治理

项目能否在其成员之外延续——巴士系数、响应能力、由谁支持,以及软件包的维护状况?

57中等 · 占总体的 24%
评分方式
9/54巴士系数 — 1 位贡献者贡献了半数提交
7.2/22.5提交分布 — 头号贡献者编写了 68% 的提交
2.7/13.5贡献者广度 — 2 位贡献者
6/10OpenSSF Scorecard:Contributors — project has 2 contributing companies or organizations -- score normalized to 6
所用输入
bus_factor1
contributors_sampled2
top_contributor_share0.681
评分方式
0/46.8议题解决 — 没有议题或无数据
38.2/38.3PR 接受 — 已裁定的 PR 中 1/1 已合并
0/15OpenSSF Scorecard:Code-Review — Found 0/26 approved changesets -- score normalized to 0
所用输入
merged_prs1
open_issues0
closed_issues0
issue_closed_ratio
closed_unmerged_prs0
已排除计分(无数据或不适用):议题解决。 其余权重已重新归一化。
评分方式
30/30所有权背书 — 组织持有
0/20已验证域名
4.3/25所有者影响力 — CircleCI-Research 有 3 位关注者
12.7/25既往记录 — 4 个公开仓库,账户约 3 年
所用输入
followers3
owner_typeOrganization
is_verified
owner_loginCircleCI-Research
public_repos4
account_age_days1,387
评分方式
25/25已发布且可解析 — go 上有 1 个软件包
35/35发布时效 — 最近一次发布于 128 天前
20/20版本历史 — 33 个已发布版本
20/20未被弃用 — 活跃,未被弃用或撤回
所用输入
packagesgithub.com/CircleCI-Research/evalbench
ecosystemsgo
any_deprecated
min_days_since_publish128

工程质量

基础的工程与文档实践是否到位?

51中等 · 占总体的 20%

工程实践

68中等
评分方式
24/24CI 工作流 — 1 个工作流
24/24存在测试
0/16Linter 配置
0/9.6Pre-commit 钩子
0/6.4.editorconfig
20/20OpenSSF Scorecard:CI-Tests — 1 out of 1 merged PRs checked by a CI test -- score normalized to 10
所用输入
has_ci
has_tests
has_editorconfig
has_linter_config
has_precommit_config

文档

25危急
评分方式
0/30README
0/25文档目录
15/15文档 / 主页站点 — https://loop.circleci.com
0/10仓库描述
10/10主题标签 — 10 个主题标签
0/10Wiki
所用输入
topicsai, ai-agents, benchmark-framework, ci-cd, circleci, devops, evaluation-framework, evaluation-metrics, llm, model-comparison
has_wiki
homepagehttps://loop.circleci.com
has_readme
has_docs_dir
has_description

安全

可见的安全与供应链实践是否稳固,且不存在未解决的高风险司法辖区暴露?

30存在风险 · 占总体的 16%

安全态势

30存在风险
评分方式
7.5/7.5Binary-Artifacts — no binaries found in the repo
3.8/7.5Branch-Protection — branch protection is not maximal on development and all release branches
2.5/2.5CI-Tests — 1 out of 1 merged PRs checked by a CI test -- score normalized to 10
0/2.5CII-Best-Practices — no effort to earn an OpenSSF best practices badge detected
0/7.5Code-Review — Found 0/26 approved changesets -- score normalized to 0
1.5/2.5Contributors — project has 2 contributing companies or organizations -- score normalized to 6
10/10Dangerous-Workflow — no dangerous workflow patterns detected
0/7.5Dependency-Update-Tool — no update tool detected
0/5Fuzzing — project is not fuzzed
2.5/2.5许可证 — license file detected
0/7.5Maintained — 0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0
0/5Packaging — 无数据
0/5Pinned-Dependencies — dependency not pinned by hash detected -- score normalized to 0
0/5SAST — SAST tool is not run on all commits -- score normalized to 0
0/5Security-Policy — security policy file not detected
0/7.5Signed-Releases — 无数据
0/7.5Token-Permissions — detected GitHub workflow tokens with excessive permissions
0/7.5Vulnerabilities — 44 existing vulnerabilities detected
所用输入
sourceopenssf_scorecard
checks_evaluated16
scorecard_versionv5.5.0
checks_inconclusive2
scorecard_aggregate3
已排除计分(无数据或不适用):packaging, signed_releases。 其余权重已重新归一化。

AI 就绪度

该仓库在多大程度上具备与 AI 编码代理协同开发与维护的条件?这是一枚独立的实验性徽章——权重为 0.0,因此单独呈现,不影响总体健康评分。

65中等 · 占总体的 0%
评分方式
45/45代理指令 — AGENTS.md
0/15机器可读文档(llms.txt)
40/40可读的提交历史 — 100 次人类提交中有 96 次说明了意图(结构化标题或解释性正文)
所用输入
has_llms_txt
legible_history_share0.96
agent_instruction_filesAGENTS.md
agent_instruction_max_bytes3,085
评分方式
12.6/18一条命令的引导启动 — go.mod(工具链约定,无任务运行器)
22/22自动化测试
0/11Lint / 格式化配置
0/11静态类型检查
10/10可复现环境 — lockfile
10/10已体现的代理实践 — 最近 100 次提交中有 11 次由代理编写或署名代理
0/8自动化维护 — 未观察到自动依赖更新
0/10OpenSSF Scorecard:Pinned-Dependencies — dependency not pinned by hash detected -- score normalized to 0
所用输入
has_nix
has_tests
lockfilesgo.sum
has_dockerfile
typed_language
bootstrap_files
has_devcontainer
has_linter_config
typecheck_configs
agent_commit_share0.11
toolchain_manifestsgo.mod
dependency_bot_commit_share0
评分方式
0/45可类型检查的代码 — HTML,未配置类型检查
54.4/55可控的文件大小 — 采样的 442 个源文件中有 5 个超过 60KB
所用输入
primary_languageHTML
largest_source_bytes73,199
source_files_sampled442
oversized_source_files5

关键数据

0GitHub 星标
2贡献者
86最近 12 个月提交数
127距最近推送天数
33发布版本数
1巴士系数(bus factor)
0开放议题
Go软件包生态系统数

数据采集警告

  • Community profile unavailable
  • GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository

更多细节

OpenSSF Scorecard 3.0 / 10
3.0综合

来自开源项目 OpenSSF Scorecard 的独立、工具无关的安全评估。每项检查奖励的是安全实践本身,而非特定供应商的工具。Scorecard 无法判定的检查项标记为 不适用,并从安全评分中剔除(绝不按零分计)。Scorecard v5.5.0 · 2026-07-25 08:58 UTC

10Binary-Artifactsno binaries found in the repo
5Branch-Protectionbranch protection is not maximal on development and all release branches
10CI-Tests1 out of 1 merged PRs checked by a CI test -- score normalized to 10
0CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
0Code-ReviewFound 0/26 approved changesets -- score normalized to 0
6Contributorsproject has 2 contributing companies or organizations -- score normalized to 6
10Dangerous-Workflowno dangerous workflow patterns detected
0Dependency-Update-Toolno update tool detected
0Fuzzingproject is not fuzzed
10Licenselicense file detected
0Maintained0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0
不适用Packagingpackaging workflow not detected
0Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
0SASTSAST tool is not run on all commits -- score normalized to 0
0Security-Policysecurity policy file not detected
不适用Signed-Releasesno releases found
0Token-Permissionsdetected GitHub workflow tokens with excessive permissions
0Vulnerabilities44 existing vulnerabilities detected
直接依赖 24
注册表软件包版本约束清单文件
Gogithub.com/charmbracelet/x/termv0.2.1go.mod
Gogithub.com/containerd/errdefsv1.0.0go.mod
Gogithub.com/docker/dockerv28.5.2+incompatiblego.mod
Gogithub.com/google/uuidv1.6.0go.mod
Gogithub.com/invopop/jsonschemav0.13.0go.mod
Gogithub.com/kaptinlin/jsonrepairv0.2.8go.mod
Gogithub.com/oklog/ulid/v2v2.1.1go.mod
Gogithub.com/openai/openai-go/v3v3.26.0go.mod
Gogithub.com/santhosh-tekuri/jsonschema/v6v6.0.2go.mod
Gogithub.com/sethvargo/go-retryv0.3.0go.mod
Gogithub.com/stretchr/testifyv1.11.1go.mod
Gogolang.org/x/timev0.14.0go.mod
Gogoogle.golang.org/genaiv1.48.0go.mod
Gogopkg.in/validator.v2v2.0.1go.mod
Gogopkg.in/yaml.v3v3.0.1go.mod
Gogithub.com/anthropics/anthropic-sdk-gov1.26.0go.mod
Gogithub.com/charmbracelet/bubbles/v2v2.0.0-beta.1go.mod
Gogithub.com/charmbracelet/bubbletea/v2v2.0.0-beta.1go.mod
Gogithub.com/charmbracelet/lipgloss/v2v2.0.0-beta.1go.mod
Gogithub.com/cohesion-org/deepseek-gov1.3.3go.mod
Gogithub.com/go-playground/validator/v10v10.30.1go.mod
Gogithub.com/rs/zerologv1.34.0go.mod
Gogithub.com/sergi/go-diffv1.4.0go.mod
Gogolang.org/x/expv0.0.0-20260218203240-3dfff04db8fago.mod
全部依赖 未采集

本报告未能采集到解析后的依赖集合:GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository

原始 JSON 报告 机器可读
{
  "data": {
    "repo": {
      "topics": [
        "ai",
        "ai-agents",
        "benchmark-framework",
        "ci-cd",
        "circleci",
        "devops",
        "evaluation-framework",
        "evaluation-metrics",
        "llm",
        "model-comparison"
      ],
      "is_fork": true,
      "size_kb": 7617,
      "has_wiki": false,
      "homepage": "https://loop.circleci.com",
      "languages": {
        "Go": 876390,
        "HTML": 20005421,
        "Shell": 1009,
        "Python": 5074,
        "Go Template": 78025
      },
      "pushed_at": "2026-03-19T15:10:09Z",
      "created_at": "2026-03-19T14:20:11Z",
      "owner_type": "Organization",
      "updated_at": "2026-03-19T16:25:31Z",
      "description": "Evaluate LLMs side-by-side. Benchmark AI models and coding agents across providers like OpenAI, Google, Anthropic, DeepSeek, and more. Supports custom tasks, structured JSON responses, tool use, and LLM-as-judge validation. Originally created by Petr Malik as MindTrial.",
      "is_archived": false,
      "is_disabled": false,
      "license_spdx": "MPL-2.0",
      "default_branch": "main",
      "license_spdx_raw": "MPL-2.0",
      "primary_language": "HTML",
      "significant_languages": [
        "HTML"
      ]
    },
    "owner": {
      "blog": "https://circleci.com",
      "name": "CircleCI Research",
      "type": "Organization",
      "login": "CircleCI-Research",
      "company": null,
      "location": "United States of America",
      "followers": 3,
      "avatar_url": "https://avatars.githubusercontent.com/u/115158100?v=4",
      "created_at": "2022-10-06T12:04:53Z",
      "is_verified": null,
      "public_repos": 4,
      "account_age_days": 1387
    },
    "license": {
      "state": "standard",
      "spdx_id": "MPL-2.0",
      "raw_spdx": "MPL-2.0",
      "file_present": true,
      "scorecard_found": true,
      "profile_has_license": false
    },
    "activity": {
      "releases": [
        {
          "tag": "v0.17.0",
          "kind": "minor",
          "published_at": "2026-03-18T22:43:28Z"
        },
        {
          "tag": "v0.16.0",
          "kind": "minor",
          "published_at": "2026-03-12T19:58:14Z"
        },
        {
          "tag": "v0.15.0",
          "kind": "minor",
          "published_at": "2026-02-27T22:17:14Z"
        },
        {
          "tag": "v0.14.1",
          "kind": "patch",
          "published_at": "2026-02-07T15:46:27Z"
        },
        {
          "tag": "v0.14.0",
          "kind": "minor",
          "published_at": "2026-02-01T03:21:58Z"
        },
        {
          "tag": "v0.13.4",
          "kind": "patch",
          "published_at": "2026-01-17T19:22:30Z"
        },
        {
          "tag": "v0.13.3",
          "kind": "patch",
          "published_at": "2025-12-23T19:35:43Z"
        },
        {
          "tag": "v0.13.2",
          "kind": "patch",
          "published_at": "2025-12-12T18:17:51Z"
        },
        {
          "tag": "v0.13.1",
          "kind": "patch",
          "published_at": "2025-12-08T00:26:37Z"
        },
        {
          "tag": "v0.13.0",
          "kind": "minor",
          "published_at": "2025-11-20T05:56:34Z"
        },
        {
          "tag": "v0.12.2",
          "kind": "patch",
          "published_at": "2025-11-12T21:39:53Z"
        },
        {
          "tag": "v0.12.1",
          "kind": "patch",
          "published_at": "2025-10-11T01:09:29Z"
        },
        {
          "tag": "v0.12.0",
          "kind": "minor",
          "published_at": "2025-10-09T18:05:48Z"
        },
        {
          "tag": "v0.11.1",
          "kind": "patch",
          "published_at": "2025-10-03T00:05:30Z"
        },
        {
          "tag": "v0.11.0",
          "kind": "minor",
          "published_at": "2025-10-02T15:35:36Z"
        },
        {
          "tag": "v0.10.1",
          "kind": "patch",
          "published_at": "2025-09-23T21:16:27Z"
        },
        {
          "tag": "v0.10.0",
          "kind": "minor",
          "published_at": "2025-09-19T16:11:27Z"
        },
        {
          "tag": "v0.9.0",
          "kind": "minor",
          "published_at": "2025-09-16T18:47:58Z"
        },
        {
          "tag": "v0.8.0",
          "kind": "minor",
          "published_at": "2025-09-03T14:00:21Z"
        },
        {
          "tag": "v0.7.2",
          "kind": "patch",
          "published_at": "2025-08-28T14:57:25Z"
        },
        {
          "tag": "v0.7.1",
          "kind": "patch",
          "published_at": "2025-08-26T16:25:26Z"
        },
        {
          "tag": "v0.7.0",
          "kind": "minor",
          "published_at": "2025-08-21T20:16:13Z"
        },
        {
          "tag": "v0.6.1",
          "kind": "patch",
          "published_at": "2025-08-16T17:01:31Z"
        },
        {
          "tag": "v0.6.0",
          "kind": "minor",
          "published_at": "2025-08-11T18:11:56Z"
        },
        {
          "tag": "v0.5.0",
          "kind": "minor",
          "published_at": "2025-07-30T03:28:50Z"
        },
        {
          "tag": "v0.4.2",
          "kind": "patch",
          "published_at": "2025-07-08T19:46:19Z"
        },
        {
          "tag": "v0.4.1",
          "kind": "patch",
          "published_at": "2025-07-07T17:30:18Z"
        },
        {
          "tag": "v0.3.2",
          "kind": "patch",
          "published_at": "2025-06-27T16:30:55Z"
        },
        {
          "tag": "v0.3.1",
          "kind": "patch",
          "published_at": "2025-06-18T04:38:04Z"
        },
        {
          "tag": "v0.3.0",
          "kind": "minor",
          "published_at": "2025-05-31T13:41:35Z"
        },
        {
          "tag": "v0.2.1",
          "kind": "patch",
          "published_at": "2025-05-24T02:21:21Z"
        },
        {
          "tag": "v0.2.0",
          "kind": "minor",
          "published_at": "2025-05-19T18:49:36Z"
        },
        {
          "tag": "v0.1.0",
          "kind": "minor",
          "published_at": "2025-04-26T18:50:00Z"
        }
      ],
      "recent_commits": [
        {
          "oid": "64b9f020db13d4629897badb64f94f85a8003544",
          "body": "fix: restore GHA workflow, fix failing tests, add AGENTS.md",
          "is_bot": false,
          "headline": "Merge pull request #1 from CircleCI-Research/fix-gha",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T15:10:03Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "d0db7a898f6f628cb467accc7850547be5c042a4",
          "body": "…sitive\n\nReverts the sequential run change — runs within a provider intentionally\nexecute in parallel for throughput. Updates the doc comment to match.\n\nFixes TestRunnerRun by sorting results by (Run, Task, Got, Want) before\ncomparing, since parallel execution produces non-deterministic ordering.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "fix: restore parallel run execution and make runner tests order-insen…",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T15:06:08Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "9a263d0487b9867f614754612e963e7044385cbe",
          "body": "…golden files\n\nRuns within a single provider were incorrectly launched in parallel\ngoroutines, causing non-deterministic result ordering that violated the\ndocumented contract (\"individual runs on a single provider are executed\nsequentially\") and broke TestRunnerRun assertions.\n\nAlso updates log formatter golden files to include the Score column\nadded to the log output format.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "fix: make runs sequential within a provider and update log formatter …",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T15:01:09Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "625baa335648b4da4d8c41ccae05e3fe2ce5a214",
          "body": "Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "docs: add AGENTS.md with repo guidance and remote enforcement",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T14:54:31Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "ef7744e213fd6ed148d7378b3cf47cd6032ef80e",
          "body": "The workflow was removed during the CircleCI migration but the README\nbadge still references it, causing a broken badge on the repo.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "fix: restore GitHub Actions workflow for build badge",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T14:44:15Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "001659b4955de4026e29d49a8077a65059c19d54",
          "body": null,
          "is_bot": false,
          "headline": "refactor: tighten up README for ease of use",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T14:30:07Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "1dd80d60d0afea570bcdf8a4f1c7ad486b06bc68",
          "body": "- Rename binary from mindtrial to evalbench (cmd/mindtrial -> cmd/evalbench)\n- Migrate Go module path from github.com/petmal/mindtrial to\n  github.com/CircleCI-Research/evalbench\n- Replace MindTrial product name with EvalBench throughout codebase\n- Update README badges, links, and install commands for CircleCI-Research\n- Add attribution: \"Originally created by Petr Malik as MindTrial\"\n- Preserve MPL 2.0 license and all original copyright headers\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "chore: rebrand MindTrial to EvalBench and migrate to CircleCI-Research",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-19T14:08:38Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "ea8a6456e0884da66f75d40e99839f35e371a5e3",
          "body": "feat: Live voice race announcer for model comparisons 🎙️🤠",
          "is_bot": false,
          "headline": "Merge pull request #2 from ryan-circleci/add-live-voice-announcer",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T15:37:16Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "cd2b9aee8d945d9b4c76951cfa3b2478d605af0c",
          "body": "Made-with: Cursor",
          "is_bot": false,
          "headline": "fix: remove 30s announcer interval option (cron minimum is 1m)",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T15:35:52Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "4a37c3fbd8815c57aff691e4fa42577a73f76a30",
          "body": "…uto-loop note\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "docs: update announcer usage section with with-announcer syntax and a…",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T09:23:01Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "3fe62cea75d6750962b83dc1877fe8c4c1ce91c9",
          "body": "- Ask user for announcer check-in frequency when with-announcer is set but no interval specified\n- Support numeric arg immediately after `with-announcer` to set interval directly\n- Fix arg parsing so config file detection uses .yaml suffix instead of positional order\n- Wire ANNOUNCER_INTERVAL through to cron schedule and loop interval file\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: add interactive announcer interval prompt and flexible arg parsing",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T09:18:08Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "8738c92a52a7be5361cb7599c96ff3d663839142",
          "body": null,
          "is_bot": false,
          "headline": "Create open-dashboard.sh",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T09:05:44Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b4d9ed5b7afd9549e45b55ecee3ef3f8a887b271",
          "body": "…r flag\n\nBoth /run-model-comparison and /simulate-model-comparison now default to\nsilent mode and ask the user upfront whether they want live voice commentary.\nPass with-announcer as an argument to skip the question. Silent mode suggests\n/loop Nm /announce-model-comparison as an easy on-ramp.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: make announcer opt-in with interactive prompt and with-announce…",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T08:59:55Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "957b0b89a47d1251c3361453feb6bd5c1b665e86",
          "body": null,
          "is_bot": false,
          "headline": "Update README.md",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T07:11:02Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9e80a0530bcfc29722884363dc9d73e6b58b8ee5",
          "body": "Background agents lose Bash permissions mid-run; /loop runs in the main\nsession where permissions are already granted, so it's the only reliable\npath. Removes the Task-based auto-launch and the Step 0b mode picker.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "refactor: drop with-announcer Task mode, always use /loop for announcer",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T06:53:06Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "530854d5ab37fd4be4c05f32444a16f7e5bc4fdf",
          "body": null,
          "is_bot": false,
          "headline": "feat: add loop interval-specific color & realism to announcer commentary",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T06:25:57Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "740b9acc71da17e605ee356c8127b9de822319a0",
          "body": "- Use YYYY-MM-DD/HH-MM-SS ISO format everywhere (simulation + eval runs)\n- Unify eval results under results/eval/ prefix across slash commands and config YAMLs\n- Pass -output-dir and -output-basename to mindtrial so CSV/HTML land alongside eval.log\n- Rename /tmp/.race_* temp files to /tmp/.eval_* for consistency\n- Add interactive config picker to run-model-comparison\n- Ignore results/ and .claude/scheduled_tasks.lock in .gitignore\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: standardize results directory naming and ignore runtime artifacts",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T04:29:16Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "f067409a54ed9fe10d6aacecfc0bf4138f8309e6",
          "body": "Accidentally changed python3 to bash when refactoring the log path.\nThe .sh file is actually a Python script.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "fix: restore python3 invocation for simulate script",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T03:42:38Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "f3adbbc16f07a7f917d465128f6ead99e93498ef",
          "body": "Store eval.log at $RESULTS_DIR/eval.log instead of logs/eval.log so\neach run has a self-contained results folder. Path is written to\n/tmp/.race_log_file at init time and read by simulate, run, announce,\nand stop commands via a logs/eval.log fallback.\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: co-locate eval.log inside results directory",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T03:40:07Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "e0bb5c1a78ba31f73ad85c9c90f956419b7dd02f",
          "body": "- Add model/provider count to opening announcement in run-model-comparison\n- Change results folder date format from YYYY-MM-DD to MM-DD-YYYY\n- Change results time subfolder format to 12h am/pm (e.g. 11-23pm)\n- Update stop-model-comparison TTS voice from af_heart to am_michael\n\nCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>",
          "is_bot": false,
          "headline": "feat: improve voice announcer UX and results folder naming",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T03:33:11Z",
          "body_truncated": false,
          "is_coding_agent": true
        },
        {
          "oid": "75c0ec4bbc927f34462b9c020a537ff8aadc01cc",
          "body": null,
          "is_bot": false,
          "headline": "update voice",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T01:57:41Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "f33a80f86cd1650e4e7a0ec33ec5beb573d5ea50",
          "body": "Rename commands to run-model-comparison / simulate-model-comparison /\nstop-model-comparison for clarity. Add simulate-model-comparison.sh\nscript that generates realistic MindTrial log output without API calls,\nenabling end-to-end testing of the voice announcer.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "refactor: rename slash commands and add race simulation",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T01:48:40Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "66e378393345bda128d79254bd5b3a987bf62dd8",
          "body": null,
          "is_bot": false,
          "headline": "Update run-eval-suite.md to default to ci/cd evals",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T01:38:10Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9092c211d236b82f95c2e1a60242907e1350396a",
          "body": "Add Claude Code slash commands that launch model-vs-model evals with a\nlive Vin Scully-style sports commentator calling the play-by-play via\nKokoro TTS. The /loop prompt parses MindTrial's zerolog output to track\nper-model standings, speed comparisons, lead changes, and race completion.\n\n- .claude/c\n[…]\nnds/stop-eval-suite.md: kills eval, finalizes transcript\n- voice_announcer.md: original design spec and reference\n- .gitignore: exclude runtime artifacts (transcript, logs, summary)\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add live voice race announcer for MindTrial evals",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T01:17:27Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "efb66f8a9fe87817002478271a8f008ba2f3aeec",
          "body": "…ents\n\nfeat: Evaluation framework enhancements — CI/CD tasks, cost tracking, CircleCI migration",
          "is_bot": false,
          "headline": "Merge pull request #1 from ryan-circleci/feat/eval-framework-enhancem…",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T00:57:06Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "d5898f5e3070ce76e4232483a0721e08bd53aa80",
          "body": "Include all HTML, CSV, JSONL, and log outputs from:\n- General intelligence eval (71 tasks x 15 models, 2026-03-12)\n- CI/CD eval v1 (19 tasks x 15 models, 2026-03-13)\n- CI/CD eval v2 (100 tasks x 15 models, 2026-03-16)\n\nResults are committed directly since CI is not yet active.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "chore: add evaluation results for general-intelligence and CI/CD runs",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-17T00:49:39Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "aeb8b49bea996ed84cab23dc43cd8b10460d6ff2",
          "body": "Add 81 new CI/CD evaluation tasks covering CircleCI (25 total),\nGitHub Actions (12), GitLab CI/CD (10), Jenkins (8), Azure DevOps (5),\nshell scripting (11), Docker (9), Kubernetes (7), Git (4), and\ndeployment strategies (5). Mixed difficulty (easy/medium/hard) to\nbetter differentiate model capabilities.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: expand CI/CD benchmark from 19 to 100 tasks",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-16T17:07:46Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "3ac13e13594bf3cde3de60e31024e79483c7dfaa",
          "body": "Add config-eval-top3-cicd.yaml pointing at tasks-cicd.yaml with\nresults directed to results/cicd/ for clean separation from\ngeneral-intelligence eval runs.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add dedicated CI/CD eval config for top-3 providers",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-14T00:22:01Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9515c123b133f82e2113e064a41b087e39866e4e",
          "body": "Made-with: Cursor",
          "is_bot": false,
          "headline": "chore: gitignore compiled binary and evaluation results",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-13T15:45:51Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "cfaae1bd5f5228f074e635368b62c3788a223fca",
          "body": "Add OpenAI Responses API provider (openai_responses.go) for GPT-5.3+\nmodels, which use a different API surface than Chat Completions.\nUpdate the top-3 eval config to include GPT-5.4, GPT-5.4 Pro, and\nrefresh the model lineup across OpenAI, Google, and Anthropic.\nExpand the pricing catalog with latest model prices (GPT-5.4/Pro,\nGemini 3 Flash, Gemini 3.1 Flash-Lite, Claude Opus 4.5/4.6,\nClaude Haiku 4.5, Claude Sonnet 4.6). Bump openai-go SDK to v3.26.0.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add GPT-5.4 Responses API support and update top-3 eval config",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-13T15:44:32Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b4b626aea37958a16e119f547d8a682164c760e5",
          "body": "Focused evaluation config targeting the 5 latest models from each of\nthe top 3 providers (15 models total):\n\n- OpenAI: GPT-5.4 Pro, GPT-5.4 Thinking, GPT-5.3-Codex, GPT-5.3\n  Instant, GPT-5.2\n- Google: Gemini 3.1 Pro, 3.1 Flash-Lite, 3 Flash, 2.5 Pro, 2.5 Flash\n- Anthropic: Claude Opus 4.6, Sonnet 4\n[…]\n Haiku 4.5, Opus 4.5,\n  Sonnet 4.5\n\nUses env var auto-fallback for API keys (OPENAI_API_KEY, GOOGLE_API_KEY,\nANTHROPIC_API_KEY). Judge uses Anthropic Claude for semantic evaluation.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add top-3 provider eval config (OpenAI, Google, Anthropic)",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:57:18Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "1567787d5f21abb899cbb6dadd578a17026ecc03",
          "body": "Two complementary mechanisms for keeping API keys out of config YAML:\n\n1. ${ENV_VAR} expansion: any value in config.yaml can reference an env\n   var using ${VAR_NAME} syntax, expanded before YAML parsing. Unset\n   variables are left as-is.\n\n2. Automatic fallback: if a provider's api-key is empty aft\n[…]\n GOOGLE_API_KEY for\n   google, etc.). Works for both providers and judges.\n\nResolution order: ${...} expansion on raw YAML, then auto-fallback for\nstill-empty keys, then validation.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add environment variable support for API keys",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:56:35Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "afdeeeb7a692d5274fde63ed865f565d4b392326",
          "body": "Extend the CircleCI config with two evaluation workflows:\n\n1. scheduled-eval: Runs every Monday at 06:00 UTC against all configured\n   models, producing HTML, CSV, and JSONL reports stored as artifacts.\n\n2. new-model-eval: API-triggered pipeline for on-demand evaluation when\n   a new model drops. Tr\n[…]\nth use a run-eval job that builds MindTrial, runs the evaluation with\nconfigurable task files, appends JSONL results to a history file, and\nstores all outputs as CircleCI artifacts.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add scheduled and new-model-drop evaluation pipelines",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:23:45Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "8f9823090e36bf0edeb3d6a500fc358869cf4a27",
          "body": "Add a new JSONL (JSON Lines) formatter that outputs one JSON object per\nline, summarizing each provider/run combination with pass rate, duration,\ntoken counts, and estimated cost. Register the formatter with a -jsonl\nCLI flag.\n\nThis format is designed for append-friendly historical result files,\nenabling trend analysis across evaluation runs over time.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add JSONL formatter for historical result tracking",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:23:20Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "df9b365c7dd128466e4b8ac7c400abb6e7b4cae8",
          "body": "Introduce a pricing catalog (pricing/pricing.go) with per-million-token\ncosts for models across OpenAI, Anthropic, Google, DeepSeek, Mistral,\nxAI, Alibaba, Moonshot, and OpenRouter. Add a Model field to RunResult\nso pricing can be looked up at report time.\n\nUpdate all formatters (CSV, HTML, summary \n[…]\nimated Cost columns. Add helper functions in\nformatters/utils.go for token aggregation and cost formatting.\n\nUpdate runner tests and regenerate golden files to match the new output.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add per-task dollar cost calculation",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:23:15Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "169e6877777a81958192300f8beb7a9a5117b4b3",
          "body": "Replace the GitHub Actions workflow (.github/workflows/go.yml) with a\nCircleCI configuration. The build-and-test workflow runs on every\npush/PR, building the project and running the test suite with race\ndetection enabled on cimg/go:1.25.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: migrate CI from GitHub Actions to CircleCI",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:23:05Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b869d9acbdf6d4f13d03cd96f5ecebdfa3fd38fe",
          "body": "Add 19 new evaluation tasks targeting CI/CD and DevOps knowledge:\nshell arithmetic, pipeline stages, semantic versioning, environment\nvariable resolution, CircleCI config debugging, Docker port mapping,\nbuild parallelism, cache key resolution, Kubernetes resource calculation,\nGit branch extraction, \n[…]\nation, shell subshell bugs, and CircleCI orb concepts.\n\nThese complement the existing general intelligence tasks with\ndomain-specific challenges that test practical CI/CD reasoning.\n\nMade-with: Cursor",
          "is_bot": false,
          "headline": "feat: add CI/CD-domain evaluation tasks",
          "author_name": "Ryan E. Hamilton",
          "author_login": "ryan-circleci",
          "committed_at": "2026-03-12T18:22:54Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "8446cc90d786989bc8ee083424b9bc0df9720547",
          "body": "Add configurable `max-turns` limit per task to prevent unbounded\nconversation loops (e.g., when a model repeatedly requests exhausted\ntools). The limit is set globally in `task-config` and can be\noverridden per task. A value of 0 means unlimited.\n\nHandle a known Google Gemini issue where the model k\n[…]\nogFinishReason` debug logging to all provider conversation\nloops for improved observability.\n\nExample config:\n\n```\ntask-config:\n  max-turns: 100\n  tasks:\n    - name: \"my task\"\n      max-turns: 200\n```",
          "is_bot": false,
          "headline": "feat: Add conversation turn limit and Gemini tool fix",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:17:14Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "c67c86e37e774d2abf208e5468ec03ec2e595431",
          "body": "Mistral text-only prompts now build Content3.String instead of always\nusing ArrayOfContentChunk, while keeping chunked content for file-based\n(multimodal) prompts. This fixes compatibility with some models,\nwhich rejects chunk arrays for text-only input.\n\nThe Mistral client has been re-generated from the latest Open API spec.\n\nAlso adds structured error logging support across provider errors and\nlogger implementations, with tests for structured field propagation.",
          "is_bot": false,
          "headline": "fix: Use string content for text-only Mistral prompts",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:17:13Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "de5d6d78d4f6b1ee0bc44be5c072df1e71729e46",
          "body": "Add support for the new Gemini 3.1 Pro models and their expanded\nthinking level capabilities in the Google provider.\n\nChanges include:\n- Add `minimal` and `medium` to supported thinking levels.\n- Add `gemini-3.1-pro-preview` and `gemini-3.1-pro-preview-customtools`\n  to the default configuration.\n- Update documentation to clarify that the `minimal` thinking level\n  does not guarantee thinking is completely disabled.",
          "is_bot": false,
          "headline": "feat: Add Gemini 3.1 support",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:17:06Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b02a965a56bb9d998d6fee0c8520e2a50be047eb",
          "body": "Standardize the conversation loop across all AI providers to safely\nhandle multi-turn tool calling and prevent infinite loops.\n\n- Introduce `isTerminalStopReason` to explicitly distinguish between\n  final responses and intermediate tool calls.\n- Add `ErrNoActionableContent` to safely break loops whe\n[…]\n chunks before unmarshaling.\n- Log skipped preamble text during non-terminal turns.\n- Make final answer unmarshal more robust by accepting JSON primitive\n  types stored directly as the answer content.",
          "is_bot": false,
          "headline": "refactor: Standardize provider conversation loops",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:14:20Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "12a188ac50774a86f470fcb343aa9f96dc57f774",
          "body": "Added `stream` configuration option to `AnthropicModelParams`.\n\nImplemented transparent response buffering using the Anthropic SDK's\nAccumulate method to support this mode while maintaining compatibility\nwith existing validation and tool logic.\n\nIntroduced retry support for streaming failures and tr\n[…]\n\nmodel-parameters:\n  max-tokens: 16384\n  effort: max\n  stream: true\n```\n\nRecommended for requests with large `max-tokens` values or extended\nthinking to prevent HTTP timeouts on long-running requests.",
          "is_bot": false,
          "headline": "feat: Add streaming support for Anthropic models",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:14:18Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "32d4d8ee4dcb6b3e0f28c8145215fc7264f3826c",
          "body": "Replace the tool-use workaround (`record_summary` tool with forced\n`tool_choice`) with the native `output_config.format` API for\nstructured JSON responses.\n\nThe old approach was incompatible with extended thinking (which\nrequires `tool_choice: auto`) and conflated the response schema\ntool with actua\n[…]\nt.\n\nAlso batch parallel tool results into a single user message per\nturn, and add an explicit error for responses with no actionable\ncontent (e.g., thinking budget exhausted without producing output).",
          "is_bot": false,
          "headline": "fix: Use native structured output for Anthropic",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:14:18Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "783103daf2e38a9ba3a424cdbed63e3b2c6bdd19",
          "body": "MoonshotAI's thinking models (kimi-k2-thinking, kimi-k2.5) send\na non-standard reasoning_content field that the Open AI SDK\nsilently drops.\n\nIntroduce `CompletionHandler` interface to let delegating providers\ncustomize how streaming chunks are accumulated and how response\nmessages are converted to r\n[…]\nequent conversation\nturns.\n\nSet `max-tokens` to 16000 in default thinking model configs per\nMoonshot AI documentation recommendation.\n\nAlso remove discontinued `kimi-latest` model from default config.",
          "is_bot": false,
          "headline": "fix: Preserve MoonshotAI reasoning content",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-27T22:12:00Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "ebd9512657c21f094fdf563244deeb65fe281536",
          "body": "Add runs for OpenRouter (INTELLECT-3, Mercury, Seed 1.6, GLM 4.6V,\nGLM 4.7, Step3), xAI (Grok 4.1 Fast), Alibaba (Qwen3-Max snapshot,\nQVQ-Max, QwQ-Plus), and Moonshot AI (Kimi K2.5).",
          "is_bot": false,
          "headline": "chore: Add new default model run configurations",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-07T15:46:27Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "18055408478e2493eb0b0ebca73921d8afad51cb",
          "body": "Add `effort` parameter to Anthropic provider for adaptive extended\nthinking (values: low, medium, high, max). When set, uses\n`thinking: {type: \"adaptive\"}` with `output_config.effort` instead\nof the deprecated fixed `budget_tokens` approach.\n\nExample config:\n\n```\nmodel-parameters:\n  max-tokens: 8192\n  effort: max\n```",
          "is_bot": false,
          "headline": "feat: Add adaptive thinking support for Claude 4.6",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-06T23:47:28Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "878da54e97e6be63943f087f34c8278d58c5a894",
          "body": "Added `stream` configuration option to `AlibabaModelParams`.\n\nImplemented transparent response buffering in the internal OpenAI\nprovider to support this mode while maintaining compatibility with\nexisting validation and tool logic.\n\nExample config:\n\n```yaml\nmodel-parameters:\n  stream: true\n```\n\nRequired for some Alibaba models, such as QwQ, QVQ, and Qwen-Omni,\nwhich mandate streaming responses.",
          "is_bot": false,
          "headline": "feat: Add streaming support for Alibaba models",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-04T02:57:40Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "990b6529f7e63079ceb5a63c71b91070f44a9fc3",
          "body": "Add `text-only` boolean option to run configurations\nthat skips tasks requiring any file attachments.\n\nExample usage:\n\n```\nruns:\n  - name: \"Text Model\"\n    model: \"model-id\"\n    text-only: true\n```\n\nUseful for text-only models like some OpenRouter-hosted LLMs.\nDefault behavior unchanged (runs all tasks).",
          "is_bot": false,
          "headline": "feat: Add text-only mode to skip tasks with files",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-01T03:21:58Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "8ae8916cb505d7e32689e502b9470bb563ee4fe2",
          "body": "Ensure prompts and usage are propagated with result objects even when\nall retry attempts fail, by capturing the last attempt's result value.",
          "is_bot": false,
          "headline": "fix: Preserve execution metadata on retry failures",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-01T03:17:23Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "ae2885fc12670db3e6483b7fcab57e5f94b06893",
          "body": "Replace the community-based client with the official Open AI SDK.\nThis improves maintainability, ensures compatibility with OpenAI's\nlatest API changes, and provides better long-term support.\n\nKey changes:\n\n- Update dependencies in go.mod and go.sum\n- Refactor OpenAI, Alibaba and MoonshotAI provider\n[…]\nng\n\nAll existing functionality is preserved, and all tests pass. The new\nimplementation handles response formatting, tool calls, and error\nhandling consistently across all OpenAI-compatible providers.",
          "is_bot": false,
          "headline": "feat: Replace OpenAI client with official SDK",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-01T03:16:48Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "7205ab691d1beb9c1734afdf617f55c3d4d8c508",
          "body": "Introduce `disable-structured-output` run-configuration flag to\ntreat model responses as plain text unstructured data,\nbypassing JSON parsing.\nWhen enabled, the entire model's response becomes the final answer,\nwith Title and Explanation populated with placeholders.\n\nKey changes:\n\n- Config: Add Disa\n[…]\nproviders:\n  - name: openai\n    runs:\n      - name: unstructured\n        model: gpt-4o\n        disable-structured-output: true\n```\n\nEnhances compatibility with models unable to generate JSON reliably.",
          "is_bot": false,
          "headline": "feat: Add flag for unstructured responses",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-02-01T03:12:22Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "885ef526df7c0c87bb77d1d7d720686775db20fc",
          "body": "Remove extra blank line after OpenAI parameters list to maintain\nconsistent style across all provider parameter sections.",
          "is_bot": false,
          "headline": "docs: fix formatting in README parameters section",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-01-17T23:41:08Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "6c5cd533ddf4bcc6e16a0d4b01d48d1b67234aeb",
          "body": "Introduce OpenRouter as a new provider for accessing models through\nthe OpenRouter API. This implementation uses the official OpenAI\nSDK v3, providing a modern and maintainable client foundation that\nmay later replace the community-based client.\n\nThe `OpenRouter` provider can be configured with:\n\n- \n[…]\nities.\n- Adds `Ptr[T]` utility function for creating pointers to values.\n- Includes comprehensive test coverage for parameter mapping.\n- Updates documentation in README.md with configuration examples.",
          "is_bot": false,
          "headline": "feat: Add OpenRouter provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2026-01-17T19:22:30Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "43c2d878b8164777f4d207e5c07a0010bb84a201",
          "body": "Expose rate metrics in summary outputs to make runs easier\nto compare at a glance.\n\nDefinitions:\n\n- Pass Rate = Passed/(Passed+Failed+Error)\n- Accuracy = Passed/(Passed+Failed)\n- Error Rate = Error/(Passed+Failed+Error)\n- Skipped tasks are excluded; rates are 0 when denominator is 0.\n\nChanges:\n\n- Add Pass Rate, Accuracy, and Error Rate to summary.log output.\n- Add the same columns to the HTML summary table (sortable).",
          "is_bot": false,
          "headline": "feat: Add pass/accuracy/error rates to summary",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-12-23T19:35:43Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "12a69da5ce2aef364d06222d28b3fd4801a5f95c",
          "body": "Extended OpenAI provider to support GPT-5.2 model with new\nreasoning_effort values (none, minimal, low, medium, high, xhigh)\nand verbosity parameter (low, medium, high) for output control.\n\nChanges:\n- Extended reasoning-effort validation to 6 values including new\n  xhigh option for GPT-5.2's maximum\n[…]\nhandling.\n- Added gpt-5.2 model configuration with xhigh reasoning effort\n  and medium verbosity.\n- Updated README documentation with complete parameter details\n  and legacy model compatibility notes.",
          "is_bot": false,
          "headline": "fix: Add support for GPT-5.2 models",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-12-12T18:17:51Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "214540c7855951e7aa8965423a6ecc450125bda8",
          "body": "Extend vision task support to latest Mistral AI models.\n\n- Support mistral-large, mistral-medium, mistral-small variants.\n- Add ministral, pixtral, magistral model families.",
          "is_bot": false,
          "headline": "fix: Update Mistral AI vision model support",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-12-08T00:26:37Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "5243da369e4d53a0fcb581dfeb9d836615d6487e",
          "body": "Add support for Google Gemini 3 models with new API features:\n\n- Add `thinking-level` parameter to control reasoning depth (low/high)\n- Add `media-resolution` parameter for image token allocation control\n- Add `text-response-format-with-tools` for backward compatibility\n  with pre-Gemini 3 models th\n[…]\ned text and tool\n  responses correctly\n\nPre-Gemini 3 models can force text response format with\ntools using `text-response-format-with-tools` parameter.\n\nAlso, fix several minor typos in task prompts.",
          "is_bot": false,
          "headline": "feat: Add Gemini 3 with thinking and media control",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-11-20T05:56:34Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "0755e3331f801643d7868ccae889d4f47870f62f",
          "body": "Introduce Moonshot AI (Kimi) as a new provider via OpenAI-compatible\nAPI. The current provider's implementation delegates request\nprocessing logic to the existing `OpenAI` provider.\n\nThe `Moonshot AI` provider can be configured with the following\nproperties:\n\n- `name`:     Must be set to \"moonshotai\n[…]\nes.\n  `LegacyJsonSchema`, adds format instruction to prompt while keeping\n  json_schema response format.\n  `LegacyJsonObject`, adds format instruction to prompt and uses\n  json_object response format.",
          "is_bot": false,
          "headline": "feat: Add Moonshot AI (Kimi) provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-11-12T21:39:53Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "e2678265b465416665fa3f8a429c84571081b045",
          "body": "Add support for persistent shared directories that enable data\nsharing across all tool invocations within a single task:\n\n- Add `shared-dir` field to `ToolConfig` for directory configuration.\n- Implement lazy directory creation on the first use.\n- Mount the shared directory to `shared-dir` path insi\n[…]\ncutor\n      shared-dir: /app/shared\n      auxiliary-dir: /app/data\n```\n\nFiles in `shared-dir` persist across all tool calls within a task,\nwhile `auxiliary-dir` contents are reset between invocations.",
          "is_bot": false,
          "headline": "feat: Add persistent shared directory for tool use",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-11T01:09:29Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "f7b296e00be0d6a52168d97e69b85c3956592833",
          "body": "Add validation to ensure Docker images for enabled tools\nare available locally before tasks begin:\n\n- Add ValidateTool method to DockerToolExecutor to check image\n  availability via Docker API.\n- Integrate validation in `defaultRunner` to validate\n  all enabled tools before execution starts.\n- Updat\n[…]\ndescription to clarify ephemeral container behavior.\n\nValidation prevents runtime failures by catching missing images\nearly, providing clear guidance on how to resolve issues before\nany tasks execute.",
          "is_bot": false,
          "headline": "feat: Validate tool images before task execution",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-11T01:04:43Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "232969edf87a3f63371174b9ed3c047ec89a714b",
          "body": "- Mount copies of all task files into `auxiliary-dir` when set.\n- Rename `file_mappings` property to `parameter-files` for consistency,\n  and to better distinguish it from auxiliary files.\n- Timeout now covers only the container run, not setup/cleanup.",
          "is_bot": false,
          "headline": "feat: Mount task files in tool container",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-09T18:05:48Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9b977533237bb132a48a498b2c0389799e1738a3",
          "body": "- Introduces unique TraceID (ULID) for each task execution result.\n  TraceIDs enable tracing and correlation across artifacts and logs.\n  CSV and text formatters include TraceID column; HTML shows\n  indicator icon with TraceID tooltip; logger prepends TraceID\n  to all result messages.\n\n- Adds tool usage indicators to HTML output with sorted tool\n  names and visual icon.",
          "is_bot": false,
          "headline": "feat: Add ID and tool usage indicators to results",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-03T00:05:30Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "2dd05e474b3138368cdb18564d6ec71a751ce61d",
          "body": "- Add Claude 4.5 Sonnet model with extended thinking\n- Update DeepSeek model names from V3.1 to V3.2",
          "is_bot": false,
          "headline": "fix: Add recent models to default config file",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-02T15:35:36Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "5ffea0b8745aaf55910a6eaf4accb13daa6a6035",
          "body": "- Add global tools config and per-task tool-selector.\n- Validate tool parameter schemas and task tool references.\n- Update all providers to support tool calling with conversation loops.\n- Implement DockerToolExecutor for secure tool execution in containers.\n- Track tool usage statistics (call count,\n[…]\napp/main.py\n```\n\nExample tool configuration:\n\n```yaml\ntool-selector:\n  tools:\n    - name: python-code-executor\n      max-calls: 10\n      timeout: 60s\n      max-memory-mb: 512\n      cpu-percent: 25\n```",
          "is_bot": false,
          "headline": "feat: Enable tool use in tasks",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-10-02T15:35:25Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "2f4a3a187dd9dea1e5804d547d936e1580464c87",
          "body": "Introduce Alibaba (Qwen) as a new provider via OpenAI-compatible API.\nThe current provider's implementation delegates the request processing\nlogic to the existing `OpenAI` provider and may not support all models\nand all options.\n\nThe `Alibaba` provider can be configured with the following properties\n[…]\nokens\n                             available to the model for generating\n                             a response.\n- Extend test coverage of `Google` model parameters.\n- Fix `DeepSeek` name references.",
          "is_bot": false,
          "headline": "feat: Add Alibaba (Qwen) provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-23T21:16:27Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "ea8a0ff0c574b4f95eaaec38f180b47f4f14eedd",
          "body": "- Add customizable judge prompt template with structured context.\n- Add verdict format and passing verdicts configuration.\n- Add default judge template with semantic equivalence logic.\n- Implement per-task validation rules resolution and caching.\n- Fix support for mapping objects in value sets.\n- En\n[…]\n      enum: [\"excellent\", \"good\", \"poor\"]\n      required: [\"quality_score\"]\n      additionalProperties: false\n    passing-verdicts:\n      - quality_score: \"excellent\"\n      - quality_score: \"good\"\n```",
          "is_bot": false,
          "headline": "feat: Judge validation with custom prompt template",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-19T16:11:27Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "d11edbc12141282ea9cafd6b2e86f98eaaa5168c",
          "body": "- `quiz - multiple choice questions - v1`:\n    accept answers without question numbers, as long as the choices are\n    correct and in sequence. (Anthropic, DeepSeek, Mistral AI)",
          "is_bot": false,
          "headline": "fix: Adjust task validation for model edge cases",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-16T18:47:58Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "ea24056e0a2f91ae2d635d16295acfdab7ae86fc",
          "body": "- Add support for structured schema-based result formats using JSON\n  schemas in task definitions alongside existing plain text formats.\n- Configure system prompt delivery per result format type with new\n  'enable-for' setting:\n  - `all`: Send for all tasks.\n  - `text`: Send for tasks with plain tex\n[…]\n Seed for deterministic generation.\n\nExample structured answer format:\n\n```\n  response-result-format:\n    type: object\n    properties:\n      answer: {type: string}\n      confidence: {type: number}\n```",
          "is_bot": false,
          "headline": "feat: Add structured JSON result format support",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-16T18:35:20Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "d9cb955aa4a36d45c7c0417391fde8705bdd3892",
          "body": "Refresh dependencies in go.mod.",
          "is_bot": false,
          "headline": "fix: Update dependencies",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-03T14:00:21Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b64109916b136ef09bc700157b09a99390eb5243",
          "body": "Fill upper triangle of run comparison matrix with disagreement\npercentages between runs. Lower triangle continues to show\nagreement data, creating a symmetrical view of run relationships.\n\nClicking disagreement cells opens a modal with tasks where the\ntwo runs produced different results. Updated legend reflects\nthe new matrix layout and color coding.",
          "is_bot": false,
          "headline": "feat: Add disagreement data to comparison matrix",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-03T13:59:10Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "42f26988b97f59a1b58459434b297ffe6347ad52",
          "body": "Add support for customizable system prompt templates in task\nconfiguration to control how response format instructions are\npresented to AI models:\n\n- Add SystemPrompt struct with Template field to config package\n- Support global system prompt template in task-config section\n- Allow per-task system p\n[…]\nhe final answer in exactly this format:\n        {{.ResponseResultFormat}}\n    tasks:\n      - name: \"my-task\"\n        system-prompt:\n          template: \"Answer format: {{.ResponseResultFormat}}\"\n  ```",
          "is_bot": false,
          "headline": "feat: Add configurable system prompt templates",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-09-03T13:58:52Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "dde5cfc4eef2682e01787c11449b52706317d33c",
          "body": "- Update the minimum required Go version to match the ```go.mod``` file.\n- Add missing blockquote.",
          "is_bot": false,
          "headline": "docs: Update README.md",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-28T16:43:00Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "4d463de2ab876596d93a0810d44e756256d42c2f",
          "body": "Add support for xAI (Grok) models as a new provider option.\n\nThe xAI provider can be configured with the following properties:\n\n  - `name`:    Must be set to \"xai\"\n  - `api-key`: API key for the xAI (Grok) models provider\n\nSupported model-specific parameters include:\n\n  - `temperature`: Controls ran\n[…]\nrok 4 is a reasoning model.\n  - ```presence-penalty``` and ```frequency-penalty``` parameters\n    are not supported by reasoning models.\n  - Grok 4 does not support a ```reasoning-effort``` parameter.",
          "is_bot": false,
          "headline": "feat: Add xAI (Grok) provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-28T14:57:25Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "d451d66def0472504548a75ca9f841c55fe7185a",
          "body": "Update response-result-format specifications in default\n```tasks.yaml``` to reduce model response formatting errors and\nimprove task validation accuracy.",
          "is_bot": false,
          "headline": "fix: Clarify answer format to reduce task failures",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-26T16:25:26Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "112b079de234d8a2a38b482ee6d7007e97f4c562",
          "body": "Align run naming with current DeepSeek models and support both\nthinking and non-thinking modes.\n\n- Rename \"DeepSeek-R1 - latest\" to \"DeepSeek-V3.1 - latest (thinking\n  mode)\" keeping deepseek-reasoner model.\n- Add \"DeepSeek-V3.1 - latest (non-thinking mode)\" using\n  deepseek-chat model.",
          "is_bot": false,
          "headline": "fix: Support DeepSeek V3.1 thinking modes",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-26T00:09:43Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "a585c5a94a470aa4b040621eeb0f34d8dec1f1dc",
          "body": "Enable clicking run comparison matrix cells to open a dialog\nlisting overlapping tasks grouped by status.",
          "is_bot": false,
          "headline": "feat: Add run details dialog to comparison matrix",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-25T20:30:15Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "8425f9e403ec8ed240ed2decf0f72b52cf649a3f",
          "body": "The summary table in the HTML report now includes checkboxes to select\nmultiple runs.\nA \"Compare Selected Runs\" button uses this selection to generate a\nvisual comparison matrix in a new window.\n\nThe matrix shows the percentage of agreement between any two runs\non their common, non-skipped tasks.\nCe\n[…]\nor-coded to show the distribution of matching results\n(passed, failed, error), with saturation indicating overlap strength.\nThe matrix is interactive, with hover effects to highlight rows and\ncolumns.",
          "is_bot": false,
          "headline": "feat: Add run comparison matrix to HTML report",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-21T20:16:13Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "71cf5a7bb8a680839589dfd633362e258fac7e94",
          "body": "Prevent duplicate names across configurations to avoid ambiguous\nresults and improve clarity in reports.\n\n- Enforce unique names within provider runs.\n- Enforce unique names within judge configurations.\n- Enforce unique names within task definitions.",
          "is_bot": false,
          "headline": "feat: Enforce unique name for runs, judges & tasks",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-21T20:16:13Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "64bde07c899ab893ff1a56f7bfd7a35994433c9e",
          "body": "Rename duplicate ```riddle - split words - v3``` to ```riddle - split words - v4```.",
          "is_bot": false,
          "headline": "fix: Fix duplicate task name",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-21T20:16:12Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "e281f8c760a646d193abdb58ad0a129f68f69ce6",
          "body": "- `quiz - multiple choice questions - v1`: tolerate trailing\n  whitespace in individual options from some models (Anthropic).\n- `quiz - analogies`: accept both verb and noun answers\n  (\"eat\" / \"food\") since \"sleep\" functions as both.\n- `riddle - first letter - v3`: fix format description\n  (4 groups = 4 letters), and add an alternative valid solution\n  with a new first letter for group #2.",
          "is_bot": false,
          "headline": "fix: Adjust task validation for model edge cases",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-16T17:01:31Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "4683f60b98e671a9840a7f57b83647165dfc1e69",
          "body": "Introduces `trim-lines` validation rule, token usage tracking,\nand enhanced HTML reports with regex filtering capabilities.\n\nNew Validation Rule:\n\n- `trim-lines` (boolean, default: false) - Trims leading/trailing\n  whitespace from each line while preserving internal spaces and\n  normalizing CRLF to \n[…]\nsections.\n- Improved filtering tooltip text with regex examples.\n\nToken Usage Tracking:\n\n- Added `TokenUsage` struct to run result details.\n- Displays usage statistics in HTML reports and CSV exports.",
          "is_bot": false,
          "headline": "feat: Add trim-lines, token usage, enhanced HTML",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-16T15:55:15Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "990411b06017959ea855ba5c9aecd85534f655f7",
          "body": "Retry on documented transient errors from OpenAI API requests.",
          "is_bot": false,
          "headline": "fix: Enhance OpenAI provider with transient error handling",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-11T18:11:56Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "1e9743f238598ef5616588e622225551ffe0c776",
          "body": "- Add OpenAI o4-mini, o3, GPT-5 mini, and GPT-5 with high reasoning\n  effort.\n- Add Google Gemini 2.5 Flash and Pro models.\n- Add Anthropic Claude 4.1 Opus model.\n- Increase rate limits for o1-mini and o3-mini to 20 requests/minute.\n- Increase request timeout for DeepSeek to 15 minutes.",
          "is_bot": false,
          "headline": "fix: Update default model list and rate limits",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-11T18:11:55Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "d29677a380b50363f3b58fa081534c42e55dfc12",
          "body": "- Add structured Details (Answer, Validation, Error) to RunResult.\n- CSV-only: Serialize Details as structured JSON.\n- HTML-only: Enrich HTML report with structured details section.\n- HTML-only: Add header filtering and click-to-filter functionality.\n- HTML-only: Add search by provider, run, and task names.\n- HTML-only: Enable column sorting.\n- HTML-only: Improve accessibility and add schema.org annotations.",
          "is_bot": false,
          "headline": "feat: Introduce structured Details and enhance HTML report",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-11T18:11:30Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "27fb07a208067e47ca8b3201e32575a244454314",
          "body": "- Regenerate Go models from Mistral AI OpenAPI spec.\n- Refresh dependencies in go.mod.",
          "is_bot": false,
          "headline": "fix: update Mistral AI client and dependencies",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-08-10T15:45:56Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "ae2080e7f1e3c4ab622540dbcf17f4e27825b460",
          "body": "This commit introduces a major refactoring to decouple task execution logic\nfrom runners and validators, and to standardize logging throughout the\napplication.\n\n- A new execution package centralizes provider execution, handling\n  retries and rate limiting through a unified Executor.\n- Default runner\n[…]\nnner is simplified by utilizing a logging.Logger\n  instance to handle logging and message emitting.\n- Providers and validators now receive a logging.Logger\n  instance, allowing for contextual logging.",
          "is_bot": false,
          "headline": "refactor: Decouple execution logic and standardize logging",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-30T03:28:50Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "28564435feba3d3ae8e01a77ca8a86ba1a17b07d",
          "body": "Add comprehensive validation framework supporting both exact value\nmatching and LLM-based semantic evaluation:\n\n- Decouple validation logic from providers and move it to a\n  new `validators` package.\n- Introduce value-match validator for traditional value matching.\n- Add judge validator using LLM pr\n[…]\nmodel: \"gpt-4o-mini\"\n\t       max-requests-per-minute: 10\n  ```\n\nJudge selector example:\n  ```yaml\n     validation-rules:\n       judge:\n         enabled: true\n\t name: \"my-judge\"\n\t variant: \"fast\"\n  ```",
          "is_bot": false,
          "headline": "feat: Add LLM judge validation feature",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-30T03:20:14Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "5354323c42d8c48bcf186a7453320e04bec6b7bf",
          "body": "Adds post-processing to ensure proper Go formatting with gofmt.\nNote: File paths must not contain spaces due to known generator issue:\nhttps://github.com/OpenAPITools/openapi-generator/issues/10839",
          "is_bot": false,
          "headline": "style: Format Go files generated by openapi-generator using gofmt",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-08T19:46:19Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "3c02134f7935c67e049047358aee051a2fc40c3c",
          "body": "Add support for image processing tasks in the Mistral AI provider.\nThe provider now supports uploading and processing image files with\nvision-capable models.\n\nSupported vision models:\n\n  - pixtral-12b-latest\n  - pixtral-large-latest\n  - mistral-medium-latest\n  - mistral-small-latest",
          "is_bot": false,
          "headline": "feat: Add vision support to Mistral AI provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-08T19:05:01Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "697bed78a77dc98adbbba1d544d8d2de5c7c91ee",
          "body": "Use v0.4.1 instead.",
          "is_bot": false,
          "headline": "fix: Retract v0.4.0 due to broken history",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-07T17:30:18Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "9505c586464848c95c0f78aae675002d1b86a26b",
          "body": "Add support for automatic retry of failed requests due to rate limiting\nor other transient errors.\nA retry policy can be configured at the provider level to apply to all\nruns, or at the individual run level to override the provider setting.\n\nThe retry policy defines the following properties:\n\n  - `m\n[…]\n backoff starting with the initial delay.\n\nRate limiting is respected during retries - if a run configuration has\n`max-requests-per-minute` set, the rate limiter will be applied to each\nretry attempt.",
          "is_bot": false,
          "headline": "feat: Add configurable retry policy",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-07T16:48:18Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "debb2c06fb5e40437ab9a89fa9f88274900e5d6c",
          "body": "Add support for Mistral AI models as a new provider option.\n\nThe Mistral AI provider can be configured with the following properties:\n\n  - `name`:    Must be set to \"mistralai\"\n  - `api-key`: API key for the Mistral AI generative models provider\n\nSupported model-specific parameters include:\n\n  - `te\n[…]\n  - `prompt-mode`: When set to \"reasoning\", instructs model to reason\n                   if supported\n  - `safe-prompt`: Enables content filtering for compliance with\n                   usage policies",
          "is_bot": false,
          "headline": "feat: Add Mistral AI provider",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-07-07T16:47:13Z",
          "body_truncated": true,
          "is_coding_agent": false
        },
        {
          "oid": "b2ca7f427c612ede5b9e8a12dae305d351e79f99",
          "body": "- Add fifteen visual tasks.",
          "is_bot": false,
          "headline": "fix: Add new visual tasks",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-06-27T16:30:55Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "671a19a85316cd43fc0234ef0a80255e9b04ae82",
          "body": "- Ensure interactive configuration selector scrolls to keep\n  cursor visible when navigating beyond the viewport bounds.\n- Improve window resize logic for consistent TUI layout adjustments.",
          "is_bot": false,
          "headline": "fix: Add scrolling to config selector and improve resize layout",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-06-27T16:27:56Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "3168044731ea9ec82fec88e97b90b7228c814f4a",
          "body": "- Add new mostly visual tasks.\n- Ensure the images have appropriate dimensions\n  to minimize token usage (adopts common tile\n  sizes of 384px or 512px).\n- Move the prompt text after the file data for\n  improved context integrity.",
          "is_bot": false,
          "headline": "fix: Improve and extend visual task coverage",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-06-18T04:38:04Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "22da64fec18f2782ad41ac5d4ec2619ef6df7283",
          "body": "- Add alternative valid answer to `riddle - web words - v2`.\n- Handle common edge cases in result validation for specific models\n  (`o4-mini` dropping spaces in visual task, `Claude 4.0` including answer\n  text in quiz responses).",
          "is_bot": false,
          "headline": "fix: Improve task validation for model-specific output formats",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-31T13:41:35Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "b989ac53e22ca8d10b18a7d078a07f08b3c5062f",
          "body": "Introduces `validation-rules` at both the global `task-config` level and\nper `task`.\nThese rules allow customization of how model-generated answers are\ncompared against expected results.\n\nSupported Rules:\n\n- `case-sensitive` (boolean, default: false)\n- `ignore-whitespace` (boolean, default: false)\n\nTask-specific rules override global settings.",
          "is_bot": false,
          "headline": "feat: Add validation rules for flexible answer checking",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-31T02:10:43Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "c8f4657c17010f2aaa65caa10b26f2190cc97bf0",
          "body": "- Add `Claude Opus 4` and `Claude Sonnet 4` models to default\n  configuration file.\n- Increase default `max-requests-per-minute` for Claude models.\n- Update Anthropic provider dependency.\n- Refactor Anthropic provider implementation to align with the updated API.",
          "is_bot": false,
          "headline": "fix: Add latest `Claude 4` models to default configuration",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-24T02:21:21Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "1d5ad621cd9625f5ff2af7f2d83818b292318a9f",
          "body": "Change `OnceWithContext` closures to receive *TaskFile as parameter\ninstead of capturing `basePath` from enclosing scope. This ensures\n`Content()` uses the current `basePath` value after `SetBasePath()` calls.",
          "is_bot": false,
          "headline": "fix: Resolve issue where attached file could not be opened",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-24T00:54:10Z",
          "body_truncated": false,
          "is_coding_agent": false
        },
        {
          "oid": "a2f7f5d8f963fcb418e4e88498418eaa887ee258",
          "body": "Replace the previous `time.Tick` based rate limiting implementation\nwith `golang.org/x/time/rate.Limiter`.\n\nThis change provides a more robust and flexible approach to rate\ncontrol. The new `rate.Limiter` allows for an initial burst of tasks up\nto the `MaxRequestsPerMinute` limit and then ensures the rate is\nmaintained. This improves throughput by executing tasks as quickly as\npossible within the defined limits.",
          "is_bot": false,
          "headline": "perf: Use `x/time/rate` for rate limiting",
          "author_name": "Petr Malik",
          "author_login": "petmal",
          "committed_at": "2025-05-19T18:49:36Z",
          "body_truncated": false,
          "is_coding_agent": false
        }
      ],
      "releases_count": 33,
      "commits_last_year": 86,
      "latest_release_at": "2026-03-18T22:43:28Z",
      "latest_release_tag": "v0.17.0",
      "releases_from_tags": true,
      "days_since_last_push": 127,
      "active_weeks_last_year": 21,
      "days_since_latest_release": 128,
      "mean_days_between_releases": 13.2
    },
    "community": {
      "has_readme": false,
      "has_license": false,
      "has_description": false,
      "has_contributing": false,
      "health_percentage": null,
      "has_issue_template": false,
      "has_code_of_conduct": false,
      "has_pull_request_template": false
    },
    "ecosystem": {
      "packages": [
        {
          "name": "github.com/CircleCI-Research/evalbench",
          "exists": true,
          "license": null,
          "keywords": [],
          "ecosystem": "go",
          "matches_repo": true,
          "registry_url": "https://pkg.go.dev/github.com/CircleCI-Research/evalbench",
          "is_deprecated": false,
          "latest_version": "v0.17.0",
          "repository_url": "https://github.com/CircleCI-Research/evalbench",
          "versions_count": 33,
          "total_downloads": null,
          "dependents_count": null,
          "deprecation_note": null,
          "maintainers_count": null,
          "monthly_downloads": null,
          "first_published_at": null,
          "latest_published_at": "2026-03-18T22:43:28Z",
          "latest_version_yanked": null,
          "days_since_latest_publish": 128
        }
      ]
    },
    "popularity": {
      "forks": 0,
      "stars": 0,
      "watchers": 0,
      "fork_history": {
        "days": [],
        "complete": true,
        "collected": 0,
        "total_forks": 0
      },
      "star_history": {
        "days": [],
        "complete": true,
        "collected": 0,
        "total_stars": 0,
        "collected_at": null
      },
      "open_issues_and_prs": 0
    },
    "ai_readiness": {
      "has_nix": false,
      "example_dirs": [],
      "has_llms_txt": false,
      "has_dockerfile": false,
      "has_mcp_signal": false,
      "bootstrap_files": [],
      "api_schema_files": [],
      "has_devcontainer": false,
      "typecheck_configs": [],
      "toolchain_manifests": [
        "go.mod"
      ],
      "largest_source_bytes": 73199,
      "source_files_sampled": 442,
      "oversized_source_files": 5,
      "agent_instruction_files": [
        "AGENTS.md"
      ],
      "agent_instruction_max_bytes": 3085
    },
    "dependencies": {
      "manifests": [
        "go.mod"
      ],
      "advisories": {
        "error": null,
        "scope": null,
        "source": null,
        "findings": [],
        "collected": false,
        "malicious": [],
        "truncated": false,
        "by_severity": {},
        "advisory_count": 0,
        "affected_count": 0,
        "assessed_count": 0,
        "malicious_count": 0,
        "assessed_package": null,
        "unassessed_count": 0,
        "direct_affected_count": 0
      },
      "ecosystems": [
        "go"
      ],
      "dependencies": [
        {
          "name": "github.com/charmbracelet/x/term",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.2.1"
        },
        {
          "name": "github.com/containerd/errdefs",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.0.0"
        },
        {
          "name": "github.com/docker/docker",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v28.5.2+incompatible"
        },
        {
          "name": "github.com/google/uuid",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.6.0"
        },
        {
          "name": "github.com/invopop/jsonschema",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.13.0"
        },
        {
          "name": "github.com/kaptinlin/jsonrepair",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.2.8"
        },
        {
          "name": "github.com/oklog/ulid/v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.1.1"
        },
        {
          "name": "github.com/openai/openai-go/v3",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v3.26.0"
        },
        {
          "name": "github.com/santhosh-tekuri/jsonschema/v6",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v6.0.2"
        },
        {
          "name": "github.com/sethvargo/go-retry",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.3.0"
        },
        {
          "name": "github.com/stretchr/testify",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.11.1"
        },
        {
          "name": "golang.org/x/time",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.14.0"
        },
        {
          "name": "google.golang.org/genai",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.48.0"
        },
        {
          "name": "gopkg.in/validator.v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.0.1"
        },
        {
          "name": "gopkg.in/yaml.v3",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v3.0.1"
        },
        {
          "name": "github.com/anthropics/anthropic-sdk-go",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.26.0"
        },
        {
          "name": "github.com/charmbracelet/bubbles/v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.0.0-beta.1"
        },
        {
          "name": "github.com/charmbracelet/bubbletea/v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.0.0-beta.1"
        },
        {
          "name": "github.com/charmbracelet/lipgloss/v2",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v2.0.0-beta.1"
        },
        {
          "name": "github.com/cohesion-org/deepseek-go",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.3.3"
        },
        {
          "name": "github.com/go-playground/validator/v10",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v10.30.1"
        },
        {
          "name": "github.com/rs/zerolog",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.34.0"
        },
        {
          "name": "github.com/sergi/go-diff",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v1.4.0"
        },
        {
          "name": "golang.org/x/exp",
          "manifest": "go.mod",
          "ecosystem": "go",
          "version_constraint": "v0.0.0-20260218203240-3dfff04db8fa"
        }
      ],
      "all_dependencies": {
        "error": "GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository",
        "source": null,
        "packages": [],
        "collected": false,
        "truncated": false,
        "total_count": null,
        "direct_count": null,
        "indirect_count": null
      }
    },
    "maintainership": {
      "issues": {
        "open_prs": 0,
        "merged_prs": 1,
        "open_issues": 0,
        "closed_ratio": null,
        "closed_issues": 0,
        "closed_unmerged_prs": 0
      },
      "bus_factor": 1,
      "bot_contributors": 0,
      "top_contributors": [
        {
          "type": "User",
          "login": "petmal",
          "commits": 79,
          "avatar_url": "https://avatars.githubusercontent.com/u/4350408?v=4"
        },
        {
          "type": "User",
          "login": "ryan-circleci",
          "commits": 37,
          "avatar_url": "https://avatars.githubusercontent.com/u/104376313?v=4"
        }
      ],
      "contributors_sampled": 2,
      "top_contributor_share": 0.681
    },
    "quality_signals": {
      "has_ci": true,
      "has_tests": true,
      "ci_workflows": [
        "go.yml"
      ],
      "has_docs_dir": false,
      "linter_configs": [],
      "has_editorconfig": false,
      "has_linter_config": false,
      "has_precommit_config": false
    },
    "security_signals": {
      "lockfiles": [
        "go.sum"
      ],
      "scorecard": {
        "checks": [
          {
            "name": "Binary-Artifacts",
            "score": 10,
            "reason": "no binaries found in the repo",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#binary-artifacts"
          },
          {
            "name": "Branch-Protection",
            "score": 5,
            "reason": "branch protection is not maximal on development and all release branches",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#branch-protection"
          },
          {
            "name": "CI-Tests",
            "score": 10,
            "reason": "1 out of 1 merged PRs checked by a CI test -- score normalized to 10",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#ci-tests"
          },
          {
            "name": "CII-Best-Practices",
            "score": 0,
            "reason": "no effort to earn an OpenSSF best practices badge detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#cii-best-practices"
          },
          {
            "name": "Code-Review",
            "score": 0,
            "reason": "Found 0/26 approved changesets -- score normalized to 0",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#code-review"
          },
          {
            "name": "Contributors",
            "score": 6,
            "reason": "project has 2 contributing companies or organizations -- score normalized to 6",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#contributors"
          },
          {
            "name": "Dangerous-Workflow",
            "score": 10,
            "reason": "no dangerous workflow patterns detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#dangerous-workflow"
          },
          {
            "name": "Dependency-Update-Tool",
            "score": 0,
            "reason": "no update tool detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#dependency-update-tool"
          },
          {
            "name": "Fuzzing",
            "score": 0,
            "reason": "project is not fuzzed",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#fuzzing"
          },
          {
            "name": "License",
            "score": 10,
            "reason": "license file detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#license"
          },
          {
            "name": "Maintained",
            "score": 0,
            "reason": "0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#maintained"
          },
          {
            "name": "Packaging",
            "score": null,
            "reason": "packaging workflow not detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#packaging"
          },
          {
            "name": "Pinned-Dependencies",
            "score": 0,
            "reason": "dependency not pinned by hash detected -- score normalized to 0",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#pinned-dependencies"
          },
          {
            "name": "SAST",
            "score": 0,
            "reason": "SAST tool is not run on all commits -- score normalized to 0",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#sast"
          },
          {
            "name": "Security-Policy",
            "score": 0,
            "reason": "security policy file not detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#security-policy"
          },
          {
            "name": "Signed-Releases",
            "score": null,
            "reason": "no releases found",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#signed-releases"
          },
          {
            "name": "Token-Permissions",
            "score": 0,
            "reason": "detected GitHub workflow tokens with excessive permissions",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#token-permissions"
          },
          {
            "name": "Vulnerabilities",
            "score": 0,
            "reason": "44 existing vulnerabilities detected",
            "documentation_url": "https://github.com/ossf/scorecard/blob/c395761df6afe1a69e476bc60a013a94bcbc153f/docs/checks.md#vulnerabilities"
          }
        ],
        "commit": "64b9f020db13d4629897badb64f94f85a8003544",
        "ran_at": "2026-07-25T08:58:05Z",
        "aggregate_score": 3,
        "scorecard_version": "v5.5.0"
      },
      "has_codeql_workflow": false,
      "has_security_policy": false,
      "has_dependabot_config": false
    },
    "contribution_flow": {
      "collected": true,
      "ci_last_run_at": "2026-03-19T15:13:01Z",
      "oldest_open_prs": [],
      "last_merged_pr_at": "2026-03-19T15:10:06Z",
      "ci_last_conclusion": "SUCCESS",
      "oldest_open_issues": []
    }
  },
  "config": {
    "disabled_metrics": [],
    "disabled_categories": [],
    "disabled_components": {}
  },
  "source": {
    "url": "https://github.com/CircleCI-Research/evalbench",
    "host": "github.com",
    "name": "evalbench",
    "owner": "CircleCI-Research"
  },
  "metrics": {
    "overall": {
      "key": "overall",
      "band": "at_risk",
      "name": "Overall health",
      "note": null,
      "notes": [],
      "value": 43,
      "inputs": {
        "security": 30,
        "vitality": 56,
        "community": 12,
        "governance": 57,
        "engineering": 51
      },
      "components": []
    },
    "categories": [
      {
        "key": "vitality",
        "band": "moderate",
        "name": "Vitality",
        "value": 56,
        "weight": 0.22,
        "metrics": [
          {
            "key": "development_activity",
            "band": "at_risk",
            "name": "Development activity",
            "note": null,
            "notes": [],
            "value": 42,
            "inputs": {
              "commits_last_year": 86,
              "human_commit_share": 1,
              "days_since_last_push": 127,
              "active_weeks_last_year": 21
            },
            "components": [
              {
                "key": "push_recency",
                "name": "Push recency",
                "detail": "last push 127 days ago",
                "points": 9.9,
                "status": "partial",
                "details": [
                  {
                    "code": "push_recency",
                    "params": {
                      "days": 127
                    }
                  }
                ],
                "max_points": 36
              },
              {
                "key": "commit_cadence",
                "name": "Commit cadence",
                "detail": "21/52 weeks with commits",
                "points": 14.5,
                "status": "partial",
                "details": [
                  {
                    "code": "commit_cadence_weeks",
                    "params": {
                      "weeks": 21
                    }
                  }
                ],
                "max_points": 36
              },
              {
                "key": "commit_volume",
                "name": "Commit volume",
                "detail": "86 commits in the last year",
                "points": 17.4,
                "status": "partial",
                "details": [
                  {
                    "code": "commits_last_year",
                    "params": {
                      "count": 86
                    }
                  }
                ],
                "max_points": 18
              },
              {
                "key": "openssf_scorecard_maintained",
                "name": "OpenSSF Scorecard: Maintained",
                "detail": "0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 10
              }
            ]
          },
          {
            "key": "release_discipline",
            "band": "good",
            "name": "Release discipline",
            "note": "Excluded from scoring (no data or not applicable): OpenSSF Scorecard: Signed-Releases. Remaining weights renormalized.",
            "notes": [
              {
                "code": "excluded_no_data",
                "params": {
                  "components": [
                    "openssf_scorecard_signed_releases"
                  ]
                }
              },
              {
                "code": "weights_renormalized",
                "params": {}
              }
            ],
            "value": 78,
            "inputs": {
              "releases_count": 33,
              "latest_release_tag": "v0.17.0",
              "releases_from_tags": true,
              "days_since_latest_release": 128,
              "mean_days_between_releases": 13.2
            },
            "components": [
              {
                "key": "ships_releases",
                "name": "Ships releases",
                "detail": "33 version tags (no GitHub releases)",
                "points": 16.2,
                "status": "partial",
                "details": [
                  {
                    "code": "version_tags_no_releases",
                    "params": {
                      "count": 33
                    }
                  }
                ],
                "max_points": 27
              },
              {
                "key": "release_recency",
                "name": "Release recency",
                "detail": "latest release 128 days ago",
                "points": 27,
                "status": "partial",
                "details": [
                  {
                    "code": "release_recency",
                    "params": {
                      "days": 128
                    }
                  }
                ],
                "max_points": 36
              },
              {
                "key": "release_cadence",
                "name": "Release cadence",
                "detail": "a release every ~13.2 days",
                "points": 27,
                "status": "met",
                "details": [
                  {
                    "code": "release_cadence",
                    "params": {
                      "gap": 13.2
                    }
                  }
                ],
                "max_points": 27
              },
              {
                "key": "openssf_scorecard_signed_releases",
                "name": "OpenSSF Scorecard: Signed-Releases",
                "detail": "no releases found",
                "points": 0,
                "status": "excluded",
                "details": [
                  {
                    "code": "no_data",
                    "params": {}
                  }
                ],
                "max_points": 10
              }
            ]
          },
          {
            "key": "abandonment",
            "band": "excellent",
            "name": "Abandonment",
            "note": null,
            "notes": [],
            "value": 100,
            "inputs": {
              "cap": null,
              "state": "unverified",
              "guards": [],
              "signals": [],
              "red_flag": false,
              "multiplier_pct": 100,
              "declared_reason": null,
              "unverified_reason": "repository_too_young",
              "unanswered_open_prs": null,
              "unanswered_open_issues": null,
              "days_since_last_merged_pr": null,
              "days_since_last_human_commit": null,
              "days_since_last_human_commit_is_floor": false
            },
            "components": [
              {
                "key": "project_is_still_maintained",
                "name": "Project is still maintained",
                "detail": "maintenance record not established from the collected data",
                "points": 100,
                "status": "met",
                "details": [
                  {
                    "code": "abandonment_unverified",
                    "params": {}
                  }
                ],
                "max_points": 100
              }
            ]
          }
        ],
        "description": "Is the project alive — is code being written and are releases shipping?"
      },
      {
        "key": "community",
        "band": "critical",
        "name": "Community & Adoption",
        "value": 12,
        "weight": 0.18,
        "metrics": [
          {
            "key": "popularity",
            "band": "critical",
            "name": "Popularity & adoption",
            "note": null,
            "notes": [],
            "value": 1,
            "inputs": {
              "forks": 0,
              "stars": 0,
              "watchers": 0,
              "growth_state": "unverified",
              "growth_factor_pct": 100,
              "growth_unverified_reason": "no_history"
            },
            "components": [
              {
                "key": "stars",
                "name": "Stars",
                "detail": "0 stars",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "stars",
                    "params": {
                      "count": 0
                    }
                  }
                ],
                "max_points": 60
              },
              {
                "key": "forks",
                "name": "Forks",
                "detail": "0 forks",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "forks",
                    "params": {
                      "count": 0
                    }
                  }
                ],
                "max_points": 25
              },
              {
                "key": "watchers",
                "name": "Watchers",
                "detail": "0 watchers",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "watchers",
                    "params": {
                      "count": 0
                    }
                  }
                ],
                "max_points": 15
              }
            ]
          },
          {
            "key": "community_health",
            "band": "critical",
            "name": "Community health",
            "note": null,
            "notes": [],
            "value": 25,
            "inputs": {
              "has_readme": false,
              "has_license": false,
              "has_contributing": false,
              "has_issue_template": false,
              "has_code_of_conduct": false,
              "has_pull_request_template": false
            },
            "components": [
              {
                "key": "readme",
                "name": "README",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 22.5
              },
              {
                "key": "license",
                "name": "License",
                "detail": "recognized license (MPL-2.0)",
                "points": 22.5,
                "status": "met",
                "details": [
                  {
                    "code": "license_standard",
                    "params": {}
                  },
                  {
                    "code": "license_spdx",
                    "params": {
                      "spdx": "MPL-2.0"
                    }
                  }
                ],
                "max_points": 22.5
              },
              {
                "key": "contributing_guide",
                "name": "CONTRIBUTING guide",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 18
              },
              {
                "key": "code_of_conduct",
                "name": "Code of conduct",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 13.5
              },
              {
                "key": "issue_template",
                "name": "Issue template",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.2
              },
              {
                "key": "pr_template",
                "name": "PR template",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 6.3
              }
            ]
          }
        ],
        "description": "Does the project have users, downloads, attention, and a welcoming setup for contributors?"
      },
      {
        "key": "governance",
        "band": "moderate",
        "name": "Sustainability & Governance",
        "value": 57,
        "weight": 0.24,
        "metrics": [
          {
            "key": "maintainer_resilience",
            "band": "critical",
            "name": "Maintainer resilience (bus factor)",
            "note": null,
            "notes": [],
            "value": 25,
            "inputs": {
              "bus_factor": 1,
              "contributors_sampled": 2,
              "top_contributor_share": 0.681
            },
            "components": [
              {
                "key": "bus_factor",
                "name": "Bus factor",
                "detail": "1 contributor(s) cover half of all commits",
                "points": 9,
                "status": "partial",
                "details": [
                  {
                    "code": "bus_factor",
                    "params": {
                      "count": 1
                    }
                  }
                ],
                "max_points": 54
              },
              {
                "key": "commit_distribution",
                "name": "Commit distribution",
                "detail": "top contributor authored 68% of commits",
                "points": 7.2,
                "status": "partial",
                "details": [
                  {
                    "code": "top_contributor_share",
                    "params": {
                      "share": 68
                    }
                  }
                ],
                "max_points": 22.5
              },
              {
                "key": "contributor_breadth",
                "name": "Contributor breadth",
                "detail": "2 contributors",
                "points": 2.7,
                "status": "partial",
                "details": [
                  {
                    "code": "contributors_sampled",
                    "params": {
                      "count": 2
                    }
                  }
                ],
                "max_points": 13.5
              },
              {
                "key": "openssf_scorecard_contributors",
                "name": "OpenSSF Scorecard: Contributors",
                "detail": "project has 2 contributing companies or organizations -- score normalized to 6",
                "points": 6,
                "status": "partial",
                "details": [],
                "max_points": 10
              }
            ]
          },
          {
            "key": "responsiveness",
            "band": "good",
            "name": "Issue & PR responsiveness",
            "note": "Excluded from scoring (no data or not applicable): Issue resolution. Remaining weights renormalized.",
            "notes": [
              {
                "code": "excluded_no_data",
                "params": {
                  "components": [
                    "issue_resolution"
                  ]
                }
              },
              {
                "code": "weights_renormalized",
                "params": {}
              }
            ],
            "value": 72,
            "inputs": {
              "merged_prs": 1,
              "open_issues": 0,
              "closed_issues": 0,
              "issue_closed_ratio": null,
              "closed_unmerged_prs": 0
            },
            "components": [
              {
                "key": "issue_resolution",
                "name": "Issue resolution",
                "detail": "no issues or no data",
                "points": 0,
                "status": "excluded",
                "details": [
                  {
                    "code": "no_issues_or_data",
                    "params": {}
                  }
                ],
                "max_points": 46.75
              },
              {
                "key": "pr_acceptance",
                "name": "PR acceptance",
                "detail": "1/1 decided PRs merged",
                "points": 38.2,
                "status": "met",
                "details": [
                  {
                    "code": "decided_prs_merged",
                    "params": {
                      "merged": 1,
                      "decided": 1
                    }
                  }
                ],
                "max_points": 38.25
              },
              {
                "key": "openssf_scorecard_code_review",
                "name": "OpenSSF Scorecard: Code-Review",
                "detail": "Found 0/26 approved changesets -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 15
              }
            ]
          },
          {
            "key": "stewardship",
            "band": "at_risk",
            "name": "Ownership & stewardship",
            "note": null,
            "notes": [],
            "value": 47,
            "inputs": {
              "followers": 3,
              "owner_type": "Organization",
              "is_verified": null,
              "owner_login": "CircleCI-Research",
              "public_repos": 4,
              "account_age_days": 1387
            },
            "components": [
              {
                "key": "ownership_backing",
                "name": "Ownership backing",
                "detail": "organization-owned",
                "points": 30,
                "status": "met",
                "details": [
                  {
                    "code": "owner_organization",
                    "params": {}
                  }
                ],
                "max_points": 30
              },
              {
                "key": "verified_domain",
                "name": "Verified domain",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 20
              },
              {
                "key": "owner_reach",
                "name": "Owner reach",
                "detail": "3 followers of CircleCI-Research",
                "points": 4.3,
                "status": "partial",
                "details": [
                  {
                    "code": "owner_followers",
                    "params": {
                      "count": 3,
                      "login": "CircleCI-Research"
                    }
                  }
                ],
                "max_points": 25
              },
              {
                "key": "track_record",
                "name": "Track record",
                "detail": "4 public repos, account ~3 yr old",
                "points": 12.7,
                "status": "partial",
                "details": [
                  {
                    "code": "public_repos",
                    "params": {
                      "count": 4
                    }
                  },
                  {
                    "code": "account_age_years",
                    "params": {
                      "years": 3
                    }
                  }
                ],
                "max_points": 25
              }
            ]
          },
          {
            "key": "package_maintenance",
            "band": "excellent",
            "name": "Package maintenance",
            "note": null,
            "notes": [],
            "value": 100,
            "inputs": {
              "packages": [
                "github.com/CircleCI-Research/evalbench"
              ],
              "ecosystems": "go",
              "any_deprecated": false,
              "min_days_since_publish": 128
            },
            "components": [
              {
                "key": "published_resolvable",
                "name": "Published & resolvable",
                "detail": "1 package(s) on go",
                "points": 25,
                "status": "met",
                "details": [
                  {
                    "code": "packages_published",
                    "params": {
                      "count": 1,
                      "ecosystems": "go"
                    }
                  }
                ],
                "max_points": 25
              },
              {
                "key": "publish_recency",
                "name": "Publish recency",
                "detail": "latest publish 128 days ago",
                "points": 35,
                "status": "met",
                "details": [
                  {
                    "code": "publish_recency",
                    "params": {
                      "days": 128
                    }
                  }
                ],
                "max_points": 35
              },
              {
                "key": "version_history",
                "name": "Version history",
                "detail": "33 published versions",
                "points": 20,
                "status": "met",
                "details": [
                  {
                    "code": "published_versions",
                    "params": {
                      "count": 33
                    }
                  }
                ],
                "max_points": 20
              },
              {
                "key": "not_deprecated",
                "name": "Not deprecated",
                "detail": "active, not deprecated or yanked",
                "points": 20,
                "status": "met",
                "details": [
                  {
                    "code": "package_not_deprecated",
                    "params": {}
                  }
                ],
                "max_points": 20
              }
            ]
          }
        ],
        "description": "Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep?"
      },
      {
        "key": "engineering",
        "band": "moderate",
        "name": "Engineering Quality",
        "value": 51,
        "weight": 0.2,
        "metrics": [
          {
            "key": "engineering_practices",
            "band": "moderate",
            "name": "Engineering practices",
            "note": null,
            "notes": [],
            "value": 68,
            "inputs": {
              "has_ci": true,
              "has_tests": true,
              "has_editorconfig": false,
              "has_linter_config": false,
              "has_precommit_config": false
            },
            "components": [
              {
                "key": "ci_workflows",
                "name": "CI workflows",
                "detail": "1 workflow(s)",
                "points": 24,
                "status": "met",
                "details": [
                  {
                    "code": "ci_workflows",
                    "params": {
                      "count": 1
                    }
                  }
                ],
                "max_points": 24
              },
              {
                "key": "tests_present",
                "name": "Tests present",
                "detail": null,
                "points": 24,
                "status": "met",
                "details": [],
                "max_points": 24
              },
              {
                "key": "linter_config",
                "name": "Linter config",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 16
              },
              {
                "key": "pre_commit_hooks",
                "name": "Pre-commit hooks",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 9.6
              },
              {
                "key": "editorconfig",
                "name": ".editorconfig",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 6.4
              },
              {
                "key": "openssf_scorecard_ci_tests",
                "name": "OpenSSF Scorecard: CI-Tests",
                "detail": "1 out of 1 merged PRs checked by a CI test -- score normalized to 10",
                "points": 20,
                "status": "met",
                "details": [],
                "max_points": 20
              }
            ]
          },
          {
            "key": "documentation",
            "band": "critical",
            "name": "Documentation",
            "note": null,
            "notes": [],
            "value": 25,
            "inputs": {
              "topics": [
                "ai",
                "ai-agents",
                "benchmark-framework",
                "ci-cd",
                "circleci",
                "devops",
                "evaluation-framework",
                "evaluation-metrics",
                "llm",
                "model-comparison"
              ],
              "has_wiki": false,
              "homepage": "https://loop.circleci.com",
              "has_readme": false,
              "has_docs_dir": false,
              "has_description": false
            },
            "components": [
              {
                "key": "readme",
                "name": "README",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 30
              },
              {
                "key": "documentation_directory",
                "name": "Documentation directory",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 25
              },
              {
                "key": "documentation_homepage_site",
                "name": "Documentation / homepage site",
                "detail": "https://loop.circleci.com",
                "points": 15,
                "status": "met",
                "details": [],
                "max_points": 15
              },
              {
                "key": "repository_description",
                "name": "Repository description",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 10
              },
              {
                "key": "topics",
                "name": "Topics",
                "detail": "10 topics",
                "points": 10,
                "status": "met",
                "details": [
                  {
                    "code": "topics_count",
                    "params": {
                      "count": 10
                    }
                  }
                ],
                "max_points": 10
              },
              {
                "key": "wiki",
                "name": "Wiki",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 10
              }
            ]
          }
        ],
        "description": "Are baseline engineering and documentation practices in place?"
      },
      {
        "key": "security",
        "band": "at_risk",
        "name": "Security",
        "value": 30,
        "weight": 0.16,
        "metrics": [
          {
            "key": "security_posture",
            "band": "at_risk",
            "name": "Security posture",
            "note": "Excluded from scoring (no data or not applicable): Packaging, Signed-Releases. Remaining weights renormalized.",
            "notes": [
              {
                "code": "excluded_no_data",
                "params": {
                  "components": [
                    "packaging",
                    "signed_releases"
                  ]
                }
              },
              {
                "code": "weights_renormalized",
                "params": {}
              }
            ],
            "value": 30,
            "inputs": {
              "source": "openssf_scorecard",
              "checks_evaluated": 16,
              "scorecard_version": "v5.5.0",
              "checks_inconclusive": 2,
              "scorecard_aggregate": 3
            },
            "components": [
              {
                "key": "binary_artifacts",
                "name": "Binary-Artifacts",
                "detail": "no binaries found in the repo",
                "points": 7.5,
                "status": "met",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "branch_protection",
                "name": "Branch-Protection",
                "detail": "branch protection is not maximal on development and all release branches",
                "points": 3.8,
                "status": "partial",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "ci_tests",
                "name": "CI-Tests",
                "detail": "1 out of 1 merged PRs checked by a CI test -- score normalized to 10",
                "points": 2.5,
                "status": "met",
                "details": [],
                "max_points": 2.5
              },
              {
                "key": "cii_best_practices",
                "name": "CII-Best-Practices",
                "detail": "no effort to earn an OpenSSF best practices badge detected",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 2.5
              },
              {
                "key": "code_review",
                "name": "Code-Review",
                "detail": "Found 0/26 approved changesets -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "contributors",
                "name": "Contributors",
                "detail": "project has 2 contributing companies or organizations -- score normalized to 6",
                "points": 1.5,
                "status": "partial",
                "details": [],
                "max_points": 2.5
              },
              {
                "key": "dangerous_workflow",
                "name": "Dangerous-Workflow",
                "detail": "no dangerous workflow patterns detected",
                "points": 10,
                "status": "met",
                "details": [],
                "max_points": 10
              },
              {
                "key": "dependency_update_tool",
                "name": "Dependency-Update-Tool",
                "detail": "no update tool detected",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "fuzzing",
                "name": "Fuzzing",
                "detail": "project is not fuzzed",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 5
              },
              {
                "key": "license",
                "name": "License",
                "detail": "license file detected",
                "points": 2.5,
                "status": "met",
                "details": [],
                "max_points": 2.5
              },
              {
                "key": "maintained",
                "name": "Maintained",
                "detail": "0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "packaging",
                "name": "Packaging",
                "detail": "packaging workflow not detected",
                "points": 0,
                "status": "excluded",
                "details": [
                  {
                    "code": "no_data",
                    "params": {}
                  }
                ],
                "max_points": 5
              },
              {
                "key": "pinned_dependencies",
                "name": "Pinned-Dependencies",
                "detail": "dependency not pinned by hash detected -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 5
              },
              {
                "key": "sast",
                "name": "SAST",
                "detail": "SAST tool is not run on all commits -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 5
              },
              {
                "key": "security_policy",
                "name": "Security-Policy",
                "detail": "security policy file not detected",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 5
              },
              {
                "key": "signed_releases",
                "name": "Signed-Releases",
                "detail": "no releases found",
                "points": 0,
                "status": "excluded",
                "details": [
                  {
                    "code": "no_data",
                    "params": {}
                  }
                ],
                "max_points": 7.5
              },
              {
                "key": "token_permissions",
                "name": "Token-Permissions",
                "detail": "detected GitHub workflow tokens with excessive permissions",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              },
              {
                "key": "vulnerabilities",
                "name": "Vulnerabilities",
                "detail": "44 existing vulnerabilities detected",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 7.5
              }
            ]
          },
          {
            "key": "high_risk_jurisdiction_exposure",
            "band": "excellent",
            "name": "High-Risk Jurisdiction Exposure",
            "note": "Only high-confidence self-published location evidence affects this multiplier. Ambiguous matches are review-only; country evidence is not proof of nationality, citizenship, legal registration, malicious intent, or sanctions status.",
            "notes": [
              {
                "code": "jurisdiction_evidence_limits",
                "params": {}
              }
            ],
            "value": 100,
            "inputs": {
              "meaning": "self-published location evidence; not nationality or citizenship",
              "red_flag": false,
              "exposures": [],
              "policy_countries": [
                "Russia",
                "Iran",
                "North Korea"
              ],
              "review_only_matches": 0,
              "assessed_self_published_locations": 4
            },
            "components": [
              {
                "key": "policy_exposure_multiplier",
                "name": "Policy exposure multiplier",
                "detail": "no confirmed policy-scope location match",
                "points": 100,
                "status": "met",
                "details": [
                  {
                    "code": "jurisdiction_no_match",
                    "params": {}
                  }
                ],
                "max_points": 100
              }
            ]
          }
        ],
        "description": "Are visible security and supply-chain practices strong, with no malicious dependency and no unresolved high-risk jurisdiction exposure?"
      },
      {
        "key": "ai_readiness",
        "band": "moderate",
        "name": "AI Readiness",
        "value": 65,
        "weight": 0,
        "metrics": [
          {
            "key": "ai_agent_context",
            "band": "excellent",
            "name": "Agent context & guidance",
            "note": null,
            "notes": [],
            "value": 85,
            "inputs": {
              "has_llms_txt": false,
              "legible_history_share": 0.96,
              "agent_instruction_files": [
                "AGENTS.md"
              ],
              "agent_instruction_max_bytes": 3085
            },
            "components": [
              {
                "key": "agent_instructions",
                "name": "Agent instructions",
                "detail": "AGENTS.md",
                "points": 45,
                "status": "met",
                "details": [
                  {
                    "code": "file_list",
                    "params": {
                      "files": "AGENTS.md"
                    }
                  }
                ],
                "max_points": 45
              },
              {
                "key": "machine_readable_docs_llms_txt",
                "name": "Machine-readable docs (llms.txt)",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 15
              },
              {
                "key": "legible_commit_history",
                "name": "Legible commit history",
                "detail": "96 of 100 human commits state their intent (structured subject or explanatory body)",
                "points": 40,
                "status": "met",
                "details": [
                  {
                    "code": "legible_history",
                    "params": {
                      "legible": 96,
                      "sampled": 100
                    }
                  }
                ],
                "max_points": 40
              }
            ]
          },
          {
            "key": "ai_verify_loop",
            "band": "moderate",
            "name": "Verify loop (build / test / typecheck)",
            "note": null,
            "notes": [],
            "value": 55,
            "inputs": {
              "has_nix": false,
              "has_tests": true,
              "lockfiles": [
                "go.sum"
              ],
              "has_dockerfile": false,
              "typed_language": false,
              "bootstrap_files": [],
              "has_devcontainer": false,
              "has_linter_config": false,
              "typecheck_configs": [],
              "agent_commit_share": 0.11,
              "toolchain_manifests": [
                "go.mod"
              ],
              "dependency_bot_commit_share": 0
            },
            "components": [
              {
                "key": "one_command_bootstrap",
                "name": "One-command bootstrap",
                "detail": "go.mod (toolchain convention, no task runner)",
                "points": 12.6,
                "status": "partial",
                "details": [
                  {
                    "code": "toolchain_convention",
                    "params": {
                      "files": "go.mod"
                    }
                  }
                ],
                "max_points": 18
              },
              {
                "key": "automated_tests",
                "name": "Automated tests",
                "detail": null,
                "points": 22,
                "status": "met",
                "details": [],
                "max_points": 22
              },
              {
                "key": "lint_format_config",
                "name": "Lint / format config",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 11
              },
              {
                "key": "static_type_checking",
                "name": "Static type checking",
                "detail": null,
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 11
              },
              {
                "key": "reproducible_environment",
                "name": "Reproducible environment",
                "detail": "lockfile",
                "points": 10,
                "status": "met",
                "details": [
                  {
                    "code": "file_list",
                    "params": {
                      "files": "lockfile"
                    }
                  }
                ],
                "max_points": 10
              },
              {
                "key": "demonstrated_agent_practice",
                "name": "Demonstrated agent practice",
                "detail": "11 of the last 100 commits agent-authored or agent-credited",
                "points": 10,
                "status": "met",
                "details": [
                  {
                    "code": "agent_authored_commits",
                    "params": {
                      "count": 11,
                      "sampled": 100
                    }
                  }
                ],
                "max_points": 10
              },
              {
                "key": "automated_maintenance",
                "name": "Automated maintenance",
                "detail": "no automated dependency updates observed",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "no_dependency_automation",
                    "params": {}
                  }
                ],
                "max_points": 8
              },
              {
                "key": "openssf_scorecard_pinned_dependencies",
                "name": "OpenSSF Scorecard: Pinned-Dependencies",
                "detail": "dependency not pinned by hash detected -- score normalized to 0",
                "points": 0,
                "status": "missed",
                "details": [],
                "max_points": 10
              }
            ]
          },
          {
            "key": "ai_code_legibility",
            "band": "moderate",
            "name": "Code legibility for models",
            "note": null,
            "notes": [],
            "value": 54,
            "inputs": {
              "primary_language": "HTML",
              "largest_source_bytes": 73199,
              "source_files_sampled": 442,
              "oversized_source_files": 5
            },
            "components": [
              {
                "key": "type_checkable_code",
                "name": "Type-checkable code",
                "detail": "HTML without a type-check config",
                "points": 0,
                "status": "missed",
                "details": [
                  {
                    "code": "no_typecheck_config_language",
                    "params": {
                      "language": "HTML"
                    }
                  }
                ],
                "max_points": 45
              },
              {
                "key": "manageable_file_sizes",
                "name": "Manageable file sizes",
                "detail": "5/442 source files over 60KB",
                "points": 54.4,
                "status": "partial",
                "details": [
                  {
                    "code": "oversized_source_files",
                    "params": {
                      "kb": 60,
                      "sampled": 442,
                      "oversized": 5
                    }
                  }
                ],
                "max_points": 55
              }
            ]
          }
        ],
        "description": "How well is the repo equipped to be developed and maintained with AI coding agents? An independent, experimental badge — weight 0.0, so it is surfaced on its own and does not affect the overall health score."
      }
    ],
    "metrics_version": "1.13.0"
  },
  "warnings": [
    "Community profile unavailable",
    "GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository"
  ],
  "report_type": "repository",
  "generated_at": "2026-07-25T08:58:13.715545Z",
  "schema_version": "0.27.0",
  "badge_url": "https://raw.githubusercontent.com/inspect-software/badges/main/v1/c/CircleCI-Research/evalbench.svg",
  "full_name": "CircleCI-Research/evalbench",
  "license_state": "standard",
  "license_spdx": "MPL-2.0"
}

评分是信号,而非担保。 评分反映的是 GitHub 上公开可见的实践——不是代码审计,也不是安全保证。

缺失数据将被剔除并重新归一化权重,绝不按零分计。方法论已版本化并公开:指标 v1.13.0、模式 v0.27.0—— 完整方法论 · 指标知识库.

单项结果在整体记录中的位置: 汇总统计Go.