All tags
Catalogue tag

#benchmarks

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

15 records
Tagged “benchmarks”Ranked by health index
PyPI · npm
93Exceptionalhealth index
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1,055↓ 406.4K/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI · npm
93Exceptionalhealth index
trycua/cua
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
HTML · Rust · Python★ 20.9K↓ 26.8K/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
crates.io · PyPI
89Excellenthealth index
constructorfabric/gears-rust
All-in-one open-source framework & middleware for enterprise-grade multi-tenant and multi-tier XaaS Services development
Rust★ 25↓ 13.5K/moAug 1, 2026
Apache-2.0Aug 1, 2026 · metrics 2.10.0
Go
80Excellenthealth index
oneclickvirt/ecs
VPS Fusion Monster Server Test GO Version Aiming to be the most comprehensive server testing project, implemented in Go with zero environment dependencies. VPS融合怪服务器测评项目 GO版本 尽量成为最全能的服务器测评项目,使用 Go 实现,无需任何环境依赖。
Go★ 2,264Jul 29, 2026
GPL-3.0Jul 29, 2026 · metrics 2.10.0
crates.io · npm
78Goodhealth index
reyamira/models
TUI and CLI for browsing AI models, benchmarks, coding agents, and statuses for AI providers.
Rust★ 483↓ 174/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
Maven
77Goodhealth index
JCTools/JCTools
No repository description published.
Java★ 3,858Jul 30, 2026
Apache-2.0Jul 30, 2026 · metrics 2.10.0
npm
77Goodhealth index
pmndrs/detect-gpu
Classifies GPUs based on their 3D rendering benchmark score allowing the developer to provide sensible default settings for graphically intensive applications.
TypeScript★ 1,211↓ 74.4K/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI
69Goodhealth index
robocurve/worldevals
A curated catalog of VLA / physical-AI benchmarks, each runnable on real robots or sims via Inspect Robots. (The Inspect Evals for robotics.)
Python★ 6Jul 30, 2026
MITJul 30, 2026 · metrics 2.10.0
PyPI
67Goodhealth index
OpenAdaptAI/openadapt-evals
Evaluation infrastructure for GUI agent benchmarks
Python★ 2↓ 2,941/moJul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
npm · PyPI
65Goodhealth index
evo-hq/evo
turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.
Python · JavaScript★ 1,348↓ 10.5K/moJul 23, 2026
Apache-2.0Jul 23, 2026 · metrics 2.10.0
63Moderatehealth index
testillano/h2agent
C++ HTTP/2 Mock Service which enables mocking HTTP/2 applications (also HTTP/1 supported).
C++ · Python · Shell★ 12Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
PyPI
56Moderatehealth index
a-r-j/ProteinWorkshop
Benchmarking framework for protein representation learning. Includes a large number of pre-training and downstream task datasets, models and training/task utilities. (ICLR 2024)
Python · Jupyter Notebook★ 275Jul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
crates.io
51Moderatehealth index
beling/bsuccinct-rs
Rust libraries and programs focused on succinct data structures
Rust★ 171↓ 343.7K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
Packagist
32At Riskhealth index
spiral/debug
[READ ONLY] Debug Toolkit. Subtree split of the Spiral Debug component (see spiral/framework)
PHP★ 2↓ 3,469/moAug 13, 2026
MITAug 13, 2026 · metrics 2.10.0