All tags
Catalogue tag

#benchmark

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

73 records
Tagged “benchmark”Ranked by health index
PyPI
95Exceptionalhealth index
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 3,362Jul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 2.10.0
npm
95Exceptionalhealth index
tinylibs/tinybench
🔎 A simple, tiny and lightweight benchmarking library!
TypeScript★ 2,354↓ 296M/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
NuGet
94Exceptionalhealth index
dotnet/BenchmarkDotNet
Powerful .NET library for benchmarking
C#★ 11.5KAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
npm
93Exceptionalhealth index
artilleryio/artillery
The complete load testing platform. Everything you need for production-grade load tests. Serverless & distributed. Load test with Playwright. Load test HTTP APIs, GraphQL, WebSocket, and more. Use any Node.js module.
TypeScript · JavaScript★ 9,059↓ 1.1M/moAug 28, 2026
MPL-2.0Aug 28, 2026 · metrics 2.10.0
npm · PyPI
91Excellenthealth index
joshuaswarren/remnic
Open-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.
TypeScript★ 176↓ 203.7K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
90Excellenthealth index
airspeed-velocity/asv
Airspeed Velocity: A simple Python benchmarking tool with web-based reporting
Python · JavaScript★ 1,008↓ 508.4K/moJul 21, 2026
BSD-3-ClauseJul 21, 2026 · metrics 2.10.0
PyPI · npm
89Excellenthealth index
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
fitbenchmarking/fitbenchmarking
Tool for comparing the run time and accuracy of minimizers on fit benchmarking problems
Python★ 16Jul 25, 2026
BSD-3-ClauseJul 25, 2026 · metrics 2.10.0
Go
89Excellenthealth index
gofiber/utils
:zap: A collection of common functions for Fiber with better performance, fewer allocations, and fewer dependencies.
Go★ 57Aug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
crates.io
89Excellenthealth index
gungraun/gungraun
High-precision, one-shot and consistent benchmarking framework/harness for Rust. All Valgrind tools at your fingertips.
Rust · Roff★ 317↓ 412.5K/moAug 22, 2026
Apache-2.0Aug 22, 2026 · metrics 2.10.0
PyPI
87Excellenthealth index
SWE-bench/SWE-bench
SWE-bench: Can Language Models Resolve Real-world Github Issues?
Python★ 5,726↓ 46.4M/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
crates.io
87Excellenthealth index
pawurb/hotpath-rs
Quickly find bottlenecks in Rust - one profiler for CPU, memory, SQL, HTTP, I/O and async code.
Rust★ 1,688↓ 174.3K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
MichaelGrupp/evo
Python package for the evaluation of odometry and SLAM
Python★ 4,309↓ 210.7K/moAug 28, 2026
GPL-3.0Aug 28, 2026 · metrics 2.10.0
crates.io · Maven · npm
86Excellenthealth index
SaaSy-Solutions/mockforge
Comprehensive mocking framework for REST, gRPC, GraphQL & WebSockets. Features intelligent RAG-driven data synthesis, latency/fault injection, HTTP bridge, WASM plugins, E2E encryption, workspace sync, and modern admin UI. Production-ready with Docker support.
Rust · TypeScript · HTML★ 11↓ 9,833/moSep 5, 2026
Custom licenseSep 5, 2026 · metrics 2.10.0
PyPI · npm
86Excellenthealth index
robocurve/inspect-robots
Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.
Python★ 269↓ 5,131/moSep 6, 2026
MITSep 6, 2026 · metrics 2.10.0
PyPI
84Excellenthealth index
CodSpeedHQ/pytest-codspeed
A pytest plugin to create benchmarks
Python★ 134↓ 2.4M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
npm
84Excellenthealth index
vava-nessa/free-coding-models
Find, benchmark and install in CLI 170+ FREE coding LLM models across 15+ providers in real time
HTML · JavaScript★ 2,186↓ 5,344/moJul 23, 2026
Custom licenseJul 23, 2026 · metrics 2.10.0
PyPI
83Excellenthealth index
SWE-bench/SWE-smith
[NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents
Python★ 752↓ 16.8M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
crates.io
81Excellenthealth index
hatoo/oha
Ohayou(おはよう), HTTP load generator, inspired by rakyll/hey with tui animation.
Rust★ 10.5K↓ 1,822/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
crates.io
81Excellenthealth index
lance0/xfr
A modern iperf3 alternative with a live TUI, multi-client server, and QUIC support. Built in Rust.
Rust★ 517↓ 179/moJul 23, 2026
Apache-2.0Jul 23, 2026 · metrics 2.10.0
crates.io
80Excellenthealth index
bencherdev/bencher
🐰 Bencher - Continuous Benchmarking
Rust · MDX★ 874↓ 11/moJul 27, 2026
Custom licenseJul 27, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
benchflow-ai/benchflow
Research infra for creating RL environments, post-training, and evals.
Python★ 317↓ 6,133/moAug 9, 2026
Apache-2.0Aug 9, 2026 · metrics 2.10.0
npm · Go
80Excellenthealth index
goptics/vizb
A tabular data visualization engine from your local to CI/CD pipeline
Go · TypeScript★ 79↓ 4,090/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
Go
80Excellenthealth index
oneclickvirt/ecs
VPS Fusion Monster Server Test GO Version Aiming to be the most comprehensive server testing project, implemented in Go with zero environment dependencies. VPS融合怪服务器测评项目 GO版本 尽量成为最全能的服务器测评项目,使用 Go 实现,无需任何环境依赖。
Go★ 2,264Jul 29, 2026
GPL-3.0Jul 29, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
Sahil170595/Chimeraforge
PyPI capacity-planning CLI for LLM deployment. pip install chimeraforge.
Python★ 2↓ 2,506/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
75Goodhealth index
soenneker/soenneker.tests.benchmark
An abstract class for benchmarking tests in .NET, integrating BenchmarkDotNet with Xunit's output helper and providing a method to log benchmark summaries asynchronously.
C#★ 0Jul 16, 2026
MITJul 16, 2026 · metrics 2.10.0
73Goodhealth index
inikep/lzbench
lzbench is an in-memory benchmark of open-source compressors
C · C++★ 1,074Jul 21, 2026
Custom licenseJul 21, 2026 · metrics 2.10.0
crates.io · PyPI
73Goodhealth index
sebastienrousseau/http-handle
Lightweight Rust HTTP server library. Sync + async, HTTP/1.1 keep-alive, HTTP/2 (h2c), HTTP/3 ALPN. 180 k req/s on Linux/arm64.
Rust★ 1↓ 2,031/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
ionelmc/pytest-benchmark
pytest fixture for benchmarking code
Python★ 1,444Jul 17, 2026
BSD-2-ClauseJul 17, 2026 · metrics 2.10.0