All tags
Catalogue tag

#benchmark

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

73 records
Tagged “benchmark”Ranked by health index
npm
71Goodhealth index
itayinbarr/little-coder
A harness optimized to smaller LLMs
TypeScript · Python · JavaScript★ 1,830↓ 3,860/moJul 23, 2026
Apache-2.0Jul 23, 2026 · metrics 2.10.0
npm
69Goodhealth index
KryptSec/oasis
Open-source AI security benchmarking CLI. Measure how AI models perform offensive security tasks with MITRE ATT&CK analysis and KSM scoring.
TypeScript★ 29↓ 39/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
crates.io
69Goodhealth index
dfinity/canbench
A benchmarking framework for canisters on the Internet Computer.
Rust★ 21↓ 16.8K/moAug 9, 2026
Apache-2.0Aug 9, 2026 · metrics 2.10.0
PyPI
67Goodhealth index
OpenAdaptAI/openadapt-evals
Evaluation infrastructure for GUI agent benchmarks
Python★ 2↓ 2,941/moJul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
Go
67Goodhealth index
RediSearch/ftsb
Full Text Search Benchmark, a tool for comparing and evaluating full-text search engines.
Python · Go★ 26Jul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
crates.io
67Goodhealth index
criterion-rs/criterion.rs
No repository description published.
Rust · HTML★ 410↓ 36.6M/moAug 8, 2026
Apache-2.0Aug 8, 2026 · metrics 2.10.0
67Goodhealth index
dadhi/FastExpressionCompiler
Fast Compiler for C# Expression Trees and the lightweight LightExpression alternative. Diagnostic and code generation tools for the expressions.
C#★ 1,370Jul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
crates.io
67Goodhealth index
imazen/zenbench
Interleaved microbenchmarking for Rust — paired statistics, CI regression testing, criterion migration
Rust★ 3↓ 6,495/moJul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
PyPI · crates.io
63Moderatehealth index
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Python★ 86↓ 6,494/moAug 20, 2026
MITAug 20, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
druide67/asiai
Multi-engine LLM benchmark & monitoring CLI for Apple Silicon
Python★ 11↓ 4,662/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
Go
63Moderatehealth index
filipecosta90/ftsb
Full Text Search Benchmark, a tool for comparing and evaluating full-text search engines.
Python · Go★ 26Jul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
google-ai-edge/LiteRT-CLI
A convenient CLI to streamline LiteRT related development workflows, including converting, quantizing, compiling, managing, running, benchmarking and visualizing LiteRT (TFLite) models on various hardwares (CPU / GPU / NPU) across platforms (desktop, mobile or cloud).
Python★ 34↓ 80/moJul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 2.10.0
npm
63Moderatehealth index
mgechev/skillgrade
"Unit tests" for your agent skills
TypeScript★ 661↓ 2,290/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
Hex
62Moderatehealth index
aryaminus/controlkeel
Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Elixir★ 10Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
crates.io
62Moderatehealth index
kornelski/dssim
Image similarity comparison simulating human perception (multiscale SSIM in Rust)
Rust★ 1,194↓ 28.6K/moSep 1, 2026
AGPL-3.0Sep 1, 2026 · metrics 2.10.0
PyPI · crates.io
59Moderatehealth index
apitap/apitap-lib
Move whole tables between databases fast — Postgres, MySQL, ClickHouse, BigQuery. Rust engine, one-line Python API, bounded memory.
Rust · Python★ 50Aug 18, 2026
MITAug 18, 2026 · metrics 2.10.0
PyPI
59Moderatehealth index
vlbthambawita/ECGBench
Reproduciable ECG Benchmark data from Open access datasets
Python★ 3↓ 2,288/moAug 6, 2026
MITAug 6, 2026 · metrics 2.10.0
npm
57Moderatehealth index
anhldh/r3f-monitor
An easy tool to monitor the performance of your R3F application.
TypeScript · CSS★ 8↓ 2,663/moJul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
crates.io
56Moderatehealth index
THeK3nger/movingai-rust
Map/Scenario parser for the MovingAI benchmark format
Roff★ 9↓ 2,286/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
56Moderatehealth index
a-r-j/ProteinWorkshop
Benchmarking framework for protein representation learning. Includes a large number of pre-training and downstream task datasets, models and training/task utilities. (ICLR 2024)
Python · Jupyter Notebook★ 275Jul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
npm
56Moderatehealth index
onury/perfy
A tiny, zero-dependency utility for measuring code execution time in high-resolution real time. Works in Node.js, browsers, Deno and Bun.
TypeScript★ 55↓ 143.7K/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
RRZE-HPC/kerncraft
Loop Kernel Analysis and Performance Modeling Toolkit
Jupyter Notebook · Python★ 99Aug 11, 2026
AGPL-3.0Aug 11, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
danaug23/harness-arena
Your model, many harnesses, many benchmarks.
Python · HTML★ 2↓ 2,988/moAug 23, 2026
Apache-2.0Aug 23, 2026 · metrics 2.10.0
npm
54Moderatehealth index
inferock/inferock-bench
Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.
TypeScript★ 123↓ 6,119/moJul 29, 2026
Custom licenseJul 29, 2026 · metrics 2.10.0
Packagist
54Moderatehealth index
infocyph/PHPForge
Shared Composer-powered QA, refactoring, benchmark, release, hook and CI tooling for PHP projects.
PHP★ 1↓ 2,564/moAug 3, 2026
MITAug 3, 2026 · metrics 2.10.0
RubyGems
54Moderatehealth index
michaelherold/benchmark-memory
Memory profiling benchmark style, for Ruby 2.1+
Ruby★ 250Jul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
RubyGems
53Moderatehealth index
pboling/gem_bench
🪑 Benchmark different versions of same or similar gems & Static Gemfile and installed gem library source code analysis
Ruby★ 94Jul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
crates.io
51Moderatehealth index
beling/bsuccinct-rs
Rust libraries and programs focused on succinct data structures
Rust★ 171↓ 343.7K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
PyPI
51Moderatehealth index
sablier-ai/finval
Rigorous validation for synthetic financial time series — 19 financial stylized-fact metrics, used as the FinBench scoring backend
Python★ 1↓ 344/moJul 31, 2026
MITJul 31, 2026 · metrics 2.10.0
PyPI
50Moderatehealth index
SidRichardsQuantum/Quantum_Backend_Bench
Run the same quantum circuit across multiple backends and compare performance, depth, gate counts, and noise effects in a single, unified workflow.
Python · Jupyter Notebook★ 1↓ 534/moAug 20, 2026
MITAug 20, 2026 · metrics 2.10.0