All tags
Catalogue tag

#evaluation

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

58 records
Tagged “evaluation”Ranked by health index
npm
69Goodhealth index
Kromatic-Innovation/panelist
Synthetic user panels for any artifact, run across multiple model providers to correct for self-preference bias — models the exact point a persona quits, dismisses, or refuses to click, not a warmth score.
JavaScript★ 0↓ 3,165/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
npm
69Goodhealth index
crewhaus/factory
Open-source compiler for AI agents. Write one crewhaus.yaml; compile it to a CLI, a Slack bot, and an eval harness from the same spec. Apache-2.0.
TypeScript★ 2↓ 25.7K/moJul 24, 2026
Apache-2.0Jul 24, 2026 · metrics 2.10.0
PyPI
69Goodhealth index
robocurve/worldevals
A curated catalog of VLA / physical-AI benchmarks, each runnable on real robots or sims via Inspect Robots. (The Inspect Evals for robotics.)
Python★ 6Jul 30, 2026
MITJul 30, 2026 · metrics 2.10.0
PyPI
67Goodhealth index
OpenAdaptAI/openadapt-evals
Evaluation infrastructure for GUI agent benchmarks
Python★ 2↓ 2,941/moJul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
npm
63Moderatehealth index
axl-sdk/axl
TypeScript SDK for orchestrating Agentic Systems — concurrency, structured output, cost control, and consensus as first-class primitives.
TypeScript★ 2↓ 3,214/moJul 27, 2026
Apache-2.0Jul 27, 2026 · metrics 2.10.0
PyPI
63Moderatehealth index
davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge
HTML · Python★ 69Sep 5, 2026
No licenseSep 5, 2026 · metrics 2.10.0
npm
63Moderatehealth index
mgechev/skillgrade
"Unit tests" for your agent skills
TypeScript★ 661↓ 2,290/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI
62Moderatehealth index
AgentX-ai/AgentX-Python
AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
Python★ 68Aug 3, 2026
MITAug 3, 2026 · metrics 2.10.0
npm
62Moderatehealth index
TypeScript · JavaScript★ 1↓ 3,920/moAug 8, 2026
No licenseAug 8, 2026 · metrics 2.10.0
PyPI · crates.io
62Moderatehealth index
nickderobertis/onejudge
A simulated interaction and evaluation loop over oneharness: drive a harness through a multi-turn conversation and score the transcript.
Rust★ 0↓ 6,844/moAug 29, 2026
MITAug 29, 2026 · metrics 2.10.0
npm
60Moderatehealth index
CarlosNZ/fig-tree-evaluator
A highly configurable custom expression tree evaluator
TypeScript★ 24↓ 3,662/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
Go · npm
60Moderatehealth index
starkSV/windows-iso-downloader
Download official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript · Go★ 14Aug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
60Moderatehealth index
toshas/torch-fidelity
High-fidelity performance metrics for generative models in PyTorch
Python★ 1,197↓ 846.4K/moAug 13, 2026
Custom licenseAug 13, 2026 · metrics 2.10.0
PyPI
57Moderatehealth index
EpsilabAI/epsilab-python
The official Python library for the Epsilab API
Python★ 0↓ 2,706/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
PyPI
56Moderatehealth index
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 607Jul 21, 2026
Custom licenseJul 21, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
danaug23/harness-arena
Your model, many harnesses, many benchmarks.
Python · HTML★ 2↓ 2,988/moAug 23, 2026
Apache-2.0Aug 23, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
lazily-hub/lazily-py
A Python library for lazy evaluation with context caching.
Python★ 1Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
Go
54Moderatehealth index
tamnd/taocp-solver
A Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 0Jul 28, 2026
MITJul 28, 2026 · metrics 2.10.0
npm
51Moderatehealth index
HolocronLab/botruntime-packages
botruntime public packages. Consumed by the botruntime platform.
TypeScript★ 0↓ 48.8K/moJul 29, 2026
No licenseJul 29, 2026 · metrics 2.10.0
PyPI
50Moderatehealth index
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2,908/moJul 22, 2026
MITJul 22, 2026 · metrics 2.10.0
PyPI
44Weakhealth index
huggingface/Math-Verify
No repository description published.
Python★ 1,176Aug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
npm
34At Riskhealth index
darks0l/modelab
Autonomous research agent SDK
TypeScript · JavaScript★ 1↓ 203/moSep 5, 2026
No licenseSep 5, 2026 · metrics 2.10.0
NuGet
34At Riskhealth index
ncalc/ncalc
NCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.
C#★ 1,142Jul 31, 2026
MITJul 31, 2026 · metrics 2.10.0
npm
34At Riskhealth index
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 353M/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
PyPI
34At Riskhealth index
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 2Jul 17, 2026
No licenseJul 17, 2026 · metrics 2.10.0
28At Riskhealth index
DanceNitra/ramr
RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)
Python★ 0Jul 29, 2026
MITJul 29, 2026 · metrics 2.10.0
PyPI
25At Riskhealth index
evfro/polara
Recommender system and evaluation framework for top-n recommendations tasks that respects polarity of feedbacks. Fast, flexible and easy to use. Written in python, boosted by scientific python stack.
Python★ 256Jul 31, 2026
MITJul 31, 2026 · metrics 2.10.0
PyPI · npm
19Criticalhealth index
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0