Усі теги
Тег каталогу

#evaluation

Усі репозиторії публічного реєстру з цим тегом — із тем GitHub або ключових слів, які публікують їхні реєстри пакетів. Здоров'я вимірюється за тією ж версіонованою методологією, що й решта реєстру.

58 записів
З тегом «evaluation»Упорядковано за індексом здоров'я
npm
69Добрийіндекс здоров'я
Kromatic-Innovation/panelist
Synthetic user panels for any artifact, run across multiple model providers to correct for self-preference bias — models the exact point a persona quits, dismisses, or refuses to click, not a warmth score.
JavaScript★ 0↓ 3 165/міс28 серп. 2026 р.
Apache-2.028 серп. 2026 р. · метрики 2.10.0
npm
69Добрийіндекс здоров'я
crewhaus/factory
Open-source compiler for AI agents. Write one crewhaus.yaml; compile it to a CLI, a Slack bot, and an eval harness from the same spec. Apache-2.0.
TypeScript★ 2↓ 25.7K/міс24 лип. 2026 р.
Apache-2.024 лип. 2026 р. · метрики 2.10.0
PyPI
69Добрийіндекс здоров'я
robocurve/worldevals
A curated catalog of VLA / physical-AI benchmarks, each runnable on real robots or sims via Inspect Robots. (The Inspect Evals for robotics.)
Python★ 630 лип. 2026 р.
MIT30 лип. 2026 р. · метрики 2.10.0
PyPI
67Добрийіндекс здоров'я
OpenAdaptAI/openadapt-evals
Evaluation infrastructure for GUI agent benchmarks
Python★ 2↓ 2 941/міс28 лип. 2026 р.
MIT28 лип. 2026 р. · метрики 2.10.0
npm
63Помірнийіндекс здоров'я
axl-sdk/axl
TypeScript SDK for orchestrating Agentic Systems — concurrency, structured output, cost control, and consensus as first-class primitives.
TypeScript★ 2↓ 3 214/міс27 лип. 2026 р.
Apache-2.027 лип. 2026 р. · метрики 2.10.0
PyPI
63Помірнийіндекс здоров'я
davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge
HTML · Python★ 695 вер. 2026 р.
Без ліцензії5 вер. 2026 р. · метрики 2.10.0
npm
63Помірнийіндекс здоров'я
mgechev/skillgrade
"Unit tests" for your agent skills
TypeScript★ 661↓ 2 290/міс5 серп. 2026 р.
MIT5 серп. 2026 р. · метрики 2.10.0
PyPI
62Помірнийіндекс здоров'я
AgentX-ai/AgentX-Python
AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
Python★ 683 серп. 2026 р.
MIT3 серп. 2026 р. · метрики 2.10.0
npm
62Помірнийіндекс здоров'я
TypeScript · JavaScript★ 1↓ 3 920/міс8 серп. 2026 р.
Без ліцензії8 серп. 2026 р. · метрики 2.10.0
PyPI · crates.io
62Помірнийіндекс здоров'я
nickderobertis/onejudge
A simulated interaction and evaluation loop over oneharness: drive a harness through a multi-turn conversation and score the transcript.
Rust★ 0↓ 6 844/міс29 серп. 2026 р.
MIT29 серп. 2026 р. · метрики 2.10.0
npm
60Помірнийіндекс здоров'я
CarlosNZ/fig-tree-evaluator
A highly configurable custom expression tree evaluator
TypeScript★ 24↓ 3 662/міс5 серп. 2026 р.
MIT5 серп. 2026 р. · метрики 2.10.0
Go · npm
60Помірнийіндекс здоров'я
starkSV/windows-iso-downloader
Download official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript · Go★ 1422 серп. 2026 р.
MIT22 серп. 2026 р. · метрики 2.10.0
PyPI
60Помірнийіндекс здоров'я
toshas/torch-fidelity
High-fidelity performance metrics for generative models in PyTorch
Python★ 1 197↓ 846.4K/міс13 серп. 2026 р.
Власна ліцензія13 серп. 2026 р. · метрики 2.10.0
PyPI
57Помірнийіндекс здоров'я
EpsilabAI/epsilab-python
The official Python library for the Epsilab API
Python★ 0↓ 2 706/міс21 лип. 2026 р.
Apache-2.021 лип. 2026 р. · метрики 2.10.0
PyPI
56Помірнийіндекс здоров'я
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 60721 лип. 2026 р.
Власна ліцензія21 лип. 2026 р. · метрики 2.10.0
PyPI
54Помірнийіндекс здоров'я
danaug23/harness-arena
Your model, many harnesses, many benchmarks.
Python · HTML★ 2↓ 2 988/міс23 серп. 2026 р.
Apache-2.023 серп. 2026 р. · метрики 2.10.0
PyPI
54Помірнийіндекс здоров'я
lazily-hub/lazily-py
A Python library for lazy evaluation with context caching.
Python★ 117 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 2.10.0
Go
54Помірнийіндекс здоров'я
tamnd/taocp-solver
A Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 028 лип. 2026 р.
MIT28 лип. 2026 р. · метрики 2.10.0
npm
51Помірнийіндекс здоров'я
HolocronLab/botruntime-packages
botruntime public packages. Consumed by the botruntime platform.
TypeScript★ 0↓ 48.8K/міс29 лип. 2026 р.
Без ліцензії29 лип. 2026 р. · метрики 2.10.0
PyPI
50Помірнийіндекс здоров'я
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2 908/міс22 лип. 2026 р.
MIT22 лип. 2026 р. · метрики 2.10.0
PyPI
44Слабкийіндекс здоров'я
huggingface/Math-Verify
Опис репозиторію не опубліковано.
Python★ 1 17613 серп. 2026 р.
Apache-2.013 серп. 2026 р. · метрики 2.10.0
npm
34У зоні ризикуіндекс здоров'я
darks0l/modelab
Autonomous research agent SDK
TypeScript · JavaScript★ 1↓ 203/міс5 вер. 2026 р.
Без ліцензії5 вер. 2026 р. · метрики 2.10.0
NuGet
34У зоні ризикуіндекс здоров'я
ncalc/ncalc
NCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.
C#★ 1 14231 лип. 2026 р.
MIT31 лип. 2026 р. · метрики 2.10.0
npm
34У зоні ризикуіндекс здоров'я
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 353M/міс4 серп. 2026 р.
MIT4 серп. 2026 р. · метрики 2.10.0
PyPI
34У зоні ризикуіндекс здоров'я
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 217 лип. 2026 р.
Без ліцензії17 лип. 2026 р. · метрики 2.10.0
28У зоні ризикуіндекс здоров'я
DanceNitra/ramr
RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)
Python★ 029 лип. 2026 р.
MIT29 лип. 2026 р. · метрики 2.10.0
PyPI
25У зоні ризикуіндекс здоров'я
evfro/polara
Recommender system and evaluation framework for top-n recommendations tasks that respects polarity of feedbacks. Fast, flexible and easy to use. Written in python, boosted by scientific python stack.
Python★ 25631 лип. 2026 р.
MIT31 лип. 2026 р. · метрики 2.10.0
PyPI · npm
19Критичнийіндекс здоров'я
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/міс5 серп. 2026 р.
Apache-2.05 серп. 2026 р. · метрики 2.10.0