Todas las etiquetas
Etiqueta del catálogo

#evaluation

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

58 registros
Con la etiqueta «evaluation»Ordenado por índice de salud
npm
69Buenoíndice de salud
Kromatic-Innovation/panelist
Synthetic user panels for any artifact, run across multiple model providers to correct for self-preference bias — models the exact point a persona quits, dismisses, or refuses to click, not a warmth score.
JavaScript★ 0↓ 3165/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
npm
69Buenoíndice de salud
crewhaus/factory
Open-source compiler for AI agents. Write one crewhaus.yaml; compile it to a CLI, a Slack bot, and an eval harness from the same spec. Apache-2.0.
TypeScript★ 2↓ 25.7K/mes24 jul 2026
Apache-2.024 jul 2026 · métricas 2.10.0
PyPI
69Buenoíndice de salud
robocurve/worldevals
A curated catalog of VLA / physical-AI benchmarks, each runnable on real robots or sims via Inspect Robots. (The Inspect Evals for robotics.)
Python★ 630 jul 2026
MIT30 jul 2026 · métricas 2.10.0
PyPI
67Buenoíndice de salud
OpenAdaptAI/openadapt-evals
Evaluation infrastructure for GUI agent benchmarks
Python★ 2↓ 2941/mes28 jul 2026
MIT28 jul 2026 · métricas 2.10.0
npm
63Moderadoíndice de salud
axl-sdk/axl
TypeScript SDK for orchestrating Agentic Systems — concurrency, structured output, cost control, and consensus as first-class primitives.
TypeScript★ 2↓ 3214/mes27 jul 2026
Apache-2.027 jul 2026 · métricas 2.10.0
PyPI
63Moderadoíndice de salud
davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge
HTML · Python★ 695 sept 2026
Sin licencia5 sept 2026 · métricas 2.10.0
npm
63Moderadoíndice de salud
mgechev/skillgrade
"Unit tests" for your agent skills
TypeScript★ 661↓ 2290/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
PyPI
62Moderadoíndice de salud
AgentX-ai/AgentX-Python
AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
Python★ 683 ago 2026
MIT3 ago 2026 · métricas 2.10.0
npm
62Moderadoíndice de salud
TypeScript · JavaScript★ 1↓ 3920/mes8 ago 2026
Sin licencia8 ago 2026 · métricas 2.10.0
PyPI · crates.io
62Moderadoíndice de salud
nickderobertis/onejudge
A simulated interaction and evaluation loop over oneharness: drive a harness through a multi-turn conversation and score the transcript.
Rust★ 0↓ 6844/mes29 ago 2026
MIT29 ago 2026 · métricas 2.10.0
npm
60Moderadoíndice de salud
CarlosNZ/fig-tree-evaluator
A highly configurable custom expression tree evaluator
TypeScript★ 24↓ 3662/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
Go · npm
60Moderadoíndice de salud
starkSV/windows-iso-downloader
Download official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript · Go★ 1422 ago 2026
MIT22 ago 2026 · métricas 2.10.0
PyPI
60Moderadoíndice de salud
toshas/torch-fidelity
High-fidelity performance metrics for generative models in PyTorch
Python★ 1197↓ 846.4K/mes13 ago 2026
Licencia propia13 ago 2026 · métricas 2.10.0
PyPI
57Moderadoíndice de salud
EpsilabAI/epsilab-python
The official Python library for the Epsilab API
Python★ 0↓ 2706/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
PyPI
56Moderadoíndice de salud
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 60721 jul 2026
Licencia propia21 jul 2026 · métricas 2.10.0
PyPI
54Moderadoíndice de salud
danaug23/harness-arena
Your model, many harnesses, many benchmarks.
Python · HTML★ 2↓ 2988/mes23 ago 2026
Apache-2.023 ago 2026 · métricas 2.10.0
PyPI
54Moderadoíndice de salud
lazily-hub/lazily-py
A Python library for lazy evaluation with context caching.
Python★ 117 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
Go
54Moderadoíndice de salud
tamnd/taocp-solver
A Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 028 jul 2026
MIT28 jul 2026 · métricas 2.10.0
npm
51Moderadoíndice de salud
HolocronLab/botruntime-packages
botruntime public packages. Consumed by the botruntime platform.
TypeScript★ 0↓ 48.8K/mes29 jul 2026
Sin licencia29 jul 2026 · métricas 2.10.0
PyPI
50Moderadoíndice de salud
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2908/mes22 jul 2026
MIT22 jul 2026 · métricas 2.10.0
PyPI
44Débilíndice de salud
huggingface/Math-Verify
El repositorio no publica descripción.
Python★ 117613 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
npm
34En riesgoíndice de salud
darks0l/modelab
Autonomous research agent SDK
TypeScript · JavaScript★ 1↓ 203/mes5 sept 2026
Sin licencia5 sept 2026 · métricas 2.10.0
NuGet
34En riesgoíndice de salud
ncalc/ncalc
NCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.
C#★ 114231 jul 2026
MIT31 jul 2026 · métricas 2.10.0
npm
34En riesgoíndice de salud
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 353M/mes4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
PyPI
34En riesgoíndice de salud
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 217 jul 2026
Sin licencia17 jul 2026 · métricas 2.10.0
28En riesgoíndice de salud
DanceNitra/ramr
RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)
Python★ 029 jul 2026
MIT29 jul 2026 · métricas 2.10.0
PyPI
25En riesgoíndice de salud
evfro/polara
Recommender system and evaluation framework for top-n recommendations tasks that respects polarity of feedbacks. Fast, flexible and easy to use. Written in python, boosted by scientific python stack.
Python★ 25631 jul 2026
MIT31 jul 2026 · métricas 2.10.0
PyPI · npm
19Críticoíndice de salud
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0