Alle Tags
Katalog-Tag

#evaluation

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

58 Einträge
Getaggt als „evaluation“Geordnet nach Gesundheitsindex
npm
69GutGesundheitsindex
Kromatic-Innovation/panelist
Synthetic user panels for any artifact, run across multiple model providers to correct for self-preference bias — models the exact point a persona quits, dismisses, or refuses to click, not a warmth score.
JavaScript★ 0↓ 3.165/Monat28. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
npm
69GutGesundheitsindex
crewhaus/factory
Open-source compiler for AI agents. Write one crewhaus.yaml; compile it to a CLI, a Slack bot, and an eval harness from the same spec. Apache-2.0.
TypeScript★ 2↓ 25.7K/Monat24. Juli 2026
Apache-2.024. Juli 2026 · Metriken 2.10.0
PyPI
69GutGesundheitsindex
robocurve/worldevals
A curated catalog of VLA / physical-AI benchmarks, each runnable on real robots or sims via Inspect Robots. (The Inspect Evals for robotics.)
Python★ 630. Juli 2026
MIT30. Juli 2026 · Metriken 2.10.0
PyPI
67GutGesundheitsindex
OpenAdaptAI/openadapt-evals
Evaluation infrastructure for GUI agent benchmarks
Python★ 2↓ 2.941/Monat28. Juli 2026
MIT28. Juli 2026 · Metriken 2.10.0
npm
63MittelGesundheitsindex
axl-sdk/axl
TypeScript SDK for orchestrating Agentic Systems — concurrency, structured output, cost control, and consensus as first-class primitives.
TypeScript★ 2↓ 3.214/Monat27. Juli 2026
Apache-2.027. Juli 2026 · Metriken 2.10.0
PyPI
63MittelGesundheitsindex
davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge
HTML · Python★ 695. Sept. 2026
Keine Lizenz5. Sept. 2026 · Metriken 2.10.0
npm
63MittelGesundheitsindex
mgechev/skillgrade
"Unit tests" for your agent skills
TypeScript★ 661↓ 2.290/Monat5. Aug. 2026
MIT5. Aug. 2026 · Metriken 2.10.0
PyPI
62MittelGesundheitsindex
AgentX-ai/AgentX-Python
AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
Python★ 683. Aug. 2026
MIT3. Aug. 2026 · Metriken 2.10.0
npm
62MittelGesundheitsindex
TypeScript · JavaScript★ 1↓ 3.920/Monat8. Aug. 2026
Keine Lizenz8. Aug. 2026 · Metriken 2.10.0
PyPI · crates.io
62MittelGesundheitsindex
nickderobertis/onejudge
A simulated interaction and evaluation loop over oneharness: drive a harness through a multi-turn conversation and score the transcript.
Rust★ 0↓ 6.844/Monat29. Aug. 2026
MIT29. Aug. 2026 · Metriken 2.10.0
npm
60MittelGesundheitsindex
CarlosNZ/fig-tree-evaluator
A highly configurable custom expression tree evaluator
TypeScript★ 24↓ 3.662/Monat5. Aug. 2026
MIT5. Aug. 2026 · Metriken 2.10.0
Go · npm
60MittelGesundheitsindex
starkSV/windows-iso-downloader
Download official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript · Go★ 1422. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
PyPI
60MittelGesundheitsindex
toshas/torch-fidelity
High-fidelity performance metrics for generative models in PyTorch
Python★ 1.197↓ 846.4K/Monat13. Aug. 2026
Eigene Lizenz13. Aug. 2026 · Metriken 2.10.0
PyPI
57MittelGesundheitsindex
EpsilabAI/epsilab-python
The official Python library for the Epsilab API
Python★ 0↓ 2.706/Monat21. Juli 2026
Apache-2.021. Juli 2026 · Metriken 2.10.0
PyPI
56MittelGesundheitsindex
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 60721. Juli 2026
Eigene Lizenz21. Juli 2026 · Metriken 2.10.0
PyPI
54MittelGesundheitsindex
danaug23/harness-arena
Your model, many harnesses, many benchmarks.
Python · HTML★ 2↓ 2.988/Monat23. Aug. 2026
Apache-2.023. Aug. 2026 · Metriken 2.10.0
PyPI
54MittelGesundheitsindex
lazily-hub/lazily-py
A Python library for lazy evaluation with context caching.
Python★ 117. Juli 2026
Apache-2.017. Juli 2026 · Metriken 2.10.0
Go
54MittelGesundheitsindex
tamnd/taocp-solver
A Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 028. Juli 2026
MIT28. Juli 2026 · Metriken 2.10.0
npm
51MittelGesundheitsindex
HolocronLab/botruntime-packages
botruntime public packages. Consumed by the botruntime platform.
TypeScript★ 0↓ 48.8K/Monat29. Juli 2026
Keine Lizenz29. Juli 2026 · Metriken 2.10.0
PyPI
50MittelGesundheitsindex
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2.908/Monat22. Juli 2026
MIT22. Juli 2026 · Metriken 2.10.0
PyPI
44SchwachGesundheitsindex
huggingface/Math-Verify
Keine Repository-Beschreibung veröffentlicht.
Python★ 1.17613. Aug. 2026
Apache-2.013. Aug. 2026 · Metriken 2.10.0
npm
34GefährdetGesundheitsindex
darks0l/modelab
Autonomous research agent SDK
TypeScript · JavaScript★ 1↓ 203/Monat5. Sept. 2026
Keine Lizenz5. Sept. 2026 · Metriken 2.10.0
NuGet
34GefährdetGesundheitsindex
ncalc/ncalc
NCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.
C#★ 1.14231. Juli 2026
MIT31. Juli 2026 · Metriken 2.10.0
npm
34GefährdetGesundheitsindex
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 353M/Monat4. Aug. 2026
MIT4. Aug. 2026 · Metriken 2.10.0
PyPI
34GefährdetGesundheitsindex
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 217. Juli 2026
Keine Lizenz17. Juli 2026 · Metriken 2.10.0
28GefährdetGesundheitsindex
DanceNitra/ramr
RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)
Python★ 029. Juli 2026
MIT29. Juli 2026 · Metriken 2.10.0
PyPI
25GefährdetGesundheitsindex
evfro/polara
Recommender system and evaluation framework for top-n recommendations tasks that respects polarity of feedbacks. Fast, flexible and easy to use. Written in python, boosted by scientific python stack.
Python★ 25631. Juli 2026
MIT31. Juli 2026 · Metriken 2.10.0
PyPI · npm
19KritischGesundheitsindex
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/Monat5. Aug. 2026
Apache-2.05. Aug. 2026 · Metriken 2.10.0