全部标签
目录标签

#evaluation

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

58 条记录
标签为“evaluation”按健康指数排序
npm
69良好健康指数
Kromatic-Innovation/panelist
Synthetic user panels for any artifact, run across multiple model providers to correct for self-preference bias — models the exact point a persona quits, dismisses, or refuses to click, not a warmth score.
JavaScript★ 0↓ 3,165/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
npm
69良好健康指数
crewhaus/factory
Open-source compiler for AI agents. Write one crewhaus.yaml; compile it to a CLI, a Slack bot, and an eval harness from the same spec. Apache-2.0.
TypeScript★ 2↓ 25.7K/月2026年7月24日
Apache-2.02026年7月24日 · 指标 2.10.0
PyPI
69良好健康指数
robocurve/worldevals
A curated catalog of VLA / physical-AI benchmarks, each runnable on real robots or sims via Inspect Robots. (The Inspect Evals for robotics.)
Python★ 62026年7月30日
MIT2026年7月30日 · 指标 2.10.0
PyPI
67良好健康指数
OpenAdaptAI/openadapt-evals
Evaluation infrastructure for GUI agent benchmarks
Python★ 2↓ 2,941/月2026年7月28日
MIT2026年7月28日 · 指标 2.10.0
npm
63中等健康指数
axl-sdk/axl
TypeScript SDK for orchestrating Agentic Systems — concurrency, structured output, cost control, and consensus as first-class primitives.
TypeScript★ 2↓ 3,214/月2026年7月27日
Apache-2.02026年7月27日 · 指标 2.10.0
PyPI
63中等健康指数
davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge
HTML · Python★ 692026年9月5日
无许可证2026年9月5日 · 指标 2.10.0
npm
63中等健康指数
mgechev/skillgrade
"Unit tests" for your agent skills
TypeScript★ 661↓ 2,290/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI
62中等健康指数
AgentX-ai/AgentX-Python
AgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
Python★ 682026年8月3日
MIT2026年8月3日 · 指标 2.10.0
npm
62中等健康指数
TypeScript · JavaScript★ 1↓ 3,920/月2026年8月8日
无许可证2026年8月8日 · 指标 2.10.0
PyPI · crates.io
62中等健康指数
nickderobertis/onejudge
A simulated interaction and evaluation loop over oneharness: drive a harness through a multi-turn conversation and score the transcript.
Rust★ 0↓ 6,844/月2026年8月29日
MIT2026年8月29日 · 指标 2.10.0
npm
60中等健康指数
CarlosNZ/fig-tree-evaluator
A highly configurable custom expression tree evaluator
TypeScript★ 24↓ 3,662/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
Go · npm
60中等健康指数
starkSV/windows-iso-downloader
Download official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript · Go★ 142026年8月22日
MIT2026年8月22日 · 指标 2.10.0
PyPI
60中等健康指数
toshas/torch-fidelity
High-fidelity performance metrics for generative models in PyTorch
Python★ 1,197↓ 846.4K/月2026年8月13日
自定义许可证2026年8月13日 · 指标 2.10.0
PyPI
57中等健康指数
EpsilabAI/epsilab-python
The official Python library for the Epsilab API
Python★ 0↓ 2,706/月2026年7月21日
Apache-2.02026年7月21日 · 指标 2.10.0
PyPI
56中等健康指数
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 6072026年7月21日
自定义许可证2026年7月21日 · 指标 2.10.0
PyPI
54中等健康指数
danaug23/harness-arena
Your model, many harnesses, many benchmarks.
Python · HTML★ 2↓ 2,988/月2026年8月23日
Apache-2.02026年8月23日 · 指标 2.10.0
PyPI
54中等健康指数
lazily-hub/lazily-py
A Python library for lazy evaluation with context caching.
Python★ 12026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
Go
54中等健康指数
tamnd/taocp-solver
A Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 02026年7月28日
MIT2026年7月28日 · 指标 2.10.0
npm
51中等健康指数
HolocronLab/botruntime-packages
botruntime public packages. Consumed by the botruntime platform.
TypeScript★ 0↓ 48.8K/月2026年7月29日
无许可证2026年7月29日 · 指标 2.10.0
PyPI
50中等健康指数
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2,908/月2026年7月22日
MIT2026年7月22日 · 指标 2.10.0
PyPI
44薄弱健康指数
huggingface/Math-Verify
该仓库未发布描述。
Python★ 1,1762026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
npm
34存在风险健康指数
darks0l/modelab
Autonomous research agent SDK
TypeScript · JavaScript★ 1↓ 203/月2026年9月5日
无许可证2026年9月5日 · 指标 2.10.0
NuGet
34存在风险健康指数
ncalc/ncalc
NCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.
C#★ 1,1422026年7月31日
MIT2026年7月31日 · 指标 2.10.0
npm
34存在风险健康指数
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 353M/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
PyPI
34存在风险健康指数
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 22026年7月17日
无许可证2026年7月17日 · 指标 2.10.0
28存在风险健康指数
DanceNitra/ramr
RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)
Python★ 02026年7月29日
MIT2026年7月29日 · 指标 2.10.0
PyPI
25存在风险健康指数
evfro/polara
Recommender system and evaluation framework for top-n recommendations tasks that respects polarity of feedbacks. Fast, flexible and easy to use. Written in python, boosted by scientific python stack.
Python★ 2562026年7月31日
MIT2026年7月31日 · 指标 2.10.0
PyPI · npm
19危急健康指数
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0