
Kromatic-Innovation/panelistSynthetic user panels for any artifact, run across multiple model providers to correct for self-preference bias — models the exact point a persona quits, dismisses, or refuses to click, not a warmth score.
JavaScript★ 0↓ 3,165/moAug 28, 2026
crewhaus/factoryOpen-source compiler for AI agents. Write one crewhaus.yaml; compile it to a CLI, a Slack bot, and an eval harness from the same spec. Apache-2.0.
TypeScript★ 2↓ 25.7K/moJul 24, 2026
robocurve/worldevalsA curated catalog of VLA / physical-AI benchmarks, each runnable on real robots or sims via Inspect Robots. (The Inspect Evals for robotics.)
Python★ 6Jul 30, 2026
Python★ 2↓ 2,941/moJul 28, 2026
npm63Moderatehealth index
axl-sdk/axlTypeScript SDK for orchestrating Agentic Systems — concurrency, structured output, cost control, and consensus as first-class primitives.
TypeScript★ 2↓ 3,214/moJul 27, 2026
PyPI63Moderatehealth index
HTML · Python★ 69Sep 5, 2026
npm63Moderatehealth index
TypeScript★ 661↓ 2,290/moAug 5, 2026
PyPI62Moderatehealth index
AgentX-ai/AgentX-PythonAgentX python SDK. Build multi-agent AI workforce. Run evaluation. Trace your agent. Full Observability.
Python★ 68Aug 3, 2026
npm62Moderatehealth index
TypeScript · JavaScript★ 1↓ 3,920/moAug 8, 2026
PyPI · crates.io62Moderatehealth index

nickderobertis/onejudgeA simulated interaction and evaluation loop over oneharness: drive a harness through a multi-turn conversation and score the transcript.
Rust★ 0↓ 6,844/moAug 29, 2026
npm60Moderatehealth index
TypeScript★ 24↓ 3,662/moAug 5, 2026
Go · npm60Moderatehealth index

starkSV/windows-iso-downloaderDownload official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript · Go★ 14Aug 22, 2026
PyPI60Moderatehealth index
Python★ 1,197↓ 846.4K/moAug 13, 2026
PyPI57Moderatehealth index
Python★ 0↓ 2,706/moJul 21, 2026
PyPI56Moderatehealth index
Python★ 607Jul 21, 2026
PyPI54Moderatehealth index
Python · HTML★ 2↓ 2,988/moAug 23, 2026
PyPI54Moderatehealth index
Python★ 1Jul 17, 2026
tamnd/taocp-solverA Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 0Jul 28, 2026
npm51Moderatehealth index
TypeScript★ 0↓ 48.8K/moJul 29, 2026
PyPI50Moderatehealth index
Python★ 0↓ 2,908/moJul 22, 2026
Python★ 1,176Aug 13, 2026
TypeScript · JavaScript★ 1↓ 203/moSep 5, 2026
ncalc/ncalcNCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.
C#★ 1,142Jul 31, 2026
JavaScript · TypeScript★ 67↓ 353M/moAug 4, 2026
PyPI34At Riskhealth index
waybarrios/crystalCRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 2Jul 17, 2026
DanceNitra/ramrRAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)
Python★ 0Jul 29, 2026
evfro/polaraRecommender system and evaluation framework for top-n recommendations tasks that respects polarity of feedbacks. Fast, flexible and easy to use. Written in python, boosted by scientific python stack.
Python★ 256Jul 31, 2026

mlflow/mlflowThe open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/moAug 5, 2026