Alle Tags
Katalog-Tag

#eval

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

22 Einträge
Getaggt als „eval“Geordnet nach Gesundheitsindex
PyPI
94AußergewöhnlichGesundheitsindex
PrimeIntellect-ai/verifiers
Our library for RL environments + evals
Python★ 4.565↓ 378.1K/Monat28. Aug. 2026
MIT28. Aug. 2026 · Metriken 2.10.0
PyPI · npm
89ExzellentGesundheitsindex
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/Monat19. Aug. 2026
Apache-2.019. Aug. 2026 · Metriken 2.10.0
npm
78GutGesundheitsindex
vercel-labs/agent-eval
Keine Repository-Beschreibung veröffentlicht.
TypeScript · JavaScript★ 229↓ 181.2K/Monat31. Juli 2026
Keine Lizenz31. Juli 2026 · Metriken 2.10.0
npm
77GutGesundheitsindex
justjake/quickjs-emscripten
Safely execute untrusted Javascript in your Javascript, and execute synchronous code that uses async functions
TypeScript · Makefile★ 1.691↓ 21.5M/Monat13. Aug. 2026
Eigene Lizenz13. Aug. 2026 · Metriken 2.10.0
PyPI
77GutGesundheitsindex
lmfit/asteval
minimalistic evaluator of python expression using ast module
Python★ 21921. Juli 2026
MIT21. Juli 2026 · Metriken 2.10.0
npm
77GutGesundheitsindex
zernie/vigiles
Like Lighthouse for your agent harness - verify your CLAUDE.md/AGENTS.md, skills & hooks are real, then test and measure they actually work. Claude Code + Codex.
TypeScript · JavaScript★ 12↓ 4.369/Monat19. Juli 2026
MIT19. Juli 2026 · Metriken 2.10.0
PyPI
73GutGesundheitsindex
phierceweb/pf-core
Python foundation for LLM apps whose prompts and spend you can actually see — versioned prompts, every call recorded and replayable, budgets, evals, jobs.
Python★ 2↓ 1.160/Monat5. Sept. 2026
MIT5. Sept. 2026 · Metriken 2.10.0
Packagist
71GutGesundheitsindex
wp-cli/eval-command
Executes arbitrary PHP code or files.
Gherkin · PHP★ 10↓ 300.6K/Monat20. Juli 2026
MIT20. Juli 2026 · Metriken 2.10.0
npm
69GutGesundheitsindex
crewhaus/factory
Open-source compiler for AI agents. Write one crewhaus.yaml; compile it to a CLI, a Slack bot, and an eval harness from the same spec. Apache-2.0.
TypeScript★ 2↓ 25.7K/Monat24. Juli 2026
Apache-2.024. Juli 2026 · Metriken 2.10.0
npm
69GutGesundheitsindex
tangle-network/agent-app
A feature-full starting point for production agent applications on Tangle.
TypeScript★ 0↓ 47.2K/Monat17. Juli 2026
MIT17. Juli 2026 · Metriken 2.10.0
npm
67GutGesundheitsindex
peerigon/angular-expressions
Angular expressions as standalone module
JavaScript★ 100↓ 569.9K/Monat5. Sept. 2026
Unlicense5. Sept. 2026 · Metriken 2.10.0
npm
65GutGesundheitsindex
PruvoNet/json-expression-eval
json serializable rule engine / boolean expression evaluator
TypeScript★ 22↓ 6.681/Monat18. Juli 2026
MIT18. Juli 2026 · Metriken 2.10.0
npm
63MittelGesundheitsindex
mgechev/skillgrade
"Unit tests" for your agent skills
TypeScript★ 661↓ 2.290/Monat5. Aug. 2026
MIT5. Aug. 2026 · Metriken 2.10.0
Packagist
63MittelGesundheitsindex
pestphp/pest-plugin-evals
Pest Browser Evals
PHP★ 4↓ 7.386/Monat4. Aug. 2026
MIT4. Aug. 2026 · Metriken 2.10.0
npm
62MittelGesundheitsindex
reggi/evalmd
:fishing_pole_and_fish: Evaluates javascript code blocks from markdown files.
JavaScript★ 47↓ 153.9K/Monat29. Juli 2026
MIT29. Juli 2026 · Metriken 2.10.0
crates.io
60MittelGesundheitsindex
AndyCappDev/stet
A PDF rendering engine and PostScript Level 3 interpreter written in pure Rust.
Rust★ 11↓ 91.6K/Monat19. Aug. 2026
Apache-2.019. Aug. 2026 · Metriken 2.10.0
npm
60MittelGesundheitsindex
gemstack-land/the-framework
Autonomous AI programming. Stop babysitting your coding agents. You make the important decisions; AI does the rest.
TypeScript★ 8↓ 5.492/Monat30. Juli 2026
MIT30. Juli 2026 · Metriken 2.10.0
RubyGems
59MittelGesundheitsindex
justi/ruby_llm-contract
Validate and retry LLM outputs for ruby_llm. Describe the JSON response you expect, fall back to a stronger model when the cheaper one fails the rules, and gate CI on regressions — all as one contract object per step.
Ruby★ 3522. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
PyPI
56MittelGesundheitsindex
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 60721. Juli 2026
Eigene Lizenz21. Juli 2026 · Metriken 2.10.0
npm
53MittelGesundheitsindex
mykim-aus/hey-llm-you-okay
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2.332/Monat30. Juli 2026
MIT30. Juli 2026 · Metriken 2.10.0
crates.io
45SchwachGesundheitsindex
bugadani/somni
Keine Repository-Beschreibung veröffentlicht.
Rust★ 2↓ 102.5K/Monat22. Aug. 2026
Apache-2.022. Aug. 2026 · Metriken 2.10.0
npm
34GefährdetGesundheitsindex
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 353M/Monat4. Aug. 2026
MIT4. Aug. 2026 · Metriken 2.10.0