Alle Tags
Katalog-Tag

#llm-eval

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

9 Einträge
Getaggt als „llm-eval“Geordnet nach Gesundheitsindex
PyPI · npm
100AußergewöhnlichGesundheitsindex
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2K28. Aug. 2026
Eigene Lizenz28. Aug. 2026 · Metriken 2.10.0
PyPI
97AußergewöhnlichGesundheitsindex
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python★ 5.77528. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
npm
96AußergewöhnlichGesundheitsindex
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/Monat5. Aug. 2026
MIT5. Aug. 2026 · Metriken 2.10.0
PyPI
90ExzellentGesundheitsindex
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3.4876. Aug. 2026
MIT6. Aug. 2026 · Metriken 2.10.0
PyPI · npm
89ExzellentGesundheitsindex
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/Monat19. Aug. 2026
Apache-2.019. Aug. 2026 · Metriken 2.10.0
PyPI
78GutGesundheitsindex
attenlabs/hotato
Find what broke in your agent calls. Pin it so it never ships again. Local voice-agent call forensics and regression guards.
Python★ 1↓ 1.680/Monat22. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
Go
69GutGesundheitsindex
lehigh-university-libraries/htr
Handwritten Text Recognition llm eval tool
Go★ 23. Aug. 2026
Apache-2.03. Aug. 2026 · Metriken 2.10.0
RubyGems
67GutGesundheitsindex
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 117. Juli 2026
Eigene Lizenz17. Juli 2026 · Metriken 2.10.0
Go · npm
60MittelGesundheitsindex
valbaudo/awf
Run agents you don't babysit, and trust the result. awf runs agentic workflows with independent gates that check every step and resumes after crashes.
Go★ 122. Juli 2026
Apache-2.022. Juli 2026 · Metriken 2.10.0