Alle Tags
Katalog-Tag

#llm-evaluation-framework

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

4 Einträge
Getaggt als „llm-evaluation-framework“Geordnet nach Gesundheitsindex
npm
96AußergewöhnlichGesundheitsindex
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/Monat5. Aug. 2026
MIT5. Aug. 2026 · Metriken 2.10.0
PyPI · npm
94AußergewöhnlichGesundheitsindex
confident-ai/deepeval
The LLM Evaluation Framework
Python · TypeScript★ 17.4K↓ 6.3M/Monat5. Aug. 2026
Apache-2.05. Aug. 2026 · Metriken 2.10.0
RubyGems
67GutGesundheitsindex
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 117. Juli 2026
Eigene Lizenz17. Juli 2026 · Metriken 2.10.0
Go
57MittelGesundheitsindex
petmal/MindTrial
MindTrial: Evaluate and compare AI language models (LLMs) on text-based tasks with optional file/image attachments and tool use. Supports multiple providers (OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, OpenRouter), custom tasks in YAML, and HTML/CSV/JSON reports.
Go★ 1730. Aug. 2026
MPL-2.030. Aug. 2026 · Metriken 2.10.0