All tags
Catalogue tag

#llm-eval

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

4 records
Tagged “llm-eval”Ranked by health index
PyPI · npm
89Excellenthealth index
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript · Jupyter Notebook★ 10.6KJul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
npm
80Goodhealth index
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.4K↓ 1.7M/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
RubyGems
61Moderatehealth index
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 1Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
Go · npm
56Moderatehealth index
valbaudo/awf
Run agents you don't babysit, and trust the result. awf runs agentic workflows with independent gates that check every step and resumes after crashes.
Go★ 1Jul 22, 2026
Apache-2.0Jul 22, 2026 · metrics 1.13.0