All tags
Catalogue tag

#llm-eval

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

9 records
Tagged “llm-eval”Ranked by health index
PyPI · npm
100Exceptionalhealth index
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2KAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
PyPI
97Exceptionalhealth index
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python★ 5,775Aug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
npm
97Exceptionalhealth index
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 25.2K↓ 2.6M/moSep 16, 2026
MITSep 16, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3,561Sep 17, 2026
MITSep 17, 2026 · metrics 2.10.0
PyPI · npm
89Excellenthealth index
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
attenlabs/hotato
Find what broke in your agent calls. Pin it so it never ships again. Local voice-agent call forensics and regression guards.
Python★ 1↓ 1,680/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
RubyGems
69Goodhealth index
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 3Sep 10, 2026
Custom licenseSep 10, 2026 · metrics 2.10.0
Go · npm
63Moderatehealth index
valbaudo/awf
Run agents you don't babysit, and trust the result. awf runs agentic workflows with independent gates that check every step and resumes after crashes.
Go★ 1Sep 12, 2026
Apache-2.0Sep 12, 2026 · metrics 2.10.0