全部标签
目录标签

#llm-eval

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

9 条记录
标签为“llm-eval”按健康指数排序
PyPI · npm
100卓越健康指数
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2K2026年8月28日
自定义许可证2026年8月28日 · 指标 2.10.0
PyPI
97卓越健康指数
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python★ 5,7752026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
npm
96卓越健康指数
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI
90优秀健康指数
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3,4872026年8月6日
MIT2026年8月6日 · 指标 2.10.0
PyPI · npm
89优秀健康指数
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/月2026年8月19日
Apache-2.02026年8月19日 · 指标 2.10.0
PyPI
78良好健康指数
attenlabs/hotato
Find what broke in your agent calls. Pin it so it never ships again. Local voice-agent call forensics and regression guards.
Python★ 1↓ 1,680/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
Go
69良好健康指数
lehigh-university-libraries/htr
Handwritten Text Recognition llm eval tool
Go★ 22026年8月3日
Apache-2.02026年8月3日 · 指标 2.10.0
RubyGems
67良好健康指数
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 12026年7月17日
自定义许可证2026年7月17日 · 指标 2.10.0
Go · npm
60中等健康指数
valbaudo/awf
Run agents you don't babysit, and trust the result. awf runs agentic workflows with independent gates that check every step and resumes after crashes.
Go★ 12026年7月22日
Apache-2.02026年7月22日 · 指标 2.10.0