npm · PyPI78良好健康指数confident-ai/deepevalThe LLM Evaluation FrameworkPython · TypeScript★ 17K↓ 17.7K/月2026年7月20日Apache-2.02026年7月20日 · 指标 1.13.0
RubyGems61中等健康指数homemade-software-inc/completion-kitYour prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.Ruby · HTML★ 12026年7月17日自定义许可证2026年7月17日 · 指标 1.13.0