PyPI · npm94卓越健康指数confident-ai/deepevalThe LLM Evaluation FrameworkPython · TypeScript★ 17.4K↓ 6.3M/月2026年8月5日Apache-2.02026年8月5日 · 指标 2.10.0
RubyGems67良好健康指数homemade-software-inc/completion-kitYour prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.Ruby · HTML★ 12026年7月17日自定义许可证2026年7月17日 · 指标 2.10.0