全部标签
目录标签

#llm-evaluation

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

25 条记录
标签为“llm-evaluation”按健康指数排序
PyPI · npm
100卓越健康指数
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2K2026年8月28日
自定义许可证2026年8月28日 · 指标 2.10.0
PyPI
97卓越健康指数
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python★ 5,7752026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
npm · PyPI
96卓越健康指数
comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
npm
96卓越健康指数
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI
94卓越健康指数
Q00/ouroboros
Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
Python★ 5,4032026年8月13日
MIT2026年8月13日 · 指标 2.10.0
PyPI · npm
94卓越健康指数
confident-ai/deepeval
The LLM Evaluation Framework
Python · TypeScript★ 17.4K↓ 6.3M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
npm
94卓越健康指数
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5K2026年8月5日
自定义许可证2026年8月5日 · 指标 2.10.0
PyPI
92优秀健康指数
JudgmentLabs/judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Python★ 1,057↓ 233.1K/月2026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
Go
91优秀健康指数
praetorian-inc/julius
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 1752026年8月8日
Apache-2.02026年8月8日 · 指标 2.10.0
npm · PyPI
90优秀健康指数
Marker-Inc-Korea/AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4,963↓ 309/月2026年8月2日
自定义许可证2026年8月2日 · 指标 2.10.0
PyPI
90优秀健康指数
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3,4872026年8月6日
MIT2026年8月6日 · 指标 2.10.0
PyPI · npm
89优秀健康指数
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/月2026年8月19日
Apache-2.02026年8月19日 · 指标 2.10.0
npm
89优秀健康指数
microsoft/prompty
Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandability, and portability for developers.
Rust · C# · TypeScript★ 1,2482026年8月13日
MIT2026年8月13日 · 指标 2.10.0
PyPI
77良好健康指数
alizahidraja/isnad
Grade every agent, scraper and model in a claim's chain — provenance, trust scoring and audit evidence for LLM pipelines
Python★ 37↓ 4,204/月2026年8月29日
Apache-2.02026年8月29日 · 指标 2.10.0
Go · Maven · npm
77良好健康指数
hugalafutro/model-hotel
Multi-Provider AI Gateway - No personal logs by design. Model autodiscovery, Failover groups, High availabilty, Android companion app, and more. "Because we have LiteLLM at home"
Go · TypeScript★ 502026年7月18日
MIT2026年7月18日 · 指标 2.10.0
PyPI
73良好健康指数
b7n0de/proofbundle
Offline cryptographic receipts for AI evaluation results — Ed25519 + RFC 6962 Merkle + optional SD-JWT. Integrity, not truth
Python★ 2↓ 6,574/月2026年7月23日
MIT2026年7月23日 · 指标 2.10.0
Hex · npm
73良好健康指数
ccarvalho-eng/aludel
LLM Evaluation for Phoenix Apps
Elixir · JavaScript · HTML★ 34↓ 19.3K/月2026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
PyPI
73良好健康指数
phierceweb/pf-core
Python foundation for LLM apps whose prompts and spend you can actually see — versioned prompts, every call recorded and replayable, budgets, evals, jobs.
Python★ 2↓ 1,160/月2026年9月5日
MIT2026年9月5日 · 指标 2.10.0
RubyGems
67良好健康指数
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 12026年7月17日
自定义许可证2026年7月17日 · 指标 2.10.0
PyPI · npm
63中等健康指数
fastaifoundry/fastaiagent-sdk
该仓库未发布描述。
Python · TypeScript★ 0↓ 2,525/月2026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
Go
54中等健康指数
tamnd/taocp-solver
A Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 02026年7月28日
MIT2026年7月28日 · 指标 2.10.0
npm
53中等健康指数
mykim-aus/hey-llm-you-okay
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2,332/月2026年7月30日
MIT2026年7月30日 · 指标 2.10.0
PyPI
50中等健康指数
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2,908/月2026年7月22日
MIT2026年7月22日 · 指标 2.10.0
PyPI
34存在风险健康指数
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 22026年7月17日
无许可证2026年7月17日 · 指标 2.10.0
PyPI · npm
19危急健康指数
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0