全部标签
目录标签

#evals

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

29 条记录
标签为“evals”按健康指数排序
PyPI · npm
100卓越健康指数
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2K2026年8月28日
自定义许可证2026年8月28日 · 指标 2.10.0
PyPI · npm
99卓越健康指数
pydantic/logfire
AI observability platform for production LLM and agent systems.
Python★ 4,442↓ 14.6M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
npm
98卓越健康指数
mastra-ai/mastra
Mastra is the modern TypeScript framework for AI-powered applications and agents.
TypeScript★ 26.9K↓ 61.2K/月2026年8月5日
自定义许可证2026年8月5日 · 指标 2.10.0
PyPI
94卓越健康指数
langchain-ai/deepagents
The batteries-included agent harness.
Python★ 27.3K↓ 210.2K/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI · npm
93卓越健康指数
harbor-framework/harbor
Framework for evaluating and improving agents
Python★ 4,6942026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
npm
90优秀健康指数
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2,069↓ 55.8K/月2026年7月18日
自定义许可证2026年7月18日 · 指标 2.10.0
PyPI
90优秀健康指数
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3,4872026年8月6日
MIT2026年8月6日 · 指标 2.10.0
PyPI · npm
89优秀健康指数
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/月2026年8月19日
Apache-2.02026年8月19日 · 指标 2.10.0
npm
84优秀健康指数
nearform/lastlight
Self-hostable, MIT Licensed, Enterprise AI Software Factory
TypeScript · Astro★ 18↓ 25.1K/月2026年7月26日
MIT2026年7月26日 · 指标 2.10.0
PyPI · npm
80优秀健康指数
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5,801↓ 248.9K/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
PyPI
80优秀健康指数
benchflow-ai/benchflow
Research infra for creating RL environments, post-training, and evals.
Python★ 317↓ 6,133/月2026年8月9日
Apache-2.02026年8月9日 · 指标 2.10.0
npm
80优秀健康指数
o-stepper/graphorin
Project Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
Go
80优秀健康指数
realkarych/catacomb
Regression testing for Claude Code and Codex agents.
Go · Shell★ 22026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
PyPI
78良好健康指数
attenlabs/hotato
Find what broke in your agent calls. Pin it so it never ships again. Local voice-agent call forensics and regression guards.
Python★ 1↓ 1,680/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm
78良好健康指数
getlarge/themoltnet
Trusted context for AI agents
TypeScript · Go★ 15↓ 7,904/月2026年7月30日
AGPL-3.02026年7月30日 · 指标 2.10.0
PyPI
78良好健康指数
superlinear-ai/raglite
🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL
Python★ 1,2002026年8月10日
MPL-2.02026年8月10日 · 指标 2.10.0
npm
77良好健康指数
zernie/vigiles
Like Lighthouse for your agent harness - verify your CLAUDE.md/AGENTS.md, skills & hooks are real, then test and measure they actually work. Claude Code + Codex.
TypeScript · JavaScript★ 12↓ 4,369/月2026年7月19日
MIT2026年7月19日 · 指标 2.10.0
PyPI
73良好健康指数
kensa-sh/kensa
Kensa turns agent traces into evals that run in CI.
Python★ 2↓ 2,171/月2026年7月23日
Apache-2.02026年7月23日 · 指标 2.10.0
Packagist
63中等健康指数
pestphp/pest-plugin-evals
Pest Browser Evals
PHP★ 4↓ 7,386/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
Hex
62中等健康指数
aryaminus/controlkeel
Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Elixir★ 102026年7月17日
自定义许可证2026年7月17日 · 指标 2.10.0
npm
62中等健康指数
shulmansj/teami
Control plane for orchestrating, evaluating, and improving agent work across a company
JavaScript★ 0↓ 2,136/月2026年7月30日
自定义许可证2026年7月30日 · 指标 2.10.0
npm
62中等健康指数
spences10/my-pi
Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
TypeScript★ 88↓ 5,088/月2026年7月18日
MIT2026年7月18日 · 指标 2.10.0
Go
57中等健康指数
farazhassan/gantry
A tiny testable, Go-native agent runtime for teams that want control, conformance, and no framework lock-ins.
Go★ 12026年7月18日
MIT2026年7月18日 · 指标 2.10.0
npm
54中等健康指数
inferock/inferock-bench
Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.
TypeScript★ 123↓ 6,119/月2026年7月29日
自定义许可证2026年7月29日 · 指标 2.10.0
npm
53中等健康指数
mykim-aus/hey-llm-you-okay
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2,332/月2026年7月30日
MIT2026年7月30日 · 指标 2.10.0
npm
51中等健康指数
HolocronLab/botruntime-packages
botruntime public packages. Consumed by the botruntime platform.
TypeScript★ 0↓ 48.8K/月2026年7月29日
无许可证2026年7月29日 · 指标 2.10.0
npm
50中等健康指数
LilMGenius/paperthin
Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
Shell · JavaScript★ 119↓ 4,136/月2026年7月31日
MIT2026年7月31日 · 指标 2.10.0
npm
47薄弱健康指数
eigenpal/cli
Create, evaluate, and deploy workflows from your terminal. Agent-ready.
TypeScript★ 2↓ 4,949/月2026年7月18日
Apache-2.02026年7月18日 · 指标 2.10.0
PyPI
39薄弱健康指数
splox-ai/python-sdk
Official Splox SDK
Python★ 0↓ 3,588/月2026年7月16日
MIT2026年7月16日 · 指标 2.10.0