All tags
Catalogue tag

#evals

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

14 records
Tagged “evals”Ranked by health index
PyPI · npm
89Excellenthealth index
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript · Jupyter Notebook★ 10.6KJul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
PyPI · npm
86Excellenthealth index
pydantic/logfire
AI observability platform for production LLM and agent systems.
Python★ 4,380↓ 23.5M/moJul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
npm
84Goodhealth index
mastra-ai/mastra
Mastra is the modern TypeScript framework for AI-powered applications and agents.
TypeScript★ 26.2K↓ 0/moJul 14, 2026
Custom licenseJul 14, 2026 · metrics 1.13.0
PyPI · npm
76Goodhealth index
harbor-framework/harbor
Framework for evaluating and improving agents
Python★ 3,333↓ 9.3M/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
npm
75Goodhealth index
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2,069↓ 55.8K/moJul 18, 2026
Custom licenseJul 18, 2026 · metrics 1.13.0
PyPI · npm
71Goodhealth index
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5,716Jul 21, 2026
MITJul 21, 2026 · metrics 1.13.0
Go
67Moderatehealth index
realkarych/catacomb
Regression testing for Claude Code and Codex agents.
Go · Shell★ 2Jul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 1.13.0
npm
66Moderatehealth index
zernie/vigiles
Like Lighthouse for your agent harness - verify your CLAUDE.md/AGENTS.md, skills & hooks are real, then test and measure they actually work. Claude Code + Codex.
TypeScript · JavaScript★ 12↓ 4,369/moJul 19, 2026
MITJul 19, 2026 · metrics 1.13.0
PyPI
64Moderatehealth index
kensa-sh/kensa
Kensa turns agent traces into evals that run in CI.
Python★ 2↓ 2,171/moJul 23, 2026
Apache-2.0Jul 23, 2026 · metrics 1.13.0
Hex
58Moderatehealth index
aryaminus/controlkeel
Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Elixir★ 10Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
npm
58Moderatehealth index
spences10/my-pi
Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
TypeScript★ 88↓ 5,088/moJul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
Go
55Moderatehealth index
farazhassan/gantry
A tiny testable, Go-native agent runtime for teams that want control, conformance, and no framework lock-ins.
Go★ 1Jul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
npm
46At riskhealth index
eigenpal/cli
Create, evaluate, and deploy workflows from your terminal. Agent-ready.
TypeScript★ 2↓ 4,949/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
43At riskhealth index
splox-ai/python-sdk
Official Splox SDK
Python★ 0↓ 3,588/moJul 16, 2026
MITJul 16, 2026 · metrics 1.13.0