All tags
Catalogue tag

#evals

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

29 records
Tagged “evals”Ranked by health index
PyPI · npm
100Exceptionalhealth index
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2KAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
PyPI · npm
99Exceptionalhealth index
pydantic/logfire
AI observability platform for production LLM and agent systems.
Python★ 4,442↓ 14.6M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
npm
98Exceptionalhealth index
mastra-ai/mastra
Mastra is the modern TypeScript framework for AI-powered applications and agents.
TypeScript★ 26.9K↓ 61.2K/moAug 5, 2026
Custom licenseAug 5, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
langchain-ai/deepagents
The batteries-included agent harness.
Python★ 27.3K↓ 210.2K/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI · npm
93Exceptionalhealth index
harbor-framework/harbor
Framework for evaluating and improving agents
Python★ 4,694Aug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
npm
90Excellenthealth index
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2,069↓ 55.8K/moJul 18, 2026
Custom licenseJul 18, 2026 · metrics 2.10.0
PyPI
90Excellenthealth index
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3,487Aug 6, 2026
MITAug 6, 2026 · metrics 2.10.0
PyPI · npm
89Excellenthealth index
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
npm
84Excellenthealth index
nearform/lastlight
Self-hostable, MIT Licensed, Enterprise AI Software Factory
TypeScript · Astro★ 18↓ 25.1K/moJul 26, 2026
MITJul 26, 2026 · metrics 2.10.0
PyPI · npm
80Excellenthealth index
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5,801↓ 248.9K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
benchflow-ai/benchflow
Research infra for creating RL environments, post-training, and evals.
Python★ 317↓ 6,133/moAug 9, 2026
Apache-2.0Aug 9, 2026 · metrics 2.10.0
npm
80Excellenthealth index
o-stepper/graphorin
Project Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
Go
80Excellenthealth index
realkarych/catacomb
Regression testing for Claude Code and Codex agents.
Go · Shell★ 2Jul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
attenlabs/hotato
Find what broke in your agent calls. Pin it so it never ships again. Local voice-agent call forensics and regression guards.
Python★ 1↓ 1,680/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm
78Goodhealth index
getlarge/themoltnet
Trusted context for AI agents
TypeScript · Go★ 15↓ 7,904/moJul 30, 2026
AGPL-3.0Jul 30, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
superlinear-ai/raglite
🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL
Python★ 1,200Aug 10, 2026
MPL-2.0Aug 10, 2026 · metrics 2.10.0
npm
77Goodhealth index
zernie/vigiles
Like Lighthouse for your agent harness - verify your CLAUDE.md/AGENTS.md, skills & hooks are real, then test and measure they actually work. Claude Code + Codex.
TypeScript · JavaScript★ 12↓ 4,369/moJul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
kensa-sh/kensa
Kensa turns agent traces into evals that run in CI.
Python★ 2↓ 2,171/moJul 23, 2026
Apache-2.0Jul 23, 2026 · metrics 2.10.0
Packagist
63Moderatehealth index
pestphp/pest-plugin-evals
Pest Browser Evals
PHP★ 4↓ 7,386/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
Hex
62Moderatehealth index
aryaminus/controlkeel
Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Elixir★ 10Jul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
npm
62Moderatehealth index
shulmansj/teami
Control plane for orchestrating, evaluating, and improving agent work across a company
JavaScript★ 0↓ 2,136/moJul 30, 2026
Custom licenseJul 30, 2026 · metrics 2.10.0
npm
62Moderatehealth index
spences10/my-pi
Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
TypeScript★ 88↓ 5,088/moJul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
Go
57Moderatehealth index
farazhassan/gantry
A tiny testable, Go-native agent runtime for teams that want control, conformance, and no framework lock-ins.
Go★ 1Jul 18, 2026
MITJul 18, 2026 · metrics 2.10.0
npm
54Moderatehealth index
inferock/inferock-bench
Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.
TypeScript★ 123↓ 6,119/moJul 29, 2026
Custom licenseJul 29, 2026 · metrics 2.10.0
npm
53Moderatehealth index
mykim-aus/hey-llm-you-okay
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2,332/moJul 30, 2026
MITJul 30, 2026 · metrics 2.10.0
npm
51Moderatehealth index
HolocronLab/botruntime-packages
botruntime public packages. Consumed by the botruntime platform.
TypeScript★ 0↓ 48.8K/moJul 29, 2026
No licenseJul 29, 2026 · metrics 2.10.0
npm
50Moderatehealth index
LilMGenius/paperthin
Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
Shell · JavaScript★ 119↓ 4,136/moJul 31, 2026
MITJul 31, 2026 · metrics 2.10.0
npm
47Weakhealth index
eigenpal/cli
Create, evaluate, and deploy workflows from your terminal. Agent-ready.
TypeScript★ 2↓ 4,949/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI
39Weakhealth index
splox-ai/python-sdk
Official Splox SDK
Python★ 0↓ 3,588/moJul 16, 2026
MITJul 16, 2026 · metrics 2.10.0