Todas las etiquetas
Etiqueta del catálogo

#evals

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

14 registros
Con la etiqueta «evals»Ordenado por índice de salud
PyPI · npm
89Excelenteíndice de salud
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript · Jupyter Notebook★ 10.6K17 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
PyPI · npm
86Excelenteíndice de salud
pydantic/logfire
AI observability platform for production LLM and agent systems.
Python★ 4380↓ 23.5M/mes18 jul 2026
MIT18 jul 2026 · métricas 1.13.0
npm
84Buenoíndice de salud
mastra-ai/mastra
Mastra is the modern TypeScript framework for AI-powered applications and agents.
TypeScript★ 26.2K↓ 0/mes14 jul 2026
Licencia propia14 jul 2026 · métricas 1.13.0
PyPI · npm
76Buenoíndice de salud
harbor-framework/harbor
Framework for evaluating and improving agents
Python★ 3333↓ 9.3M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
npm
75Buenoíndice de salud
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2069↓ 55.8K/mes18 jul 2026
Licencia propia18 jul 2026 · métricas 1.13.0
PyPI · npm
71Buenoíndice de salud
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 571621 jul 2026
MIT21 jul 2026 · métricas 1.13.0
Go
67Moderadoíndice de salud
realkarych/catacomb
Regression testing for Claude Code and Codex agents.
Go · Shell★ 220 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0
npm
66Moderadoíndice de salud
zernie/vigiles
Like Lighthouse for your agent harness - verify your CLAUDE.md/AGENTS.md, skills & hooks are real, then test and measure they actually work. Claude Code + Codex.
TypeScript · JavaScript★ 12↓ 4369/mes19 jul 2026
MIT19 jul 2026 · métricas 1.13.0
PyPI
64Moderadoíndice de salud
kensa-sh/kensa
Kensa turns agent traces into evals that run in CI.
Python★ 2↓ 2171/mes23 jul 2026
Apache-2.023 jul 2026 · métricas 1.13.0
Hex
58Moderadoíndice de salud
aryaminus/controlkeel
Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Elixir★ 1017 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
npm
58Moderadoíndice de salud
spences10/my-pi
Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
TypeScript★ 88↓ 5088/mes18 jul 2026
MIT18 jul 2026 · métricas 1.13.0
Go
55Moderadoíndice de salud
farazhassan/gantry
A tiny testable, Go-native agent runtime for teams that want control, conformance, and no framework lock-ins.
Go★ 118 jul 2026
MIT18 jul 2026 · métricas 1.13.0
npm
46En riesgoíndice de salud
eigenpal/cli
Create, evaluate, and deploy workflows from your terminal. Agent-ready.
TypeScript★ 2↓ 4949/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
43En riesgoíndice de salud
splox-ai/python-sdk
Official Splox SDK
Python★ 0↓ 3588/mes16 jul 2026
MIT16 jul 2026 · métricas 1.13.0