Todas las etiquetas
Etiqueta del catálogo

#evals

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

29 registros
Con la etiqueta «evals»Ordenado por índice de salud
PyPI · npm
100Excepcionalíndice de salud
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2K28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI · npm
99Excepcionalíndice de salud
pydantic/logfire
AI observability platform for production LLM and agent systems.
Python★ 4442↓ 14.6M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
npm
98Excepcionalíndice de salud
mastra-ai/mastra
Mastra is the modern TypeScript framework for AI-powered applications and agents.
TypeScript★ 26.9K↓ 61.2K/mes5 ago 2026
Licencia propia5 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
langchain-ai/deepagents
The batteries-included agent harness.
Python★ 27.3K↓ 210.2K/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
PyPI · npm
93Excepcionalíndice de salud
harbor-framework/harbor
Framework for evaluating and improving agents
Python★ 469427 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
npm
90Excelenteíndice de salud
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2069↓ 55.8K/mes18 jul 2026
Licencia propia18 jul 2026 · métricas 2.10.0
PyPI
90Excelenteíndice de salud
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 34876 ago 2026
MIT6 ago 2026 · métricas 2.10.0
PyPI · npm
89Excelenteíndice de salud
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/mes19 ago 2026
Apache-2.019 ago 2026 · métricas 2.10.0
npm
84Excelenteíndice de salud
nearform/lastlight
Self-hostable, MIT Licensed, Enterprise AI Software Factory
TypeScript · Astro★ 18↓ 25.1K/mes26 jul 2026
MIT26 jul 2026 · métricas 2.10.0
PyPI · npm
80Excelenteíndice de salud
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5801↓ 248.9K/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
PyPI
80Excelenteíndice de salud
benchflow-ai/benchflow
Research infra for creating RL environments, post-training, and evals.
Python★ 317↓ 6133/mes9 ago 2026
Apache-2.09 ago 2026 · métricas 2.10.0
npm
80Excelenteíndice de salud
o-stepper/graphorin
Project Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
Go
80Excelenteíndice de salud
realkarych/catacomb
Regression testing for Claude Code and Codex agents.
Go · Shell★ 220 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
attenlabs/hotato
Find what broke in your agent calls. Pin it so it never ships again. Local voice-agent call forensics and regression guards.
Python★ 1↓ 1680/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm
78Buenoíndice de salud
getlarge/themoltnet
Trusted context for AI agents
TypeScript · Go★ 15↓ 7904/mes30 jul 2026
AGPL-3.030 jul 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
superlinear-ai/raglite
🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL
Python★ 120010 ago 2026
MPL-2.010 ago 2026 · métricas 2.10.0
npm
77Buenoíndice de salud
zernie/vigiles
Like Lighthouse for your agent harness - verify your CLAUDE.md/AGENTS.md, skills & hooks are real, then test and measure they actually work. Claude Code + Codex.
TypeScript · JavaScript★ 12↓ 4369/mes19 jul 2026
MIT19 jul 2026 · métricas 2.10.0
PyPI
73Buenoíndice de salud
kensa-sh/kensa
Kensa turns agent traces into evals that run in CI.
Python★ 2↓ 2171/mes23 jul 2026
Apache-2.023 jul 2026 · métricas 2.10.0
Packagist
63Moderadoíndice de salud
pestphp/pest-plugin-evals
Pest Browser Evals
PHP★ 4↓ 7386/mes4 ago 2026
MIT4 ago 2026 · métricas 2.10.0
Hex
62Moderadoíndice de salud
aryaminus/controlkeel
Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Elixir★ 1017 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
npm
62Moderadoíndice de salud
shulmansj/teami
Control plane for orchestrating, evaluating, and improving agent work across a company
JavaScript★ 0↓ 2136/mes30 jul 2026
Licencia propia30 jul 2026 · métricas 2.10.0
npm
62Moderadoíndice de salud
spences10/my-pi
Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
TypeScript★ 88↓ 5088/mes18 jul 2026
MIT18 jul 2026 · métricas 2.10.0
Go
57Moderadoíndice de salud
farazhassan/gantry
A tiny testable, Go-native agent runtime for teams that want control, conformance, and no framework lock-ins.
Go★ 118 jul 2026
MIT18 jul 2026 · métricas 2.10.0
npm
54Moderadoíndice de salud
inferock/inferock-bench
Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.
TypeScript★ 123↓ 6119/mes29 jul 2026
Licencia propia29 jul 2026 · métricas 2.10.0
npm
53Moderadoíndice de salud
mykim-aus/hey-llm-you-okay
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2332/mes30 jul 2026
MIT30 jul 2026 · métricas 2.10.0
npm
51Moderadoíndice de salud
HolocronLab/botruntime-packages
botruntime public packages. Consumed by the botruntime platform.
TypeScript★ 0↓ 48.8K/mes29 jul 2026
Sin licencia29 jul 2026 · métricas 2.10.0
npm
50Moderadoíndice de salud
LilMGenius/paperthin
Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
Shell · JavaScript★ 119↓ 4136/mes31 jul 2026
MIT31 jul 2026 · métricas 2.10.0
npm
47Débilíndice de salud
eigenpal/cli
Create, evaluate, and deploy workflows from your terminal. Agent-ready.
TypeScript★ 2↓ 4949/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
PyPI
39Débilíndice de salud
splox-ai/python-sdk
Official Splox SDK
Python★ 0↓ 3588/mes16 jul 2026
MIT16 jul 2026 · métricas 2.10.0