Alle Tags
Katalog-Tag

#evals

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

29 Einträge
Getaggt als „evals“Geordnet nach Gesundheitsindex
PyPI · npm
100AußergewöhnlichGesundheitsindex
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2K28. Aug. 2026
Eigene Lizenz28. Aug. 2026 · Metriken 2.10.0
PyPI · npm
99AußergewöhnlichGesundheitsindex
pydantic/logfire
AI observability platform for production LLM and agent systems.
Python★ 4.442↓ 14.6M/Monat27. Aug. 2026
MIT27. Aug. 2026 · Metriken 2.10.0
npm
98AußergewöhnlichGesundheitsindex
mastra-ai/mastra
Mastra is the modern TypeScript framework for AI-powered applications and agents.
TypeScript★ 26.9K↓ 61.2K/Monat5. Aug. 2026
Eigene Lizenz5. Aug. 2026 · Metriken 2.10.0
PyPI
94AußergewöhnlichGesundheitsindex
langchain-ai/deepagents
The batteries-included agent harness.
Python★ 27.3K↓ 210.2K/Monat5. Aug. 2026
MIT5. Aug. 2026 · Metriken 2.10.0
PyPI · npm
93AußergewöhnlichGesundheitsindex
harbor-framework/harbor
Framework for evaluating and improving agents
Python★ 4.69427. Aug. 2026
Apache-2.027. Aug. 2026 · Metriken 2.10.0
npm
90ExzellentGesundheitsindex
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2.069↓ 55.8K/Monat18. Juli 2026
Eigene Lizenz18. Juli 2026 · Metriken 2.10.0
PyPI
90ExzellentGesundheitsindex
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3.4876. Aug. 2026
MIT6. Aug. 2026 · Metriken 2.10.0
PyPI · npm
89ExzellentGesundheitsindex
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/Monat19. Aug. 2026
Apache-2.019. Aug. 2026 · Metriken 2.10.0
npm
84ExzellentGesundheitsindex
nearform/lastlight
Self-hostable, MIT Licensed, Enterprise AI Software Factory
TypeScript · Astro★ 18↓ 25.1K/Monat26. Juli 2026
MIT26. Juli 2026 · Metriken 2.10.0
PyPI · npm
80ExzellentGesundheitsindex
AgentOps-AI/agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Python · TypeScript★ 5.801↓ 248.9K/Monat28. Aug. 2026
MIT28. Aug. 2026 · Metriken 2.10.0
PyPI
80ExzellentGesundheitsindex
benchflow-ai/benchflow
Research infra for creating RL environments, post-training, and evals.
Python★ 317↓ 6.133/Monat9. Aug. 2026
Apache-2.09. Aug. 2026 · Metriken 2.10.0
npm
80ExzellentGesundheitsindex
o-stepper/graphorin
Project Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/Monat1. Aug. 2026
MIT1. Aug. 2026 · Metriken 2.10.0
Go
80ExzellentGesundheitsindex
realkarych/catacomb
Regression testing for Claude Code and Codex agents.
Go · Shell★ 220. Juli 2026
Apache-2.020. Juli 2026 · Metriken 2.10.0
PyPI
78GutGesundheitsindex
attenlabs/hotato
Find what broke in your agent calls. Pin it so it never ships again. Local voice-agent call forensics and regression guards.
Python★ 1↓ 1.680/Monat22. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
npm
78GutGesundheitsindex
getlarge/themoltnet
Trusted context for AI agents
TypeScript · Go★ 15↓ 7.904/Monat30. Juli 2026
AGPL-3.030. Juli 2026 · Metriken 2.10.0
PyPI
78GutGesundheitsindex
superlinear-ai/raglite
🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL
Python★ 1.20010. Aug. 2026
MPL-2.010. Aug. 2026 · Metriken 2.10.0
npm
77GutGesundheitsindex
zernie/vigiles
Like Lighthouse for your agent harness - verify your CLAUDE.md/AGENTS.md, skills & hooks are real, then test and measure they actually work. Claude Code + Codex.
TypeScript · JavaScript★ 12↓ 4.369/Monat19. Juli 2026
MIT19. Juli 2026 · Metriken 2.10.0
PyPI
73GutGesundheitsindex
kensa-sh/kensa
Kensa turns agent traces into evals that run in CI.
Python★ 2↓ 2.171/Monat23. Juli 2026
Apache-2.023. Juli 2026 · Metriken 2.10.0
Packagist
63MittelGesundheitsindex
pestphp/pest-plugin-evals
Pest Browser Evals
PHP★ 4↓ 7.386/Monat4. Aug. 2026
MIT4. Aug. 2026 · Metriken 2.10.0
Hex
62MittelGesundheitsindex
aryaminus/controlkeel
Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Elixir★ 1017. Juli 2026
Eigene Lizenz17. Juli 2026 · Metriken 2.10.0
npm
62MittelGesundheitsindex
shulmansj/teami
Control plane for orchestrating, evaluating, and improving agent work across a company
JavaScript★ 0↓ 2.136/Monat30. Juli 2026
Eigene Lizenz30. Juli 2026 · Metriken 2.10.0
npm
62MittelGesundheitsindex
spences10/my-pi
Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
TypeScript★ 88↓ 5.088/Monat18. Juli 2026
MIT18. Juli 2026 · Metriken 2.10.0
Go
57MittelGesundheitsindex
farazhassan/gantry
A tiny testable, Go-native agent runtime for teams that want control, conformance, and no framework lock-ins.
Go★ 118. Juli 2026
MIT18. Juli 2026 · Metriken 2.10.0
npm
54MittelGesundheitsindex
inferock/inferock-bench
Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.
TypeScript★ 123↓ 6.119/Monat29. Juli 2026
Eigene Lizenz29. Juli 2026 · Metriken 2.10.0
npm
53MittelGesundheitsindex
mykim-aus/hey-llm-you-okay
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2.332/Monat30. Juli 2026
MIT30. Juli 2026 · Metriken 2.10.0
npm
51MittelGesundheitsindex
HolocronLab/botruntime-packages
botruntime public packages. Consumed by the botruntime platform.
TypeScript★ 0↓ 48.8K/Monat29. Juli 2026
Keine Lizenz29. Juli 2026 · Metriken 2.10.0
npm
50MittelGesundheitsindex
LilMGenius/paperthin
Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
Shell · JavaScript★ 119↓ 4.136/Monat31. Juli 2026
MIT31. Juli 2026 · Metriken 2.10.0
npm
47SchwachGesundheitsindex
eigenpal/cli
Create, evaluate, and deploy workflows from your terminal. Agent-ready.
TypeScript★ 2↓ 4.949/Monat18. Juli 2026
Apache-2.018. Juli 2026 · Metriken 2.10.0
PyPI
39SchwachGesundheitsindex
splox-ai/python-sdk
Official Splox SDK
Python★ 0↓ 3.588/Monat16. Juli 2026
MIT16. Juli 2026 · Metriken 2.10.0