Todas las etiquetas
Etiqueta del catálogo

#evaluation

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

58 registros
Con la etiqueta «evaluation»Ordenado por índice de salud
PyPI · npm
99Excepcionalíndice de salud
langchain-ai/langsmith-sdk
LangSmith Client SDK Implementations
Python · TypeScript★ 1039↓ 148M/mes27 ago 2026
MIT27 ago 2026 · métricas 2.10.0
PyPI · npm
98Excepcionalíndice de salud
Agenta-AI/agenta
Agenta is a workspace where you and your team build agents and automations.
TypeScript · Python★ 4573↓ 17.6K/mes28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
npm · Go
98Excepcionalíndice de salud
langwatch/langwatch
The platform for LLM evaluations and AI agent testing
TypeScript★ 3487↓ 1393/mes13 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
npm · PyPI
96Excepcionalíndice de salud
comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
npm
96Excepcionalíndice de salud
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
Go · PyPI · npm
95Excepcionalíndice de salud
Tencent/WeKnora
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 19.4K5 ago 2026
Licencia propia5 ago 2026 · métricas 2.10.0
PyPI
95Excepcionalíndice de salud
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 336219 jul 2026
Apache-2.019 jul 2026 · métricas 2.10.0
Go
95Excepcionalíndice de salud
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 156119 jul 2026
Apache-2.019 jul 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
langchain-ai/deepagents
The batteries-included agent harness.
Python★ 27.3K↓ 210.2K/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
npm
94Excepcionalíndice de salud
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5K5 ago 2026
Licencia propia5 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Python · TypeScript★ 3172↓ 67.9K/mes1 ago 2026
Apache-2.01 ago 2026 · métricas 2.10.0
PyPI · npm
93Excepcionalíndice de salud
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1055↓ 406.4K/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
npm · PyPI
91Excelenteíndice de salud
joshuaswarren/remnic
Open-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.
TypeScript★ 176↓ 203.7K/mes22 ago 2026
MIT22 ago 2026 · métricas 2.10.0
npm
90Excelenteíndice de salud
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2069↓ 55.8K/mes18 jul 2026
Licencia propia18 jul 2026 · métricas 2.10.0
npm · PyPI
90Excelenteíndice de salud
Marker-Inc-Korea/AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4963↓ 309/mes2 ago 2026
Licencia propia2 ago 2026 · métricas 2.10.0
PyPI · npm
89Excelenteíndice de salud
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/mes19 ago 2026
Apache-2.019 ago 2026 · métricas 2.10.0
PyPI · Go · npm
89Excelenteíndice de salud
gooddata/gooddata-python-sdk
GoodData Cloud Python SDK
Python★ 35↓ 153.9K/mes22 ago 2026
Licencia propia22 ago 2026 · métricas 2.10.0
PyPI
89Excelenteíndice de salud
vibrantlabsai/ragas
Supercharge Your LLM Application Evaluations 🚀
Python · Jupyter Notebook★ 15.2K↓ 1.6M/mes8 ago 2026
Apache-2.08 ago 2026 · métricas 2.10.0
PyPI · npm
88Excelenteíndice de salud
hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Python★ 124↓ 2207/mes26 jul 2026
Apache-2.026 jul 2026 · métricas 2.10.0
PyPI
86Excelenteíndice de salud
BrainLesion/panoptica
panoptica -- instance-wise evaluation of 3D semantic and instance segmentation maps
Python · Jupyter Notebook★ 3331 jul 2026
Apache-2.031 jul 2026 · métricas 2.10.0
PyPI
86Excelenteíndice de salud
MichaelGrupp/evo
Python package for the evaluation of odometry and SLAM
Python★ 4309↓ 210.7K/mes28 ago 2026
GPL-3.028 ago 2026 · métricas 2.10.0
NuGet
86Excelenteíndice de salud
asc-community/AngouriMath
Open-source cross-platform symbolic algebra library for C# and F#. Can be used for both production and research purposes.
C#★ 82723 ago 2026
MIT23 ago 2026 · métricas 2.10.0
PyPI · npm
86Excelenteíndice de salud
robocurve/inspect-robots
Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.
Python★ 269↓ 5131/mes6 sept 2026
MIT6 sept 2026 · métricas 2.10.0
npm · crates.io · PyPI
84Excelenteíndice de salud
PSU3D0/formualizer
Embeddable spreadsheet engine - parse, evaluate & mutate Excel workbooks from Rust, Python, or the browser. Arrow-powered, 400+ functions.
Rust★ 177↓ 10.9K/mes5 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
npm · Go · crates.io +2
81Excelenteíndice de salud
ops-ai/Toggly.FeatureManagement
Enables teams to release software faster and safer, and with better results.
TypeScript · C#★ 5↓ 3643/mes3 sept 2026
MIT3 sept 2026 · métricas 2.10.0
npm · PyPI
80Excelenteíndice de salud
aikdna/kdna
KDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/mes22 jul 2026
Apache-2.022 jul 2026 · métricas 2.10.0
npm
80Excelenteíndice de salud
o-stepper/graphorin
Project Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/mes1 ago 2026
MIT1 ago 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 246518 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
PyPI
73Buenoíndice de salud
mjpost/sacrebleu
Reference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1259↓ 4.3M/mes27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
npm
71Buenoíndice de salud
JudgmentLabs/judgeval-js
The open source post-building layer for agents.
TypeScript★ 5↓ 92.5K/mes1 ago 2026
Sin licencia1 ago 2026 · métricas 2.10.0