All tags
Catalogue tag

#evaluation

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

58 records
Tagged “evaluation”Ranked by health index
PyPI · npm
99Exceptionalhealth index
langchain-ai/langsmith-sdk
LangSmith Client SDK Implementations
Python · TypeScript★ 1,039↓ 148M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI · npm
98Exceptionalhealth index
Agenta-AI/agenta
Agenta is a workspace where you and your team build agents and automations.
TypeScript · Python★ 4,573↓ 17.6K/moAug 28, 2026
Custom licenseAug 28, 2026 · metrics 2.10.0
npm · Go
98Exceptionalhealth index
langwatch/langwatch
The platform for LLM evaluations and AI agent testing
TypeScript★ 3,487↓ 1,393/moAug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
npm · PyPI
96Exceptionalhealth index
comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
npm
96Exceptionalhealth index
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
Go · PyPI · npm
95Exceptionalhealth index
Tencent/WeKnora
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 19.4KAug 5, 2026
Custom licenseAug 5, 2026 · metrics 2.10.0
PyPI
95Exceptionalhealth index
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 3,362Jul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 2.10.0
Go
95Exceptionalhealth index
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 1,561Jul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
langchain-ai/deepagents
The batteries-included agent harness.
Python★ 27.3K↓ 210.2K/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
npm
94Exceptionalhealth index
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5KAug 5, 2026
Custom licenseAug 5, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Python · TypeScript★ 3,172↓ 67.9K/moAug 1, 2026
Apache-2.0Aug 1, 2026 · metrics 2.10.0
PyPI · npm
93Exceptionalhealth index
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1,055↓ 406.4K/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
npm · PyPI
91Excellenthealth index
joshuaswarren/remnic
Open-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.
TypeScript★ 176↓ 203.7K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
npm
90Excellenthealth index
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2,069↓ 55.8K/moJul 18, 2026
Custom licenseJul 18, 2026 · metrics 2.10.0
npm · PyPI
90Excellenthealth index
Marker-Inc-Korea/AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4,963↓ 309/moAug 2, 2026
Custom licenseAug 2, 2026 · metrics 2.10.0
PyPI · npm
89Excellenthealth index
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/moAug 19, 2026
Apache-2.0Aug 19, 2026 · metrics 2.10.0
PyPI · Go · npm
89Excellenthealth index
gooddata/gooddata-python-sdk
GoodData Cloud Python SDK
Python★ 35↓ 153.9K/moAug 22, 2026
Custom licenseAug 22, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
vibrantlabsai/ragas
Supercharge Your LLM Application Evaluations 🚀
Python · Jupyter Notebook★ 15.2K↓ 1.6M/moAug 8, 2026
Apache-2.0Aug 8, 2026 · metrics 2.10.0
PyPI · npm
88Excellenthealth index
hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Python★ 124↓ 2,207/moJul 26, 2026
Apache-2.0Jul 26, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
BrainLesion/panoptica
panoptica -- instance-wise evaluation of 3D semantic and instance segmentation maps
Python · Jupyter Notebook★ 33Jul 31, 2026
Apache-2.0Jul 31, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
MichaelGrupp/evo
Python package for the evaluation of odometry and SLAM
Python★ 4,309↓ 210.7K/moAug 28, 2026
GPL-3.0Aug 28, 2026 · metrics 2.10.0
NuGet
86Excellenthealth index
asc-community/AngouriMath
Open-source cross-platform symbolic algebra library for C# and F#. Can be used for both production and research purposes.
C#★ 827Aug 23, 2026
MITAug 23, 2026 · metrics 2.10.0
PyPI · npm
86Excellenthealth index
robocurve/inspect-robots
Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.
Python★ 269↓ 5,131/moSep 6, 2026
MITSep 6, 2026 · metrics 2.10.0
npm · crates.io · PyPI
84Excellenthealth index
PSU3D0/formualizer
Embeddable spreadsheet engine - parse, evaluate & mutate Excel workbooks from Rust, Python, or the browser. Arrow-powered, 400+ functions.
Rust★ 177↓ 10.9K/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
npm · Go · crates.io +2
81Excellenthealth index
ops-ai/Toggly.FeatureManagement
Enables teams to release software faster and safer, and with better results.
TypeScript · C#★ 5↓ 3,643/moSep 3, 2026
MITSep 3, 2026 · metrics 2.10.0
npm · PyPI
80Excellenthealth index
aikdna/kdna
KDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/moJul 22, 2026
Apache-2.0Jul 22, 2026 · metrics 2.10.0
npm
80Excellenthealth index
o-stepper/graphorin
Project Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/moAug 1, 2026
MITAug 1, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 2,465Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
mjpost/sacrebleu
Reference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1,259↓ 4.3M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
npm
71Goodhealth index
JudgmentLabs/judgeval-js
The open source post-building layer for agents.
TypeScript★ 5↓ 92.5K/moAug 1, 2026
No licenseAug 1, 2026 · metrics 2.10.0