Усі теги
Тег каталогу

#evaluation

Усі репозиторії публічного реєстру з цим тегом — із тем GitHub або ключових слів, які публікують їхні реєстри пакетів. Здоров'я вимірюється за тією ж версіонованою методологією, що й решта реєстру.

58 записів
З тегом «evaluation»Упорядковано за індексом здоров'я
PyPI · npm
99Винятковийіндекс здоров'я
langchain-ai/langsmith-sdk
LangSmith Client SDK Implementations
Python · TypeScript★ 1 039↓ 148M/міс27 серп. 2026 р.
MIT27 серп. 2026 р. · метрики 2.10.0
PyPI · npm
98Винятковийіндекс здоров'я
Agenta-AI/agenta
Agenta is a workspace where you and your team build agents and automations.
TypeScript · Python★ 4 573↓ 17.6K/міс28 серп. 2026 р.
Власна ліцензія28 серп. 2026 р. · метрики 2.10.0
npm · Go
98Винятковийіндекс здоров'я
langwatch/langwatch
The platform for LLM evaluations and AI agent testing
TypeScript★ 3 487↓ 1 393/міс13 серп. 2026 р.
Apache-2.013 серп. 2026 р. · метрики 2.10.0
npm · PyPI
96Винятковийіндекс здоров'я
comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/міс5 серп. 2026 р.
Apache-2.05 серп. 2026 р. · метрики 2.10.0
npm
96Винятковийіндекс здоров'я
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/міс5 серп. 2026 р.
MIT5 серп. 2026 р. · метрики 2.10.0
Go · PyPI · npm
95Винятковийіндекс здоров'я
Tencent/WeKnora
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 19.4K5 серп. 2026 р.
Власна ліцензія5 серп. 2026 р. · метрики 2.10.0
PyPI
95Винятковийіндекс здоров'я
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 3 36219 лип. 2026 р.
Apache-2.019 лип. 2026 р. · метрики 2.10.0
Go
95Винятковийіндекс здоров'я
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 1 56119 лип. 2026 р.
Apache-2.019 лип. 2026 р. · метрики 2.10.0
PyPI
94Винятковийіндекс здоров'я
langchain-ai/deepagents
The batteries-included agent harness.
Python★ 27.3K↓ 210.2K/міс5 серп. 2026 р.
MIT5 серп. 2026 р. · метрики 2.10.0
npm
94Винятковийіндекс здоров'я
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5K5 серп. 2026 р.
Власна ліцензія5 серп. 2026 р. · метрики 2.10.0
PyPI
94Винятковийіндекс здоров'я
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Python · TypeScript★ 3 172↓ 67.9K/міс1 серп. 2026 р.
Apache-2.01 серп. 2026 р. · метрики 2.10.0
PyPI · npm
93Винятковийіндекс здоров'я
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1 055↓ 406.4K/міс18 лип. 2026 р.
Apache-2.018 лип. 2026 р. · метрики 2.10.0
npm · PyPI
91Відміннийіндекс здоров'я
joshuaswarren/remnic
Open-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.
TypeScript★ 176↓ 203.7K/міс22 серп. 2026 р.
MIT22 серп. 2026 р. · метрики 2.10.0
npm
90Відміннийіндекс здоров'я
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2 069↓ 55.8K/міс18 лип. 2026 р.
Власна ліцензія18 лип. 2026 р. · метрики 2.10.0
npm · PyPI
90Відміннийіндекс здоров'я
Marker-Inc-Korea/AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4 963↓ 309/міс2 серп. 2026 р.
Власна ліцензія2 серп. 2026 р. · метрики 2.10.0
PyPI · npm
89Відміннийіндекс здоров'я
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/міс19 серп. 2026 р.
Apache-2.019 серп. 2026 р. · метрики 2.10.0
PyPI · Go · npm
89Відміннийіндекс здоров'я
gooddata/gooddata-python-sdk
GoodData Cloud Python SDK
Python★ 35↓ 153.9K/міс22 серп. 2026 р.
Власна ліцензія22 серп. 2026 р. · метрики 2.10.0
PyPI
89Відміннийіндекс здоров'я
vibrantlabsai/ragas
Supercharge Your LLM Application Evaluations 🚀
Python · Jupyter Notebook★ 15.2K↓ 1.6M/міс8 серп. 2026 р.
Apache-2.08 серп. 2026 р. · метрики 2.10.0
PyPI · npm
88Відміннийіндекс здоров'я
hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Python★ 124↓ 2 207/міс26 лип. 2026 р.
Apache-2.026 лип. 2026 р. · метрики 2.10.0
PyPI
86Відміннийіндекс здоров'я
BrainLesion/panoptica
panoptica -- instance-wise evaluation of 3D semantic and instance segmentation maps
Python · Jupyter Notebook★ 3331 лип. 2026 р.
Apache-2.031 лип. 2026 р. · метрики 2.10.0
PyPI
86Відміннийіндекс здоров'я
MichaelGrupp/evo
Python package for the evaluation of odometry and SLAM
Python★ 4 309↓ 210.7K/міс28 серп. 2026 р.
GPL-3.028 серп. 2026 р. · метрики 2.10.0
NuGet
86Відміннийіндекс здоров'я
asc-community/AngouriMath
Open-source cross-platform symbolic algebra library for C# and F#. Can be used for both production and research purposes.
C#★ 82723 серп. 2026 р.
MIT23 серп. 2026 р. · метрики 2.10.0
PyPI · npm
86Відміннийіндекс здоров'я
robocurve/inspect-robots
Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.
Python★ 269↓ 5 131/міс6 вер. 2026 р.
MIT6 вер. 2026 р. · метрики 2.10.0
npm · crates.io · PyPI
84Відміннийіндекс здоров'я
PSU3D0/formualizer
Embeddable spreadsheet engine - parse, evaluate & mutate Excel workbooks from Rust, Python, or the browser. Arrow-powered, 400+ functions.
Rust★ 177↓ 10.9K/міс5 вер. 2026 р.
Apache-2.05 вер. 2026 р. · метрики 2.10.0
npm · Go · crates.io +2
81Відміннийіндекс здоров'я
ops-ai/Toggly.FeatureManagement
Enables teams to release software faster and safer, and with better results.
TypeScript · C#★ 5↓ 3 643/міс3 вер. 2026 р.
MIT3 вер. 2026 р. · метрики 2.10.0
npm · PyPI
80Відміннийіндекс здоров'я
aikdna/kdna
KDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/міс22 лип. 2026 р.
Apache-2.022 лип. 2026 р. · метрики 2.10.0
npm
80Відміннийіндекс здоров'я
o-stepper/graphorin
Project Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/міс1 серп. 2026 р.
MIT1 серп. 2026 р. · метрики 2.10.0
PyPI
78Добрийіндекс здоров'я
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 2 46518 лип. 2026 р.
Apache-2.018 лип. 2026 р. · метрики 2.10.0
PyPI
73Добрийіндекс здоров'я
mjpost/sacrebleu
Reference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1 259↓ 4.3M/міс27 серп. 2026 р.
Apache-2.027 серп. 2026 р. · метрики 2.10.0
npm
71Добрийіндекс здоров'я
JudgmentLabs/judgeval-js
The open source post-building layer for agents.
TypeScript★ 5↓ 92.5K/міс1 серп. 2026 р.
Без ліцензії1 серп. 2026 р. · метрики 2.10.0