All tags
Catalogue tag

#evaluation

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

22 records
Tagged “evaluation”Ranked by health index
npm · PyPI
88Excellenthealth index
langchain-ai/langsmith-sdk
LangSmith Client SDK Implementations
Python · TypeScript★ 971↓ 22.6M/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI · npm
83Goodhealth index
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.1K↓ 40.1M/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
82Goodhealth index
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 3,362Jul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0
npm
81Goodhealth index
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 31.3KJul 17, 2026
Custom licenseJul 17, 2026 · metrics 1.13.0
Go
81Goodhealth index
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 1,561Jul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 1.13.0
npm
80Goodhealth index
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.4K↓ 1.7M/moJul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI · npm
79Goodhealth index
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1,055↓ 406.4K/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
Go · npm · PyPI
77Goodhealth index
Tencent/WeKnora
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 18.7KJul 21, 2026
Custom licenseJul 21, 2026 · metrics 1.13.0
npm
75Goodhealth index
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2,069↓ 55.8K/moJul 18, 2026
Custom licenseJul 18, 2026 · metrics 1.13.0
npm · PyPI
68Moderatehealth index
aikdna/kdna
KDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/moJul 22, 2026
Apache-2.0Jul 22, 2026 · metrics 1.13.0
PyPI
67Moderatehealth index
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 2,465Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
PyPI
66Moderatehealth index
robocurve/inspect-robots
Evaluation framework for VLA / physical-AI models: define a benchmark once, run any policy on any robot or sim. (The Inspect AI for robotics.)
Python★ 5Jul 15, 2026
MITJul 15, 2026 · metrics 1.13.0
PyPI
60Moderatehealth index
mjpost/sacrebleu
Reference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1,253↓ 3.9M/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 1.13.0
Go · npm
58Moderatehealth index
starksv/windows-iso-downloader
Download official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript★ 7↓ 0/moJul 14, 2026
MITJul 14, 2026 · metrics 1.13.0
PyPI
55Moderatehealth index
EpsilabAI/epsilab-python
The official Python library for the Epsilab API
Python★ 0↓ 2,706/moJul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 1.13.0
PyPI
52Moderatehealth index
attenlabs/hotato
Conversation QA for voice agents. Catch the calls that pass every text check but talk over the caller, skip a disclosure, or claim a task that never happened. Self-hosted, offline, MIT.
HTML★ 0↓ 2,575/moJul 13, 2026
MITJul 13, 2026 · metrics 1.13.0
PyPI
51Moderatehealth index
lazily-hub/lazily-py
A Python library for lazy evaluation with context caching.
Python★ 1Jul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 1.13.0
PyPI
50Moderatehealth index
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2,908/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
PyPI
47At riskhealth index
davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge
HTML · Python★ 66↓ 0/moJul 14, 2026
No licenseJul 14, 2026 · metrics 1.13.0
npm
40At riskhealth index
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 316M/moJul 22, 2026
MITJul 22, 2026 · metrics 1.13.0
PyPI
39At riskhealth index
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 2Jul 17, 2026
No licenseJul 17, 2026 · metrics 1.13.0
PyPI
26Criticalhealth index
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 607Jul 21, 2026
Custom licenseJul 21, 2026 · metrics 1.13.0