Alle Tags
Katalog-Tag

#evaluation

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

58 Einträge
Getaggt als „evaluation“Geordnet nach Gesundheitsindex
PyPI · npm
99AußergewöhnlichGesundheitsindex
langchain-ai/langsmith-sdk
LangSmith Client SDK Implementations
Python · TypeScript★ 1.039↓ 148M/Monat27. Aug. 2026
MIT27. Aug. 2026 · Metriken 2.10.0
PyPI · npm
98AußergewöhnlichGesundheitsindex
Agenta-AI/agenta
Agenta is a workspace where you and your team build agents and automations.
TypeScript · Python★ 4.573↓ 17.6K/Monat28. Aug. 2026
Eigene Lizenz28. Aug. 2026 · Metriken 2.10.0
npm · Go
98AußergewöhnlichGesundheitsindex
langwatch/langwatch
The platform for LLM evaluations and AI agent testing
TypeScript★ 3.487↓ 1.393/Monat13. Aug. 2026
Apache-2.013. Aug. 2026 · Metriken 2.10.0
npm · PyPI
96AußergewöhnlichGesundheitsindex
comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/Monat5. Aug. 2026
Apache-2.05. Aug. 2026 · Metriken 2.10.0
npm
96AußergewöhnlichGesundheitsindex
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/Monat5. Aug. 2026
MIT5. Aug. 2026 · Metriken 2.10.0
Go · PyPI · npm
95AußergewöhnlichGesundheitsindex
Tencent/WeKnora
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 19.4K5. Aug. 2026
Eigene Lizenz5. Aug. 2026 · Metriken 2.10.0
PyPI
95AußergewöhnlichGesundheitsindex
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 3.36219. Juli 2026
Apache-2.019. Juli 2026 · Metriken 2.10.0
Go
95AußergewöhnlichGesundheitsindex
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 1.56119. Juli 2026
Apache-2.019. Juli 2026 · Metriken 2.10.0
PyPI
94AußergewöhnlichGesundheitsindex
langchain-ai/deepagents
The batteries-included agent harness.
Python★ 27.3K↓ 210.2K/Monat5. Aug. 2026
MIT5. Aug. 2026 · Metriken 2.10.0
npm
94AußergewöhnlichGesundheitsindex
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5K5. Aug. 2026
Eigene Lizenz5. Aug. 2026 · Metriken 2.10.0
PyPI
94AußergewöhnlichGesundheitsindex
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Python · TypeScript★ 3.172↓ 67.9K/Monat1. Aug. 2026
Apache-2.01. Aug. 2026 · Metriken 2.10.0
PyPI · npm
93AußergewöhnlichGesundheitsindex
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1.055↓ 406.4K/Monat18. Juli 2026
Apache-2.018. Juli 2026 · Metriken 2.10.0
npm · PyPI
91ExzellentGesundheitsindex
joshuaswarren/remnic
Open-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.
TypeScript★ 176↓ 203.7K/Monat22. Aug. 2026
MIT22. Aug. 2026 · Metriken 2.10.0
npm
90ExzellentGesundheitsindex
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2.069↓ 55.8K/Monat18. Juli 2026
Eigene Lizenz18. Juli 2026 · Metriken 2.10.0
npm · PyPI
90ExzellentGesundheitsindex
Marker-Inc-Korea/AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4.963↓ 309/Monat2. Aug. 2026
Eigene Lizenz2. Aug. 2026 · Metriken 2.10.0
PyPI · npm
89ExzellentGesundheitsindex
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/Monat19. Aug. 2026
Apache-2.019. Aug. 2026 · Metriken 2.10.0
PyPI · Go · npm
89ExzellentGesundheitsindex
gooddata/gooddata-python-sdk
GoodData Cloud Python SDK
Python★ 35↓ 153.9K/Monat22. Aug. 2026
Eigene Lizenz22. Aug. 2026 · Metriken 2.10.0
PyPI
89ExzellentGesundheitsindex
vibrantlabsai/ragas
Supercharge Your LLM Application Evaluations 🚀
Python · Jupyter Notebook★ 15.2K↓ 1.6M/Monat8. Aug. 2026
Apache-2.08. Aug. 2026 · Metriken 2.10.0
PyPI · npm
88ExzellentGesundheitsindex
hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Python★ 124↓ 2.207/Monat26. Juli 2026
Apache-2.026. Juli 2026 · Metriken 2.10.0
PyPI
86ExzellentGesundheitsindex
BrainLesion/panoptica
panoptica -- instance-wise evaluation of 3D semantic and instance segmentation maps
Python · Jupyter Notebook★ 3331. Juli 2026
Apache-2.031. Juli 2026 · Metriken 2.10.0
PyPI
86ExzellentGesundheitsindex
MichaelGrupp/evo
Python package for the evaluation of odometry and SLAM
Python★ 4.309↓ 210.7K/Monat28. Aug. 2026
GPL-3.028. Aug. 2026 · Metriken 2.10.0
NuGet
86ExzellentGesundheitsindex
asc-community/AngouriMath
Open-source cross-platform symbolic algebra library for C# and F#. Can be used for both production and research purposes.
C#★ 82723. Aug. 2026
MIT23. Aug. 2026 · Metriken 2.10.0
PyPI · npm
86ExzellentGesundheitsindex
robocurve/inspect-robots
Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.
Python★ 269↓ 5.131/Monat6. Sept. 2026
MIT6. Sept. 2026 · Metriken 2.10.0
npm · crates.io · PyPI
84ExzellentGesundheitsindex
PSU3D0/formualizer
Embeddable spreadsheet engine - parse, evaluate & mutate Excel workbooks from Rust, Python, or the browser. Arrow-powered, 400+ functions.
Rust★ 177↓ 10.9K/Monat5. Sept. 2026
Apache-2.05. Sept. 2026 · Metriken 2.10.0
npm · Go · crates.io +2
81ExzellentGesundheitsindex
ops-ai/Toggly.FeatureManagement
Enables teams to release software faster and safer, and with better results.
TypeScript · C#★ 5↓ 3.643/Monat3. Sept. 2026
MIT3. Sept. 2026 · Metriken 2.10.0
npm · PyPI
80ExzellentGesundheitsindex
aikdna/kdna
KDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/Monat22. Juli 2026
Apache-2.022. Juli 2026 · Metriken 2.10.0
npm
80ExzellentGesundheitsindex
o-stepper/graphorin
Project Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/Monat1. Aug. 2026
MIT1. Aug. 2026 · Metriken 2.10.0
PyPI
78GutGesundheitsindex
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 2.46518. Juli 2026
Apache-2.018. Juli 2026 · Metriken 2.10.0
PyPI
73GutGesundheitsindex
mjpost/sacrebleu
Reference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1.259↓ 4.3M/Monat27. Aug. 2026
Apache-2.027. Aug. 2026 · Metriken 2.10.0
npm
71GutGesundheitsindex
JudgmentLabs/judgeval-js
The open source post-building layer for agents.
TypeScript★ 5↓ 92.5K/Monat1. Aug. 2026
Keine Lizenz1. Aug. 2026 · Metriken 2.10.0