Alle Tags
Katalog-Tag

#evaluation

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

22 Einträge
Getaggt als „evaluation“Geordnet nach Gesundheitsindex
npm · PyPI
88ExzellentGesundheitsindex
langchain-ai/langsmith-sdk
LangSmith Client SDK Implementations
Python · TypeScript★ 971↓ 22.6M/Monat17. Juli 2026
MIT17. Juli 2026 · Metriken 1.13.0
PyPI · npm
83GutGesundheitsindex
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.1K↓ 40.1M/Monat18. Juli 2026
Apache-2.018. Juli 2026 · Metriken 1.13.0
PyPI
82GutGesundheitsindex
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 3.36219. Juli 2026
Apache-2.019. Juli 2026 · Metriken 1.13.0
npm
81GutGesundheitsindex
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 31.3K17. Juli 2026
Eigene Lizenz17. Juli 2026 · Metriken 1.13.0
Go
81GutGesundheitsindex
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 1.56119. Juli 2026
Apache-2.019. Juli 2026 · Metriken 1.13.0
npm
80GutGesundheitsindex
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.4K↓ 1.7M/Monat17. Juli 2026
MIT17. Juli 2026 · Metriken 1.13.0
PyPI · npm
79GutGesundheitsindex
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1.055↓ 406.4K/Monat18. Juli 2026
Apache-2.018. Juli 2026 · Metriken 1.13.0
Go · npm · PyPI
77GutGesundheitsindex
Tencent/WeKnora
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 18.7K21. Juli 2026
Eigene Lizenz21. Juli 2026 · Metriken 1.13.0
npm
75GutGesundheitsindex
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2.069↓ 55.8K/Monat18. Juli 2026
Eigene Lizenz18. Juli 2026 · Metriken 1.13.0
npm · PyPI
68MittelGesundheitsindex
aikdna/kdna
KDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/Monat22. Juli 2026
Apache-2.022. Juli 2026 · Metriken 1.13.0
PyPI
67MittelGesundheitsindex
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 2.46518. Juli 2026
Apache-2.018. Juli 2026 · Metriken 1.13.0
PyPI
66MittelGesundheitsindex
robocurve/inspect-robots
Evaluation framework for VLA / physical-AI models: define a benchmark once, run any policy on any robot or sim. (The Inspect AI for robotics.)
Python★ 515. Juli 2026
MIT15. Juli 2026 · Metriken 1.13.0
PyPI
60MittelGesundheitsindex
mjpost/sacrebleu
Reference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1.253↓ 3.9M/Monat18. Juli 2026
Apache-2.018. Juli 2026 · Metriken 1.13.0
Go · npm
58MittelGesundheitsindex
starksv/windows-iso-downloader
Download official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript★ 7↓ 0/Monat14. Juli 2026
MIT14. Juli 2026 · Metriken 1.13.0
PyPI
55MittelGesundheitsindex
EpsilabAI/epsilab-python
The official Python library for the Epsilab API
Python★ 0↓ 2.706/Monat21. Juli 2026
Apache-2.021. Juli 2026 · Metriken 1.13.0
PyPI
52MittelGesundheitsindex
attenlabs/hotato
Conversation QA for voice agents. Catch the calls that pass every text check but talk over the caller, skip a disclosure, or claim a task that never happened. Self-hosted, offline, MIT.
HTML★ 0↓ 2.575/Monat13. Juli 2026
MIT13. Juli 2026 · Metriken 1.13.0
PyPI
51MittelGesundheitsindex
lazily-hub/lazily-py
A Python library for lazy evaluation with context caching.
Python★ 117. Juli 2026
Apache-2.017. Juli 2026 · Metriken 1.13.0
PyPI
50MittelGesundheitsindex
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2.908/Monat22. Juli 2026
MIT22. Juli 2026 · Metriken 1.13.0
PyPI
47GefährdetGesundheitsindex
davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge
HTML · Python★ 66↓ 0/Monat14. Juli 2026
Keine Lizenz14. Juli 2026 · Metriken 1.13.0
npm
40GefährdetGesundheitsindex
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 316M/Monat22. Juli 2026
MIT22. Juli 2026 · Metriken 1.13.0
PyPI
39GefährdetGesundheitsindex
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 217. Juli 2026
Keine Lizenz17. Juli 2026 · Metriken 1.13.0
PyPI
26KritischGesundheitsindex
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 60721. Juli 2026
Eigene Lizenz21. Juli 2026 · Metriken 1.13.0