Todas las etiquetas
Etiqueta del catálogo

#evaluation

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

22 registros
Con la etiqueta «evaluation»Ordenado por índice de salud
npm · PyPI
88Excelenteíndice de salud
langchain-ai/langsmith-sdk
LangSmith Client SDK Implementations
Python · TypeScript★ 971↓ 22.6M/mes17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
PyPI · npm
83Buenoíndice de salud
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.1K↓ 40.1M/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
82Buenoíndice de salud
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 336219 jul 2026
Apache-2.019 jul 2026 · métricas 1.13.0
npm
81Buenoíndice de salud
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 31.3K17 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
Go
81Buenoíndice de salud
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 156119 jul 2026
Apache-2.019 jul 2026 · métricas 1.13.0
npm
80Buenoíndice de salud
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.4K↓ 1.7M/mes17 jul 2026
MIT17 jul 2026 · métricas 1.13.0
PyPI · npm
79Buenoíndice de salud
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1055↓ 406.4K/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
Go · npm · PyPI
77Buenoíndice de salud
Tencent/WeKnora
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 18.7K21 jul 2026
Licencia propia21 jul 2026 · métricas 1.13.0
npm
75Buenoíndice de salud
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2069↓ 55.8K/mes18 jul 2026
Licencia propia18 jul 2026 · métricas 1.13.0
npm · PyPI
68Moderadoíndice de salud
aikdna/kdna
KDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/mes22 jul 2026
Apache-2.022 jul 2026 · métricas 1.13.0
PyPI
67Moderadoíndice de salud
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 246518 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
66Moderadoíndice de salud
robocurve/inspect-robots
Evaluation framework for VLA / physical-AI models: define a benchmark once, run any policy on any robot or sim. (The Inspect AI for robotics.)
Python★ 515 jul 2026
MIT15 jul 2026 · métricas 1.13.0
PyPI
60Moderadoíndice de salud
mjpost/sacrebleu
Reference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1253↓ 3.9M/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
Go · npm
58Moderadoíndice de salud
starksv/windows-iso-downloader
Download official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript★ 7↓ 0/mes14 jul 2026
MIT14 jul 2026 · métricas 1.13.0
PyPI
55Moderadoíndice de salud
EpsilabAI/epsilab-python
The official Python library for the Epsilab API
Python★ 0↓ 2706/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI
52Moderadoíndice de salud
attenlabs/hotato
Conversation QA for voice agents. Catch the calls that pass every text check but talk over the caller, skip a disclosure, or claim a task that never happened. Self-hosted, offline, MIT.
HTML★ 0↓ 2575/mes13 jul 2026
MIT13 jul 2026 · métricas 1.13.0
PyPI
51Moderadoíndice de salud
lazily-hub/lazily-py
A Python library for lazy evaluation with context caching.
Python★ 117 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI
50Moderadoíndice de salud
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2908/mes22 jul 2026
MIT22 jul 2026 · métricas 1.13.0
PyPI
47En riesgoíndice de salud
davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge
HTML · Python★ 66↓ 0/mes14 jul 2026
Sin licencia14 jul 2026 · métricas 1.13.0
npm
40En riesgoíndice de salud
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 316M/mes22 jul 2026
MIT22 jul 2026 · métricas 1.13.0
PyPI
39En riesgoíndice de salud
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 217 jul 2026
Sin licencia17 jul 2026 · métricas 1.13.0
PyPI
26Críticoíndice de salud
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 60721 jul 2026
Licencia propia21 jul 2026 · métricas 1.13.0