Todas las etiquetas
Etiqueta del catálogo

#llm-evaluation

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

25 registros
Con la etiqueta «llm-evaluation»Ordenado por índice de salud
PyPI · npm
100Excepcionalíndice de salud
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2K28 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python★ 577528 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
npm · PyPI
96Excepcionalíndice de salud
comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
npm
96Excepcionalíndice de salud
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/mes5 ago 2026
MIT5 ago 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
Q00/ouroboros
Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
Python★ 540313 ago 2026
MIT13 ago 2026 · métricas 2.10.0
PyPI · npm
94Excepcionalíndice de salud
confident-ai/deepeval
The LLM Evaluation Framework
Python · TypeScript★ 17.4K↓ 6.3M/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
npm
94Excepcionalíndice de salud
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5K5 ago 2026
Licencia propia5 ago 2026 · métricas 2.10.0
PyPI
92Excelenteíndice de salud
JudgmentLabs/judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Python★ 1057↓ 233.1K/mes13 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
Go
91Excelenteíndice de salud
praetorian-inc/julius
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 1758 ago 2026
Apache-2.08 ago 2026 · métricas 2.10.0
npm · PyPI
90Excelenteíndice de salud
Marker-Inc-Korea/AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4963↓ 309/mes2 ago 2026
Licencia propia2 ago 2026 · métricas 2.10.0
PyPI
90Excelenteíndice de salud
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 34876 ago 2026
MIT6 ago 2026 · métricas 2.10.0
PyPI · npm
89Excelenteíndice de salud
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/mes19 ago 2026
Apache-2.019 ago 2026 · métricas 2.10.0
npm
89Excelenteíndice de salud
microsoft/prompty
Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandability, and portability for developers.
Rust · C# · TypeScript★ 124813 ago 2026
MIT13 ago 2026 · métricas 2.10.0
PyPI
77Buenoíndice de salud
alizahidraja/isnad
Grade every agent, scraper and model in a claim's chain — provenance, trust scoring and audit evidence for LLM pipelines
Python★ 37↓ 4204/mes29 ago 2026
Apache-2.029 ago 2026 · métricas 2.10.0
Go · Maven · npm
77Buenoíndice de salud
hugalafutro/model-hotel
Multi-Provider AI Gateway - No personal logs by design. Model autodiscovery, Failover groups, High availabilty, Android companion app, and more. "Because we have LiteLLM at home"
Go · TypeScript★ 5018 jul 2026
MIT18 jul 2026 · métricas 2.10.0
PyPI
73Buenoíndice de salud
b7n0de/proofbundle
Offline cryptographic receipts for AI evaluation results — Ed25519 + RFC 6962 Merkle + optional SD-JWT. Integrity, not truth
Python★ 2↓ 6574/mes23 jul 2026
MIT23 jul 2026 · métricas 2.10.0
Hex · npm
73Buenoíndice de salud
ccarvalho-eng/aludel
LLM Evaluation for Phoenix Apps
Elixir · JavaScript · HTML★ 34↓ 19.3K/mes17 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI
73Buenoíndice de salud
phierceweb/pf-core
Python foundation for LLM apps whose prompts and spend you can actually see — versioned prompts, every call recorded and replayable, budgets, evals, jobs.
Python★ 2↓ 1160/mes5 sept 2026
MIT5 sept 2026 · métricas 2.10.0
RubyGems
67Buenoíndice de salud
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 117 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
PyPI · npm
63Moderadoíndice de salud
fastaifoundry/fastaiagent-sdk
El repositorio no publica descripción.
Python · TypeScript★ 0↓ 2525/mes17 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
Go
54Moderadoíndice de salud
tamnd/taocp-solver
A Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 028 jul 2026
MIT28 jul 2026 · métricas 2.10.0
npm
53Moderadoíndice de salud
mykim-aus/hey-llm-you-okay
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2332/mes30 jul 2026
MIT30 jul 2026 · métricas 2.10.0
PyPI
50Moderadoíndice de salud
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2908/mes22 jul 2026
MIT22 jul 2026 · métricas 2.10.0
PyPI
34En riesgoíndice de salud
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 217 jul 2026
Sin licencia17 jul 2026 · métricas 2.10.0
PyPI · npm
19Críticoíndice de salud
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/mes5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0