PyPI · npm100Exceptionalhealth index
Python · TypeScript★ 11.2KAug 28, 2026
PyPI97Exceptionalhealth index
Python★ 5,775Aug 28, 2026
npm · PyPI96Exceptionalhealth index

comet-ml/opikDebug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/moAug 5, 2026
npm96Exceptionalhealth index

promptfoo/promptfooTest your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/moAug 5, 2026
PyPI94Exceptionalhealth index

Q00/ouroborosAgent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
Python★ 5,403Aug 13, 2026
PyPI · npm94Exceptionalhealth index
Python · TypeScript★ 17.4K↓ 6.3M/moAug 5, 2026
npm94Exceptionalhealth index

langfuse/langfuse🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5KAug 5, 2026
PyPI92Excellenthealth index

JudgmentLabs/judgevalThe Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Python★ 1,057↓ 233.1K/moAug 13, 2026
Go91Excellenthealth index

praetorian-inc/juliusSimple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 175Aug 8, 2026
npm · PyPI90Excellenthealth index
Marker-Inc-Korea/AutoRAGAutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4,963↓ 309/moAug 2, 2026
PyPI90Excellenthealth index

truera/trulensEvaluation and Tracking for LLM Experiments and AI Agents
Python★ 3,487Aug 6, 2026
PyPI · npm89Excellenthealth index

UiPath/coder_evalTest that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/moAug 19, 2026
npm89Excellenthealth index

microsoft/promptyPrompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandability, and portability for developers.
Rust · C# · TypeScript★ 1,248Aug 13, 2026

alizahidraja/isnadGrade every agent, scraper and model in a claim's chain — provenance, trust scoring and audit evidence for LLM pipelines
Python★ 37↓ 4,204/moAug 29, 2026
Go · Maven · npm77Goodhealth index
hugalafutro/model-hotelMulti-Provider AI Gateway - No personal logs by design. Model autodiscovery, Failover groups, High availabilty, Android companion app, and more. "Because we have LiteLLM at home"
Go · TypeScript★ 50Jul 18, 2026
b7n0de/proofbundleOffline cryptographic receipts for AI evaluation results — Ed25519 + RFC 6962 Merkle + optional SD-JWT. Integrity, not truth
Python★ 2↓ 6,574/moJul 23, 2026
Hex · npm73Goodhealth index
Elixir · JavaScript · HTML★ 34↓ 19.3K/moJul 17, 2026

phierceweb/pf-corePython foundation for LLM apps whose prompts and spend you can actually see — versioned prompts, every call recorded and replayable, budgets, evals, jobs.
Python★ 2↓ 1,160/moSep 5, 2026
RubyGems67Goodhealth index
homemade-software-inc/completion-kitYour prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 1Jul 17, 2026
PyPI · npm63Moderatehealth index
Python · TypeScript★ 0↓ 2,525/moJul 17, 2026
tamnd/taocp-solverA Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 0Jul 28, 2026
npm53Moderatehealth index
mykim-aus/hey-llm-you-okayHey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2,332/moJul 30, 2026
PyPI50Moderatehealth index
Python★ 0↓ 2,908/moJul 22, 2026
PyPI34At Riskhealth index
waybarrios/crystalCRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 2Jul 17, 2026

mlflow/mlflowThe open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/moAug 5, 2026