Усі теги
Тег каталогу

#llm-evaluation

Усі репозиторії публічного реєстру з цим тегом — із тем GitHub або ключових слів, які публікують їхні реєстри пакетів. Здоров'я вимірюється за тією ж версіонованою методологією, що й решта реєстру.

25 записів
З тегом «llm-evaluation»Упорядковано за індексом здоров'я
PyPI · npm
100Винятковийіндекс здоров'я
Arize-ai/phoenix
AI Observability & Evaluation
Python · TypeScript★ 11.2K28 серп. 2026 р.
Власна ліцензія28 серп. 2026 р. · метрики 2.10.0
PyPI
97Винятковийіндекс здоров'я
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python★ 5 77528 серп. 2026 р.
Apache-2.028 серп. 2026 р. · метрики 2.10.0
npm · PyPI
96Винятковийіндекс здоров'я
comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/міс5 серп. 2026 р.
Apache-2.05 серп. 2026 р. · метрики 2.10.0
npm
96Винятковийіндекс здоров'я
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/міс5 серп. 2026 р.
MIT5 серп. 2026 р. · метрики 2.10.0
PyPI
94Винятковийіндекс здоров'я
Q00/ouroboros
Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
Python★ 5 40313 серп. 2026 р.
MIT13 серп. 2026 р. · метрики 2.10.0
PyPI · npm
94Винятковийіндекс здоров'я
confident-ai/deepeval
The LLM Evaluation Framework
Python · TypeScript★ 17.4K↓ 6.3M/міс5 серп. 2026 р.
Apache-2.05 серп. 2026 р. · метрики 2.10.0
npm
94Винятковийіндекс здоров'я
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5K5 серп. 2026 р.
Власна ліцензія5 серп. 2026 р. · метрики 2.10.0
PyPI
92Відміннийіндекс здоров'я
JudgmentLabs/judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Python★ 1 057↓ 233.1K/міс13 серп. 2026 р.
Apache-2.013 серп. 2026 р. · метрики 2.10.0
Go
91Відміннийіндекс здоров'я
praetorian-inc/julius
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
Go★ 1758 серп. 2026 р.
Apache-2.08 серп. 2026 р. · метрики 2.10.0
npm · PyPI
90Відміннийіндекс здоров'я
Marker-Inc-Korea/AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4 963↓ 309/міс2 серп. 2026 р.
Власна ліцензія2 серп. 2026 р. · метрики 2.10.0
PyPI
90Відміннийіндекс здоров'я
truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
Python★ 3 4876 серп. 2026 р.
MIT6 серп. 2026 р. · метрики 2.10.0
PyPI · npm
89Відміннийіндекс здоров'я
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/міс19 серп. 2026 р.
Apache-2.019 серп. 2026 р. · метрики 2.10.0
npm
89Відміннийіндекс здоров'я
microsoft/prompty
Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandability, and portability for developers.
Rust · C# · TypeScript★ 1 24813 серп. 2026 р.
MIT13 серп. 2026 р. · метрики 2.10.0
PyPI
77Добрийіндекс здоров'я
alizahidraja/isnad
Grade every agent, scraper and model in a claim's chain — provenance, trust scoring and audit evidence for LLM pipelines
Python★ 37↓ 4 204/міс29 серп. 2026 р.
Apache-2.029 серп. 2026 р. · метрики 2.10.0
Go · Maven · npm
77Добрийіндекс здоров'я
hugalafutro/model-hotel
Multi-Provider AI Gateway - No personal logs by design. Model autodiscovery, Failover groups, High availabilty, Android companion app, and more. "Because we have LiteLLM at home"
Go · TypeScript★ 5018 лип. 2026 р.
MIT18 лип. 2026 р. · метрики 2.10.0
PyPI
73Добрийіндекс здоров'я
b7n0de/proofbundle
Offline cryptographic receipts for AI evaluation results — Ed25519 + RFC 6962 Merkle + optional SD-JWT. Integrity, not truth
Python★ 2↓ 6 574/міс23 лип. 2026 р.
MIT23 лип. 2026 р. · метрики 2.10.0
Hex · npm
73Добрийіндекс здоров'я
ccarvalho-eng/aludel
LLM Evaluation for Phoenix Apps
Elixir · JavaScript · HTML★ 34↓ 19.3K/міс17 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 2.10.0
PyPI
73Добрийіндекс здоров'я
phierceweb/pf-core
Python foundation for LLM apps whose prompts and spend you can actually see — versioned prompts, every call recorded and replayable, budgets, evals, jobs.
Python★ 2↓ 1 160/міс5 вер. 2026 р.
MIT5 вер. 2026 р. · метрики 2.10.0
RubyGems
67Добрийіндекс здоров'я
homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby · HTML★ 117 лип. 2026 р.
Власна ліцензія17 лип. 2026 р. · метрики 2.10.0
PyPI · npm
63Помірнийіндекс здоров'я
fastaifoundry/fastaiagent-sdk
Опис репозиторію не опубліковано.
Python · TypeScript★ 0↓ 2 525/міс17 лип. 2026 р.
Apache-2.017 лип. 2026 р. · метрики 2.10.0
Go
54Помірнийіндекс здоров'я
tamnd/taocp-solver
A Go library and CLI for complete TAOCP solutions, fast and audited solving modes, reproducible model evaluation, and detailed token and list-cost accounting.
Go★ 028 лип. 2026 р.
MIT28 лип. 2026 р. · метрики 2.10.0
npm
53Помірнийіндекс здоров'я
mykim-aus/hey-llm-you-okay
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
TypeScript · JavaScript★ 1↓ 2 332/міс30 лип. 2026 р.
MIT30 лип. 2026 р. · метрики 2.10.0
PyPI
50Помірнийіндекс здоров'я
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2 908/міс22 лип. 2026 р.
MIT22 лип. 2026 р. · метрики 2.10.0
PyPI
34У зоні ризикуіндекс здоров'я
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 217 лип. 2026 р.
Без ліцензії17 лип. 2026 р. · метрики 2.10.0
PyPI · npm
19Критичнийіндекс здоров'я
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.4K↓ 41.5M/міс5 серп. 2026 р.
Apache-2.05 серп. 2026 р. · метрики 2.10.0