全部标签
目录标签

#evaluation

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

58 条记录
标签为“evaluation”按健康指数排序
PyPI · npm
99卓越健康指数
langchain-ai/langsmith-sdk
LangSmith Client SDK Implementations
Python · TypeScript★ 1,039↓ 148M/月2026年8月27日
MIT2026年8月27日 · 指标 2.10.0
PyPI · npm
98卓越健康指数
Agenta-AI/agenta
Agenta is a workspace where you and your team build agents and automations.
TypeScript · Python★ 4,573↓ 17.6K/月2026年8月28日
自定义许可证2026年8月28日 · 指标 2.10.0
npm · Go
98卓越健康指数
langwatch/langwatch
The platform for LLM evaluations and AI agent testing
TypeScript★ 3,487↓ 1,393/月2026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
npm · PyPI
96卓越健康指数
comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python · TypeScript★ 21.1K↓ 111.9K/月2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
npm
96卓越健康指数
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.9K↓ 2.1M/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
Go · PyPI · npm
95卓越健康指数
Tencent/WeKnora
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 19.4K2026年8月5日
自定义许可证2026年8月5日 · 指标 2.10.0
PyPI
95卓越健康指数
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 3,3622026年7月19日
Apache-2.02026年7月19日 · 指标 2.10.0
Go
95卓越健康指数
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 1,5612026年7月19日
Apache-2.02026年7月19日 · 指标 2.10.0
PyPI
94卓越健康指数
langchain-ai/deepagents
The batteries-included agent harness.
Python★ 27.3K↓ 210.2K/月2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
npm
94卓越健康指数
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 32.5K2026年8月5日
自定义许可证2026年8月5日 · 指标 2.10.0
PyPI
94卓越健康指数
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Python · TypeScript★ 3,172↓ 67.9K/月2026年8月1日
Apache-2.02026年8月1日 · 指标 2.10.0
PyPI · npm
93卓越健康指数
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1,055↓ 406.4K/月2026年7月18日
Apache-2.02026年7月18日 · 指标 2.10.0
npm · PyPI
91优秀健康指数
joshuaswarren/remnic
Open-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.
TypeScript★ 176↓ 203.7K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm
90优秀健康指数
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2,069↓ 55.8K/月2026年7月18日
自定义许可证2026年7月18日 · 指标 2.10.0
npm · PyPI
90优秀健康指数
Marker-Inc-Korea/AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript · Python★ 4,963↓ 309/月2026年8月2日
自定义许可证2026年8月2日 · 指标 2.10.0
PyPI · npm
89优秀健康指数
UiPath/coder_eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Python · TypeScript★ 116↓ 13.9K/月2026年8月19日
Apache-2.02026年8月19日 · 指标 2.10.0
PyPI · Go · npm
89优秀健康指数
gooddata/gooddata-python-sdk
GoodData Cloud Python SDK
Python★ 35↓ 153.9K/月2026年8月22日
自定义许可证2026年8月22日 · 指标 2.10.0
PyPI
89优秀健康指数
vibrantlabsai/ragas
Supercharge Your LLM Application Evaluations 🚀
Python · Jupyter Notebook★ 15.2K↓ 1.6M/月2026年8月8日
Apache-2.02026年8月8日 · 指标 2.10.0
PyPI · npm
88优秀健康指数
hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Python★ 124↓ 2,207/月2026年7月26日
Apache-2.02026年7月26日 · 指标 2.10.0
PyPI
86优秀健康指数
BrainLesion/panoptica
panoptica -- instance-wise evaluation of 3D semantic and instance segmentation maps
Python · Jupyter Notebook★ 332026年7月31日
Apache-2.02026年7月31日 · 指标 2.10.0
PyPI
86优秀健康指数
MichaelGrupp/evo
Python package for the evaluation of odometry and SLAM
Python★ 4,309↓ 210.7K/月2026年8月28日
GPL-3.02026年8月28日 · 指标 2.10.0
NuGet
86优秀健康指数
asc-community/AngouriMath
Open-source cross-platform symbolic algebra library for C# and F#. Can be used for both production and research purposes.
C#★ 8272026年8月23日
MIT2026年8月23日 · 指标 2.10.0
PyPI · npm
86优秀健康指数
robocurve/inspect-robots
Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.
Python★ 269↓ 5,131/月2026年9月6日
MIT2026年9月6日 · 指标 2.10.0
npm · crates.io · PyPI
84优秀健康指数
PSU3D0/formualizer
Embeddable spreadsheet engine - parse, evaluate & mutate Excel workbooks from Rust, Python, or the browser. Arrow-powered, 400+ functions.
Rust★ 177↓ 10.9K/月2026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
npm · Go · crates.io +2
81优秀健康指数
ops-ai/Toggly.FeatureManagement
Enables teams to release software faster and safer, and with better results.
TypeScript · C#★ 5↓ 3,643/月2026年9月3日
MIT2026年9月3日 · 指标 2.10.0
npm · PyPI
80优秀健康指数
aikdna/kdna
KDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/月2026年7月22日
Apache-2.02026年7月22日 · 指标 2.10.0
npm
80优秀健康指数
o-stepper/graphorin
Project Graphorin is a TypeScript framework for personal AI assistants and long-living agents with rich memory, durable workflow, and observability out of the box.
TypeScript★ 3↓ 43.5K/月2026年8月1日
MIT2026年8月1日 · 指标 2.10.0
PyPI
78良好健康指数
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 2,4652026年7月18日
Apache-2.02026年7月18日 · 指标 2.10.0
PyPI
73良好健康指数
mjpost/sacrebleu
Reference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1,259↓ 4.3M/月2026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
npm
71良好健康指数
JudgmentLabs/judgeval-js
The open source post-building layer for agents.
TypeScript★ 5↓ 92.5K/月2026年8月1日
无许可证2026年8月1日 · 指标 2.10.0