全部标签
目录标签

#evaluation

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

22 条记录
标签为“evaluation”按健康指数排序
npm · PyPI
88优秀健康指数
langchain-ai/langsmith-sdk
LangSmith Client SDK Implementations
Python · TypeScript★ 971↓ 22.6M/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
PyPI · npm
83良好健康指数
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
Python · TypeScript★ 27.1K↓ 40.1M/月2026年7月18日
Apache-2.02026年7月18日 · 指标 1.13.0
PyPI
82良好健康指数
embeddings-benchmark/mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Python · Jupyter Notebook★ 3,3622026年7月19日
Apache-2.02026年7月19日 · 指标 1.13.0
npm
81良好健康指数
langfuse/langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
TypeScript★ 31.3K2026年7月17日
自定义许可证2026年7月17日 · 指标 1.13.0
Go
81良好健康指数
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Go★ 1,5612026年7月19日
Apache-2.02026年7月19日 · 指标 1.13.0
npm
80良好健康指数
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
TypeScript★ 23.4K↓ 1.7M/月2026年7月17日
MIT2026年7月17日 · 指标 1.13.0
PyPI · npm
79良好健康指数
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1,055↓ 406.4K/月2026年7月18日
Apache-2.02026年7月18日 · 指标 1.13.0
Go · npm · PyPI
77良好健康指数
Tencent/WeKnora
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go · Vue · TypeScript★ 18.7K2026年7月21日
自定义许可证2026年7月21日 · 指标 1.13.0
npm
75良好健康指数
MCPJam/inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
TypeScript★ 2,069↓ 55.8K/月2026年7月18日
自定义许可证2026年7月18日 · 指标 1.13.0
npm · PyPI
68中等健康指数
aikdna/kdna
KDNA protocol and Core runtime for versioned, verifiable, encrypted, authorized judgment assets.
JavaScript★ 28↓ 20.7K/月2026年7月22日
Apache-2.02026年7月22日 · 指标 1.13.0
PyPI
67中等健康指数
huggingface/evaluate
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
Python★ 2,4652026年7月18日
Apache-2.02026年7月18日 · 指标 1.13.0
PyPI
66中等健康指数
robocurve/inspect-robots
Evaluation framework for VLA / physical-AI models: define a benchmark once, run any policy on any robot or sim. (The Inspect AI for robotics.)
Python★ 52026年7月15日
MIT2026年7月15日 · 指标 1.13.0
PyPI
60中等健康指数
mjpost/sacrebleu
Reference BLEU implementation that auto-downloads test sets and reports a version string to facilitate cross-lab comparisons
Python★ 1,253↓ 3.9M/月2026年7月18日
Apache-2.02026年7月18日 · 指标 1.13.0
Go · npm
58中等健康指数
starksv/windows-iso-downloader
Download official Windows ISOs directly from Microsoft's CDN. Web app + CLI tool. No account, no ads, no browser required.
TypeScript★ 7↓ 0/月2026年7月14日
MIT2026年7月14日 · 指标 1.13.0
PyPI
55中等健康指数
EpsilabAI/epsilab-python
The official Python library for the Epsilab API
Python★ 0↓ 2,706/月2026年7月21日
Apache-2.02026年7月21日 · 指标 1.13.0
PyPI
52中等健康指数
attenlabs/hotato
Conversation QA for voice agents. Catch the calls that pass every text check but talk over the caller, skip a disclosure, or claim a task that never happened. Self-hosted, offline, MIT.
HTML★ 0↓ 2,575/月2026年7月13日
MIT2026年7月13日 · 指标 1.13.0
PyPI
51中等健康指数
lazily-hub/lazily-py
A Python library for lazy evaluation with context caching.
Python★ 12026年7月17日
Apache-2.02026年7月17日 · 指标 1.13.0
PyPI
50中等健康指数
JarJarBeatyourattitude/evalt
Budget-bounded LLM routing that finds the cheapest model and prompt meeting your accuracy target.
Python★ 0↓ 2,908/月2026年7月22日
MIT2026年7月22日 · 指标 1.13.0
PyPI
47存在风险健康指数
davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge
HTML · Python★ 66↓ 0/月2026年7月14日
无许可证2026年7月14日 · 指标 1.13.0
npm
40存在风险健康指数
sindresorhus/define-lazy-prop
Define a lazily evaluated property on an object
JavaScript · TypeScript★ 67↓ 316M/月2026年7月22日
MIT2026年7月22日 · 指标 1.13.0
PyPI
39存在风险健康指数
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 22026年7月17日
无许可证2026年7月17日 · 指标 1.13.0
PyPI
26危急健康指数
danthedeckie/simpleeval
Simple Safe Sandboxed Extensible Expression Evaluator for Python
Python★ 6072026年7月21日
自定义许可证2026年7月21日 · 指标 1.13.0