PyPI99卓越健康指数huggingface/trlTrain transformer language models with reinforcement learning.Python★ 19.1K↓ 4.1M/月2026年8月24日Apache-2.02026年8月24日 · 指标 2.10.0
PyPI94卓越健康指数PrimeIntellect-ai/verifiersOur library for RL environments + evalsPython★ 4,565↓ 378.1K/月2026年8月28日MIT2026年8月28日 · 指标 2.10.0
PyPI94卓越健康指数modelscope/ms-swiftUse PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).Python★ 15.4K↓ 111K/月2026年8月27日Apache-2.02026年8月27日 · 指标 2.10.0
PyPI92优秀健康指数JudgmentLabs/judgevalThe Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.Python★ 1,057↓ 233.1K/月2026年8月13日Apache-2.02026年8月13日 · 指标 2.10.0
PyPI75良好健康指数freesolo-co/flashLoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.Python★ 2↓ 5,834/月2026年8月29日Apache-2.02026年8月29日 · 指标 2.10.0
PyPI34存在风险健康指数waybarrios/crystalCRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099Python★ 22026年7月17日无许可证2026年7月17日 · 指标 2.10.0