PyPI99卓越健康指数huggingface/trlTrain transformer language models with reinforcement learning.Python★ 19.1K↓ 4.1M/月2026年8月24日Apache-2.02026年8月24日 · 指标 2.10.0
PyPI75良好健康指数freesolo-co/flashLoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.Python★ 2↓ 5,834/月2026年8月29日Apache-2.02026年8月29日 · 指标 2.10.0