All tags
Catalogue tag

#reinforcement-learning

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

39 records
Tagged “reinforcement-learning”Ranked by health index
Maven · PyPI
100Exceptionalhealth index
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.4KAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · Go · crates.io
100Exceptionalhealth index
wandb/wandb
The AI developer platform. Use Weights & Biases to train and fine-tune models, and manage models from experimentation to production.
Python · Go★ 11.2K↓ 26.6M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI
98Exceptionalhealth index
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
C++ · Python★ 6,256Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI · crates.io
98Exceptionalhealth index
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 31.3K↓ 144M/moAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI
97Exceptionalhealth index
google-deepmind/optax
Optax is a gradient processing and optimization library for JAX.
Python★ 2,301Jul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI
96Exceptionalhealth index
Farama-Foundation/Gymnasium
A standard API for single-agent reinforcement learning environments, with popular reference environments and related utilities (formerly Gym)
Python★ 12.4K↓ 5.7M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
PyPI
95Exceptionalhealth index
google-deepmind/open_spiel
OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.
C++ · Python★ 5,440Aug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
94Exceptionalhealth index
PrimeIntellect-ai/verifiers
Our library for RL environments + evals
Python★ 4,565↓ 378.1K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
PyPI · npm
94Exceptionalhealth index
unslothai/unsloth
Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.
Python · TypeScript★ 69.6KAug 4, 2026
Apache-2.0Aug 4, 2026 · metrics 2.10.0
PyPI
93Exceptionalhealth index
AgileRL/AgileRL
Streamlining reinforcement learning with RLOps. State-of-the-art RL algorithms and tools, with 10x faster training through evolutionary hyperparameter optimization.
Python★ 938↓ 1,021/moJul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 2.10.0
PyPI · npm
93Exceptionalhealth index
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1,055↓ 406.4K/moJul 18, 2026
Apache-2.0Jul 18, 2026 · metrics 2.10.0
PyPI
93Exceptionalhealth index
mujocolab/mjlab
Isaac Lab API, powered by MuJoCo-Warp, for RL and robotics research
Python★ 2,709Jul 22, 2026
Apache-2.0Jul 22, 2026 · metrics 2.10.0
npm
93Exceptionalhealth index
proffesor-for-testing/agentic-qe
Agentic QE Fleet is an open-source AI-powered QA/QE platform designed for use with Coding Agents (works best with Claude Code) featuring specialized agents and skills to support testing activities for a product at any stage of the SDLC. Free to use, fork, build, and contribute. Based on the Agentic QE Framework created by Dragan Spiridonov.
TypeScript★ 474↓ 56.4K/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI · npm
93Exceptionalhealth index
trycua/cua
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
HTML · Rust · Python★ 20.9K↓ 26.8K/moAug 5, 2026
MITAug 5, 2026 · metrics 2.10.0
PyPI
92Excellenthealth index
JudgmentLabs/judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Python★ 1,057↓ 233.1K/moAug 13, 2026
Apache-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
92Excellenthealth index
munich-quantum-toolkit/predictor
MQT Predictor - A Tool for Automatic Device Selection with Device-Specific Circuit Compilation for Quantum Computing
Python★ 86↓ 285/moJul 21, 2026
MITJul 21, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
DLR-RM/stable-baselines3
PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.
Python★ 13.7KAug 12, 2026
MITAug 12, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
Farama-Foundation/Arcade-Learning-Environment
A simple framework that allows researchers and hobbyists to develop AI agents for Atari 2600 games
C++ · IDL★ 2,441Aug 13, 2026
GPL-2.0Aug 13, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
pytorch/rl
A modular, primitive-first, python-first PyTorch library for Reinforcement Learning.
Python★ 3,548↓ 84.8K/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI
90Excellenthealth index
Farama-Foundation/PettingZoo
A standard API for multi-agent reinforcement learning environments, with popular reference environments and related utilities
Python★ 3,489Aug 13, 2026
MITAug 13, 2026 · metrics 2.10.0
npm
90Excellenthealth index
enactic/openarm
A fully open-source humanoid arm for physical AI research and deployment in contact-rich environments.
MDX · TypeScript★ 2,924Sep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
PyPI
86Excellenthealth index
google-deepmind/dm_control
Google DeepMind's software stack for physics-based simulation and Reinforcement Learning environments, using MuJoCo.
Python★ 4,663Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI
84Excellenthealth index
ttktjmt/mjswan
MuJoco Simulation on Web Assembly with Neural netwroks
Python · TypeScript★ 309Jul 15, 2026
Apache-2.0Jul 15, 2026 · metrics 2.10.0
PyPI
81Excellenthealth index
desy-ml/cheetah
Fast and differentiable particle accelerator optics simulation for reinforcement learning and optimisation applications.
Python · Jupyter Notebook★ 69↓ 2,065/moAug 16, 2026
GPL-3.0Aug 16, 2026 · metrics 2.10.0
PyPI
80Excellenthealth index
ugr-sail/sinergym
Gym environment for building simulation and control using reinforcement learning
Python★ 234Aug 15, 2026
MITAug 15, 2026 · metrics 2.10.0
PyPI
75Goodhealth index
freesolo-co/flash
LoRA post-training for open-weight models: SFT, GRPO, and on-policy distillation. Describe a run in TOML; Flash allocates a GPU, trains, streams checkpoints, and serves the adapter.
Python★ 2↓ 5,834/moAug 29, 2026
Apache-2.0Aug 29, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
miskibin/py-draughts
Fastest Python draughts/checkers library — bitboards, 8 variants, alpha-beta engine, web UI. ~200x faster than pydraughts.
Python★ 16↓ 673/moAug 17, 2026
GPL-3.0Aug 17, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
araffin/sbx
SBX: Stable Baselines Jax (SB3 + Jax) RL algorithms
Python★ 608↓ 1,960/moSep 6, 2026
MITSep 6, 2026 · metrics 2.10.0
npm · PyPI
67Goodhealth index
alex-jb/orallexa-ai-trading-agent
Self-tuning multi-agent AI trading system. 8-source signal fusion, Bull/Bear/Judge debate on Claude Opus 4.7, Kelly + ATR position sizing. Python · Kalshi + Polymarket adapters.
Python · TypeScript★ 60Jul 26, 2026
MITJul 26, 2026 · metrics 2.10.0
65Goodhealth index
JuliaPOMDP/POMDPs.jl
MDPs and POMDPs in Julia - An interface for defining, solving, and simulating fully and partially observable Markov decision processes on discrete and continuous spaces.
Julia★ 764Aug 8, 2026
Custom licenseAug 8, 2026 · metrics 2.10.0