Todas las etiquetas
Etiqueta del catálogo

#reinforcement-learning

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

18 registros
Con la etiqueta «reinforcement-learning»Ordenado por índice de salud
PyPI · Go · crates.io
94Excelenteíndice de salud
wandb/wandb
The AI developer platform. Use Weights & Biases to train and fine-tune models, and manage models from experimentation to production.
Python · Go★ 11.2K↓ 23M/mes20 jul 2026
MIT20 jul 2026 · métricas 1.13.0
Maven · PyPI
90Excelenteíndice de salud
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python · C++★ 43.3K17 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI
85Excelenteíndice de salud
Farama-Foundation/Gymnasium
A standard API for single-agent reinforcement learning environments, with popular reference environments and related utilities (formerly Gym)
Python★ 12.2K↓ 10.1M/mes18 jul 2026
MIT18 jul 2026 · métricas 1.13.0
PyPI
85Excelenteíndice de salud
google-deepmind/optax
Optax is a gradient processing and optimization library for JAX.
Python★ 230118 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
crates.io · PyPI
82Buenoíndice de salud
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python★ 30.3K↓ 274M/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
PyPI · npm
82Buenoíndice de salud
unslothai/unsloth
Unsloth is a local UI for training and running Gemma 4, Qwen3.6, DeepSeek, Kimi, GLM and other models.
Python · TypeScript★ 68.5K↓ 2.3M/mes20 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0
PyPI · npm
79Buenoíndice de salud
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
Python · MDX★ 1055↓ 406.4K/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
79Buenoíndice de salud
pytorch/rl
A modular, primitive-first, python-first PyTorch library for Reinforcement Learning.
Python★ 349115 jul 2026
MIT15 jul 2026 · métricas 1.13.0
PyPI
78Buenoíndice de salud
AgileRL/AgileRL
Streamlining reinforcement learning with RLOps. State-of-the-art RL algorithms and tools, with 10x faster training through evolutionary hyperparameter optimization.
Python★ 938↓ 1021/mes20 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0
PyPI
78Buenoíndice de salud
mujocolab/mjlab
Isaac Lab API, powered by MuJoCo-Warp, for RL and robotics research
Python★ 270922 jul 2026
Apache-2.022 jul 2026 · métricas 1.13.0
PyPI
77Buenoíndice de salud
munich-quantum-toolkit/predictor
MQT Predictor - A Tool for Automatic Device Selection with Device-Specific Circuit Compilation for Quantum Computing
Python★ 86↓ 285/mes21 jul 2026
MIT21 jul 2026 · métricas 1.13.0
npm · PyPI
76Buenoíndice de salud
trycua/cua
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
HTML · Python · Rust★ 20.5K↓ 153/mes22 jul 2026
MIT22 jul 2026 · métricas 1.13.0
npm
75Buenoíndice de salud
enactic/openarm
A fully open-source humanoid arm for physical AI research and deployment in contact-rich environments.
MDX · TypeScript★ 2718↓ 0/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
PyPI
71Buenoíndice de salud
ttktjmt/mjswan
MuJoco Simulation on Web Assembly with Neural netwroks
Python · TypeScript★ 30915 jul 2026
Apache-2.015 jul 2026 · métricas 1.13.0
PyPI · crates.io
55Moderadoíndice de salud
tsilva/SuperMarioBros-Nes-turbo
🚀 Blazing fast SuperMarioBros-Nes environment for RL research 🍄
Python · Rust★ 0↓ 2748/mes16 jul 2026
MIT16 jul 2026 · métricas 1.13.0
npm
49En riesgoíndice de salud
ruvnet/agentdb
Vector memory that gets smarter every time your agent uses it.
TypeScript · HTML · JavaScript★ 80↓ 649.1K/mes23 jul 2026
MIT23 jul 2026 · métricas 1.13.0
PyPI
39En riesgoíndice de salud
waybarrios/crystal
CRYSTAL: Beyond Final Answers: Benchmark for Transparent Multimodal Reasoning Evaluation | arXiv 2603.13099
Python★ 217 jul 2026
Sin licencia17 jul 2026 · métricas 1.13.0