New to Claude Skills? Learn how to install them →

Reinforcement Learning skills

Free agent skills tagged reinforcement learning, ready to install into any SKILL.md-compatible agent.

P

PufferLib

k-dense-ai

Guidance for reinforcement learning environments and policies.

Developer ToolsintermediatePython · Shell33.2k repo

Stable Baselines3

k-dense-ai

Implement and train reinforcement learning agents with ease.

Developer ToolsintermediatePython · Shell33.2k repo

Fine-Tuning with TRL

davila7

Align language models with human preferences using reinforcement learning.

Developer ToolsintermediatePython · Shell30.2k repo

Launch NeMo-RL

nvidia

Efficiently manage NeMo-RL recipes on Kubernetes.

Developer ToolsintermediateShell2.8k repo
B

Brev Etiquette

nvidia

Optimize storage and workflow for NeMo-RL agents.

Developer ToolsintermediateShell2.8k repo

TRL Fine-Tuning

nousresearch

Align language models with human preferences using TRL.

Developer ToolsintermediatePython · Shell228.5k repo

GRPO & RLVR Training

wshobson

Enhance model reasoning with verifiable rewards.

Developer Toolsintermediate38.7k repo

Torchforge RL Training

davila7

Streamline your PyTorch-native reinforcement learning workflows.

Developer ToolsintermediateShell30.2k repo

Miles RL Training

davila7

Enterprise-grade reinforcement learning for large models.

Developer ToolsadvancedPython · Shell30.2k repo

AgentDB Learning Plugins

ruvnet

Create and train AI learning plugins with ease.

AI & AgentsintermediateNode · Shell67.6k repo

ReasoningBank with AgentDB

ruvnet

Accelerate adaptive learning for self-learning agents.

AI & AgentsintermediateNode · Shell67.6k repo
G

GRPO/RL Training

davila7

Fine-tune models with structured output and custom rewards.

Developer ToolsadvancedPython30.2k repo

Auto Research

Autonomous NeMo-RL research agent workflow for directed hypothesis testing and open-ended discovery. Guides agents through the full experiment lifecycle: understanding recipes and environments, wiring RL or NeMo-gym runs, launching reproducible baselines and iterations, analyzing results, preserving human oversight, and using git plus TSV logs as the research ledger. Do NOT use for: bug fixes, code review, documentation, refactoring, dependency updates, or single-file changes.

TRL Training

Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning). Supports SFT, DPO, GRPO, KTO, RLOO and Reward Model training via CLI commands.

Volcano Engine RL Training

Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.

Slime RL Training

Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.