Reinforcement Learning skills
Free agent skills tagged reinforcement learning, ready to install into any SKILL.md-compatible agent.
16 skills
PufferLib
k-dense-ai
Guidance for reinforcement learning environments and policies.
Stable Baselines3
k-dense-ai
Implement and train reinforcement learning agents with ease.
Fine-Tuning with TRL
davila7
Align language models with human preferences using reinforcement learning.
Launch NeMo-RL
nvidia
Efficiently manage NeMo-RL recipes on Kubernetes.
Brev Etiquette
nvidia
Optimize storage and workflow for NeMo-RL agents.
TRL Fine-Tuning
nousresearch
Align language models with human preferences using TRL.
GRPO & RLVR Training
wshobson
Enhance model reasoning with verifiable rewards.
Torchforge RL Training
davila7
Streamline your PyTorch-native reinforcement learning workflows.
Miles RL Training
davila7
Enterprise-grade reinforcement learning for large models.
AgentDB Learning Plugins
ruvnet
Create and train AI learning plugins with ease.
ReasoningBank with AgentDB
ruvnet
Accelerate adaptive learning for self-learning agents.
GRPO/RL Training
davila7
Fine-tune models with structured output and custom rewards.
Auto Research
Autonomous NeMo-RL research agent workflow for directed hypothesis testing and open-ended discovery. Guides agents through the full experiment lifecycle: understanding recipes and environments, wiring RL or NeMo-gym runs, launching reproducible baselines and iterations, analyzing results, preserving human oversight, and using git plus TSV logs as the research ledger. Do NOT use for: bug fixes, code review, documentation, refactoring, dependency updates, or single-file changes.
TRL Training
Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning). Supports SFT, DPO, GRPO, KTO, RLOO and Reward Model training via CLI commands.
Volcano Engine RL Training
Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.
Slime RL Training
Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.
