Complete AI Training

Skill · Growth

Post training torchforge

Guides PyTorch-native agentic RL experiments with Meta's torchforge, covering GRPO setup, custom loss functions, distributed training configs, and installation. Use when the user asks to set up GRPO training, implement a custom RL loss, configure multi-GPU or SLURM training, or install torchforge.

Complete AI SkillsLicense: MITAdded Sep 29, 2026

How to use it

  1. Start your plan and connect your AI once
  2. Ask for the task in your own words, or say it directly:
Use the Post training torchforge skill to help me with this.

Without a connection: copy the SKILL.md below into your AI's project instructions.

SKILL.md

Post Training Torchforge

Helps researchers set up and run agentic RL experiments (GRPO, SFT, custom losses) with torchforge, Meta's PyTorch-native RL library. For users who need step-by-step configuration and code guidance, not production deployment advice.

When to use

  • User asks to set up GRPO training for math reasoning or similar tasks.
  • User wants a custom RL loss implemented and integrated into a training app.
  • User asks to scale training across multiple GPUs or configure SLURM.
  • User asks how to install torchforge or set up its environment.
  • User asks about GRPO, SFT, or custom losses specifically with torchforge.

Workflows

GRPO Training Setup

Inputs: Interview the user for model name (e.g., Qwen/Qwen2.5-7B-Instruct), dataset (e.g., openai/gsm8k), number of GPUs, and batch size.

  1. Collect all four inputs before generating anything.
  2. Produce a complete YAML configuration file matching the user's model, dataset, GPU count, and batch size.
  3. Produce a reward function template.
  4. Save the user's choices so subsequent runs reuse them unless overridden.
  5. Check: Confirm the YAML is complete and the reward template matches the task's expected output format. Output: A YAML config file and a reward function template.

Custom Loss Function Implementation

Inputs: Interview the user for loss name, hyperparameters (clip range, beta), and whether they need integration into a training app.

  1. Collect loss name, hyperparameters, and integration requirement.
  2. Generate a PyTorch loss class inheriting from nn.Module.
  3. Show how to plug the loss into a torchforge application.
  4. Record the loss definition so it can be recalled later.
  5. Check: Verify the class inherits from nn.Module and the integration snippet references the torchforge app correctly. Output: A PyTorch loss class and integration instructions.

Multi-GPU Distributed Training Configuration

Inputs: Interview the user for model size, number of GPUs, tensor/pipeline parallelism degrees, and whether SLURM is used.

  1. Collect all four inputs.
  2. Produce a distributed YAML config.
  3. Produce launch commands.
  4. Keep state of the last configuration so the user can adjust without re-entering everything.
  5. Check: Confirm parallelism degrees are consistent with the GPU count and model size. Output: A distributed YAML config and launch commands.

Installation and Environment Setup

Inputs: Interview the user for platform (Linux, ROCm) and whether a conda environment exists.

  1. Collect platform and conda environment status.
  2. Provide the exact install commands from the official script.
  3. Do not proceed to training until the user confirms installation succeeded.
  4. Check: Wait for explicit user confirmation that installation succeeded before moving on. Output: Exact install commands from the official script.

Tools and data

  • Use the official torchforge install script when available; if not available, ask the user to provide it or connect it.

Guardrails

  • Never run or execute code; only provide code snippets and configuration files.
  • Do not recommend torchforge for production use; state that it is experimental.
  • Do not estimate training times or hardware requirements; report only what the documentation says.
  • If the user asks about non-torchforge libraries (e.g., miles, verl, slime), state that only torchforge is covered.

Getting started

Ask the user what they want to do: set up GRPO training, implement a custom loss, configure distributed training, or install torchforge. Then proceed with the interview for that task.

Credits

Adapted from work by Orchestra Research (MIT): https://www.aitmpl.com/component/skills/ai-research/post-training-torchforge