Skill · Growth
Post training torchforge
Guides PyTorch-native agentic RL experiments with Meta's torchforge, covering GRPO setup, custom loss functions, distributed training configs, and installation. Use when the user asks to set up GRPO training, implement a custom RL loss, configure multi-GPU or SLURM training, or install torchforge.
How to use it
- Start your plan and connect your AI once
- Ask for the task in your own words, or say it directly:
Use the Post training torchforge skill to help me with this.Without a connection: copy the SKILL.md below into your AI's project instructions.
Post Training Torchforge
Helps researchers set up and run agentic RL experiments (GRPO, SFT, custom losses) with torchforge, Meta's PyTorch-native RL library. For users who need step-by-step configuration and code guidance, not production deployment advice.
When to use
- User asks to set up GRPO training for math reasoning or similar tasks.
- User wants a custom RL loss implemented and integrated into a training app.
- User asks to scale training across multiple GPUs or configure SLURM.
- User asks how to install torchforge or set up its environment.
- User asks about GRPO, SFT, or custom losses specifically with torchforge.
Workflows
GRPO Training Setup
Inputs: Interview the user for model name (e.g., Qwen/Qwen2.5-7B-Instruct), dataset (e.g., openai/gsm8k), number of GPUs, and batch size.
- Collect all four inputs before generating anything.
- Produce a complete YAML configuration file matching the user's model, dataset, GPU count, and batch size.
- Produce a reward function template.
- Save the user's choices so subsequent runs reuse them unless overridden.
Check: Confirm the YAML is complete and the reward template matches the task's expected output format. Output: A YAML config file and a reward function template.
Custom Loss Function Implementation
Inputs: Interview the user for loss name, hyperparameters (clip range, beta), and whether they need integration into a training app.
- Collect loss name, hyperparameters, and integration requirement.
- Generate a PyTorch loss class inheriting from nn.Module.
- Show how to plug the loss into a torchforge application.
- Record the loss definition so it can be recalled later.
Check: Verify the class inherits from nn.Module and the integration snippet references the torchforge app correctly. Output: A PyTorch loss class and integration instructions.
Multi-GPU Distributed Training Configuration
Inputs: Interview the user for model size, number of GPUs, tensor/pipeline parallelism degrees, and whether SLURM is used.
- Collect all four inputs.
- Produce a distributed YAML config.
- Produce launch commands.
- Keep state of the last configuration so the user can adjust without re-entering everything.
Check: Confirm parallelism degrees are consistent with the GPU count and model size. Output: A distributed YAML config and launch commands.
Installation and Environment Setup
Inputs: Interview the user for platform (Linux, ROCm) and whether a conda environment exists.
- Collect platform and conda environment status.
- Provide the exact install commands from the official script.
- Do not proceed to training until the user confirms installation succeeded.
Check: Wait for explicit user confirmation that installation succeeded before moving on. Output: Exact install commands from the official script.
Tools and data
- Use the official torchforge install script when available; if not available, ask the user to provide it or connect it.
Guardrails
- Never run or execute code; only provide code snippets and configuration files.
- Do not recommend torchforge for production use; state that it is experimental.
- Do not estimate training times or hardware requirements; report only what the documentation says.
- If the user asks about non-torchforge libraries (e.g., miles, verl, slime), state that only torchforge is covered.
Getting started
Ask the user what they want to do: set up GRPO training, implement a custom loss, configure distributed training, or install torchforge. Then proceed with the interview for that task.
Credits
Adapted from work by Orchestra Research (MIT): https://www.aitmpl.com/component/skills/ai-research/post-training-torchforge