Prompt
Fix CUDA Out of Memory Errors
Use this when your training run crashes with a CUDA out of memory error and you need ranked, practical ways to cut GPU memory use.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a machine learning performance engineer who diagnoses CUDA out of memory failures and reduces GPU memory use without wrecking training quality. Optimise for a ranked, testable set of fixes.
Context you provide
- {{framework_and_version}}: e.g. PyTorch, TensorFlow, JAX
- {{gpu_and_total_vram}}: card model and memory size
- {{error_message}}: full OOM text including allocated and reserved figures
- {{model_summary}}: parameter count, layer types, attention type
- {{training_config}}: batch size, sequence length, precision, optimiser, gradient accumulation
- {{code_snippet}}: the forward or training step where it fails
- {{what_you_tried}}: changes already made and their effect
Instructions
- Ask for any missing input, then continue.
- Name the buffers most likely to dominate memory for this configuration.
- Classify the cause: activations, optimiser state, fragmentation, a leaked reference, or another process on the GPU.
- Rank fixes by memory saved against risk to training quality, smallest code change first.
- For each fix give the exact code or config edit and a one-line check that confirms it worked.
- Add instrumentation for allocated and reserved memory per step.
- State what to try if the top fixes do not clear the error.
Output format Numbered sections matching the instructions, short code blocks, and one ranked table of fixes with memory estimate, effort and quality cost. Under 600 words. No GPU background, no praise, no filler.
Guardrails
- Do not invent memory figures, flags or API names. If a flag may differ in the user's version, say so and point them to the framework documentation.
- Flag every fix that changes numerical results or training dynamics.
- Tell the user to confirm the GPU is not shared or holding stale memory before assuming the model is at fault.
Example framework_and_version: PyTorch 2.3; gpu_and_total_vram: A100 40GB; training_config: batch 16, seq 2048, fp16, AdamW; what_you_tried: reduced batch to 8, still fails after a few hundred steps.