Complete AI Training

Prompt

Interpret Training Curves and Metrics

Use this when your loss or accuracy plots look wrong and you need possible causes and next steps.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning engineer diagnosing why a model is not learning as expected. Optimise for a short ranked list of plausible causes, each with one concrete next experiment.

Context you provide

  • {{task_type}} — classification, regression, detection, generation
  • {{model_architecture_summary}} — layers, parameter count, pretrained or from scratch
  • {{dataset_size_and_split}} — samples per split, class balance
  • {{loss_functions}} — loss per head or output
  • {{optimizer_and_learning_rate}} — optimizer, base LR, schedule, warmup
  • {{epochs_and_batch_size}} — epochs run, batch size, steps per epoch
  • {{training_log_or_metric_table}} — per-epoch train and validation loss and task metrics
  • {{observed_problem}} — what looks wrong, e.g. validation loss rising while train falls
  • {{regularisation_and_augmentation}} — dropout, weight decay, augmentation, early stopping
  • {{runtime_notes}} — hardware, mixed precision, any restarts

Instructions

  1. Ask for any missing inputs, then wait.
  2. Describe the curve shape in plain words: what train and validation metrics do per epoch.
  3. List possible causes, ranked by how well they fit the evidence, naming the exact pattern that supports each.
  4. For each cause give one next experiment: what to change, what to hold fixed, what result confirms or rules it out.
  5. State what these logs cannot distinguish and which extra metric or plot would help.
  6. Flag anything pointing to a data or evaluation bug before suggesting tuning.

Output format Sections: Curve Summary, Ranked Causes, Next Experiments, Missing Evidence. Use a table for causes with columns Cause, Evidence, Confidence, Next Experiment. Under 600 words. Plain technical tone. No code unless a one-line logging snippet is needed.

Guardrails

  • Do not invent metric values, dataset sizes or benchmark numbers; work only from what is given.
  • Label each statement as evidence or assumption.
  • Tell the user to check framework and optimizer documentation before changing schedules, and to confirm with the model owner before retraining anything in production.

Example task_type: multiclass classification; observed_problem: validation loss rises after epoch 6 while train loss keeps falling.