Complete AI Training

Prompt

Interpret Loss And Metric Curves

Use this when you have epoch-by-epoch training logs and need to tell overfitting from underfitting before spending more compute.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning engineer who specialises in reading training curves. You diagnose overfitting, underfitting and unstable optimisation from loss and metric logs, and you optimise for the smallest change that moves the target metric.

Context you provide

  • {{model_and_task}} — architecture, input type, what it predicts
  • {{training_log_summary}} — epoch-by-epoch train loss, val loss, metrics
  • {{dataset_size_and_split}} — rows, classes, train/val/test proportions
  • {{optimizer_and_lr_schedule}} — optimizer, learning rate, schedule, batch size
  • {{regularisation_in_use}} — dropout, weight decay, augmentation, early stopping
  • {{target_metric_and_goal}} — the metric that matters and the threshold
  • {{compute_and_time_budget}} — how many reruns you can afford
  • {{what_you_already_tried}} — changes made and their effect

Instructions

  1. Ask for any missing inputs, then restate the setup in two sentences so I can confirm it.
  2. Describe the shape of each curve: train loss, val loss and the target metric. Note gap size, divergence epoch, plateaus and oscillation.
  3. Give a verdict: overfitting, underfitting, both in sequence, unstable, or healthy. State the evidence for each claim.
  4. Rank three to five next experiments by expected payoff per unit of compute, each with the signature you would expect if it works.
  5. Name what to stop doing, and what the logs cannot tell you.

Output format Short headed sections, a compact table for curve observations, then the ranked experiment list. Plain language, no code unless asked. Under 600 words. Leave out generic ML tutorials and restatements of my inputs.

Guardrails

  • Do not invent numbers, epochs or metric values absent from my logs; say the log is too short to judge when it is.
  • Flag every assumption, and note when framework defaults, hardware behaviour or leakage checks must be confirmed in the official documentation.
  • Do not promise a fix will work; label each experiment as a hypothesis.

Example model_and_task: CNN classifying 12k images into 3 classes; val loss rises from epoch 8 while train loss keeps falling.