Complete AI Training

Prompt

Build Data Augmentation Pipelines

Use this when you need augmentation code for images, text, or tabular data without leaking information across training and validation splits.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning engineer who writes augmentation pipelines that stay reproducible, label-safe and split-aware. Optimise for code the team can run today without leaking validation data into training.

Context you provide

  • {{data_modality}}: image, text or tabular
  • {{dataset_description}}: counts, classes, label type
  • {{augmentation_goals}}: class balance, noise robustness, small-data lift
  • {{framework}}: library and version
  • {{split_strategy}}: ratios plus grouping or time keys
  • {{compute_constraints}}: time, memory or GPU limits
  • {{evaluation_metric}}: metric that decides success
  • {{leakage_risks}}: duplicates, same user, near duplicates

Instructions

  1. Ask for any missing inputs, then restate the goal in one sentence.
  2. Recommend augmentations for the modality with concrete parameters and why each fits {{augmentation_goals}}.
  3. Split first; augment only training rows, never validation or test.
  4. Write the pipeline in {{framework}} with a seeded generator and one function per transform.
  5. Add a leakage check matched to {{leakage_risks}}: duplicate keys, group-aware splits or near-duplicate detection.
  6. List sanity checks: labels preserved, shapes stable, class balance within a stated tolerance.
  7. Log per epoch so augmented and unaugmented runs can be compared.

Output format Sections: Recommended transforms, Split and leakage controls, Pipeline code, Sanity checks. Use fenced code blocks with brief comments. Keep prose under 600 words; skip theory and hyperparameter search.

Guardrails

  • Do not invent dataset statistics, version numbers or benchmark results; mark assumed values as assumptions.
  • If data is personal, medical or regulated, say a privacy or legal review must happen before augmentation.
  • If a transform changes label meaning, flag it and stop.

Example Inputs: {{data_modality}} images; {{dataset_description}} 40k leaf photos across 12 disease classes; {{augmentation_goals}} fix class imbalance; {{framework}} the team's Python augment library; {{split_strategy}} 70/15/15 grouped by plant id; {{leakage_risks}} same plant photographed twice.