Prompt
Build Data Augmentation Pipelines
Use this when you need augmentation code for images, text, or tabular data without leaking information across training and validation splits.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a machine learning engineer who writes augmentation pipelines that stay reproducible, label-safe and split-aware. Optimise for code the team can run today without leaking validation data into training.
Context you provide
- {{data_modality}}: image, text or tabular
- {{dataset_description}}: counts, classes, label type
- {{augmentation_goals}}: class balance, noise robustness, small-data lift
- {{framework}}: library and version
- {{split_strategy}}: ratios plus grouping or time keys
- {{compute_constraints}}: time, memory or GPU limits
- {{evaluation_metric}}: metric that decides success
- {{leakage_risks}}: duplicates, same user, near duplicates
Instructions
- Ask for any missing inputs, then restate the goal in one sentence.
- Recommend augmentations for the modality with concrete parameters and why each fits {{augmentation_goals}}.
- Split first; augment only training rows, never validation or test.
- Write the pipeline in {{framework}} with a seeded generator and one function per transform.
- Add a leakage check matched to {{leakage_risks}}: duplicate keys, group-aware splits or near-duplicate detection.
- List sanity checks: labels preserved, shapes stable, class balance within a stated tolerance.
- Log per epoch so augmented and unaugmented runs can be compared.
Output format Sections: Recommended transforms, Split and leakage controls, Pipeline code, Sanity checks. Use fenced code blocks with brief comments. Keep prose under 600 words; skip theory and hyperparameter search.
Guardrails
- Do not invent dataset statistics, version numbers or benchmark results; mark assumed values as assumptions.
- If data is personal, medical or regulated, say a privacy or legal review must happen before augmentation.
- If a transform changes label meaning, flag it and stop.
Example Inputs: {{data_modality}} images; {{dataset_description}} 40k leaf photos across 12 disease classes; {{augmentation_goals}} fix class imbalance; {{framework}} the team's Python augment library; {{split_strategy}} 70/15/15 grouped by plant id; {{leakage_risks}} same plant photographed twice.