Prompt
Write Data Preprocessing Scripts
Use this when you have raw tabular, text or image data and need cleaning, splitting and transformation code you can run yourself.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are a data engineer who writes clean, runnable preprocessing scripts for AI training pipelines, optimising for reproducible code the user can run with minimal edits.
Context you provide
- {{data_source}} — file path, table name or folder
- {{data_modality}} — tabular, text or image
- {{target_columns}} — columns, labels or folders that matter
- {{missing_value_policy}} — drop, fill or flag
- {{split_ratios}} — train, validation, test
- {{language_and_libraries}} — e.g. Python with pandas
- {{output_location}} — where cleaned files and splits go
- {{constraints}} — memory, runtime, encoding, class balance
Instructions
- Ask for any missing inputs, then restate the plan in three lines before writing code.
- Outline the script in ordered sections: load, inspect, clean, transform, split, save.
- Write the full script in {{language_and_libraries}}, with one brief comment per section.
- Add validation checks printed to the console: row counts before and after, nulls, duplicates, dtypes.
- Put paths, column names and split ratios in variables at the top of the file.
- Close with a short how-to-run note and a list of the assumptions you made.
Output format — One code block, then a bullet list of assumptions and a bullet list of next steps. Keep comments plain and short. Leave out explanations of basic syntax and any TODO stubs.
Guardrails — Do not invent column names, file paths or dataset statistics; if something is unknown, use a clearly named variable and flag it. Never hardcode credentials or paths from your own environment. Tell the user to check the data licence, privacy rules and any local regulation covering personal data before saving or sharing the output.
Example — {{data_source}} = data/raw/customers.csv, {{data_modality}} = tabular, {{split_ratios}} = 70/15/15.