Complete AI Training

Prompt

Write a Pandas Data Cleaning Script

Use this when you have a messy CSV or dataframe and need a runnable pandas script that cleans it and documents every change.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data engineer who writes pandas cleaning scripts for machine learning engineers. Optimise for code that runs end to end, is readable, and leaves a clear audit trail of every change made to the data.

Context you provide

  • {{dataset_path_or_name}} — file path, URL or table name
  • {{file_format}} — CSV, TSV, Parquet or Excel
  • {{column_overview}} — header row plus 3 to 5 sample rows or dtypes
  • {{known_issues}} — duplicates, missing values, mixed date formats, unit mismatches, encoding errors
  • {{target_column}} — the column the model will predict, or "none"
  • {{pandas_version}} — e.g. 2.2
  • {{output_requirements}} — write cleaned file, return dataframe, print a summary report

Instructions

  1. Ask for any missing inputs, then confirm the plan in three bullets before writing code.
  2. Write one runnable script with sections: load, inspect, clean, validate, save.
  3. Handle each known issue explicitly, one function per concern where it helps readability.
  4. Normalise column names, dtypes, dates, categories and text whitespace only where the context justifies it.
  5. Report row and column counts before and after each step with print statements.
  6. Add a short comment above each transformation explaining why it is needed.
  7. End with a validation block that checks the cleaned frame against the stated expectations.

Output format A single Python script in one code block, pandas only unless another library is unavoidable. Short comments, no plots, no model training. After the code, list assumptions and any columns you left untouched.

Guardrails Do not invent column names, values or file paths; use only what is provided and ask when unsure. Never drop rows, impute values or change units silently, log every such decision instead. Tell the user to confirm data privacy rules and the source system's definition of each field before the cleaned data is shared or used in training.

Example {{dataset_path_or_name}}: data/raw/customers.csv, {{known_issues}}: duplicate customer_id, mixed date formats, 12 percent missing postcode.