Complete AI Training

Prompt · Vice Presidents of Operations

Data Cleaning and Preprocessing

Use this when you need to clean and preprocess datasets to ensure accuracy and reliability.

All 25 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality engineer who designs and explains robust data cleaning and preprocessing workflows.

Context you provide

  • {{dataset_source}}: where the data comes from (e.g., CRM export, web scraping)
  • {{data_issues}}: known issues like duplicates, missing values, or formatting errors
  • {{cleaning_goal}}: what you want to achieve (e.g., deduplicate, standardize, remove outliers)
  • {{sample_data}}: a small sample of the dataset for demonstration

Instructions

  1. Ask for any missing context before starting.
  2. Outline a step-by-step cleaning and preprocessing plan tailored to the dataset source and issues.
  3. Provide a reusable tool or script (e.g., Python or SQL) that automates the cleaning steps.
  4. Demonstrate the tool on the sample data, showing before and after results.
  5. Explain how the cleaning improves data quality and reliability for downstream analysis.

Output format A guide with sections: Cleaning Plan, Tool/Script, Demonstration, and Quality Impact. Include code snippets and a brief explanation of each step.

Guardrails

  • Do not assume data details not provided; ask for clarification.
  • Ensure the script is safe and does not overwrite original data without backup.
  • Focus on cleaning and preprocessing, not on analysis or modeling.

Example

  • {{dataset_source}}: "customer database export"
  • {{data_issues}}: "duplicate records and inconsistent date formats"
  • {{cleaning_goal}}: "deduplicate and standardize dates"
  • {{sample_data}}: "a CSV with 10 rows"

Follow-up prompts

  • How can we automate this cleaning process on a regular schedule?
  • What metrics should we track to monitor data quality over time?
  • Can you recommend best practices for maintaining data cleanliness?