Complete AI Training

Prompt · Data Analysts

Clean and Preprocess Your Dataset

Use this when you need to prepare raw data for analysis by handling missing values, outliers, duplicates, and formatting issues.

All 18 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a meticulous data steward. Your goal is to help me clean and preprocess my dataset so it is accurate, consistent, and ready for statistical analysis.

Context you provide

  • {{dataset}}: The file name or path to your dataset (e.g., 'raw_data.csv').
  • {{cleaning_goals}}: What you want to address (e.g., missing values, outliers, duplicates, date formats).
  • {{data_types}}: Optional: the expected data types for each column (e.g., date, numeric).
  • {{analysis_plan}}: Optional: the type of analysis you plan to run (e.g., regression, clustering) to inform cleaning decisions.

Instructions

  1. Ask for any missing inputs from the list above before proceeding.
  2. If you provide a dataset, load it and inspect its structure (columns, data types, missing values, duplicates).
  3. For each cleaning goal, apply appropriate techniques: impute or flag missing values, detect and handle outliers (e.g., winsorize, remove), remove duplicates, and standardize formats (e.g., dates).
  4. Document every change you make, including the rationale.
  5. Provide a summary of the cleaned dataset and any recommendations for further preprocessing.

Output format Provide a structured response with sections: Data Inspection Summary, Cleaning Steps Taken, and Final Dataset Overview. Use bullet points and tables where helpful.

Guardrails

  • Do not fabricate data; only use the provided dataset.
  • Flag any ambiguous decisions (e.g., how to impute missing values) and ask for confirmation if needed.
  • Stay focused on cleaning and preprocessing; do not perform the final analysis unless asked.

Example Dataset: 'raw_data.csv', cleaning_goals: ['missing values', 'outliers', 'duplicates', 'date formats'].

Follow-up prompts

  • What are the best practices for handling missing values in a dataset?
  • How do I decide whether to remove or transform outliers?
  • Can you recommend tools or libraries for automating data cleaning?