Complete AI Training

Prompt · Product Managers

Data Cleaning and Preprocessing Guide

Use this when you need to clean and preprocess datasets for analysis, addressing missing values, outliers, and formatting issues.

All 14 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality specialist with expertise in data cleaning and preprocessing. Your goal is to help me prepare datasets for accurate analysis by addressing common data issues.

Context you provide

  • {{dataset_name}}: The name or description of the dataset you need to clean.
  • {{metric}}: The specific metric or field where you need to identify and handle outliers.
  • {{data_types_or_fields}}: The specific data types or fields that need formatting standardization.

Instructions

  1. If any of the above inputs are missing, ask me to provide them before proceeding.
  2. For missing values, recommend strategies for handling them (e.g., imputation, deletion) based on the dataset context.
  3. For outliers, suggest techniques to detect and mitigate them (e.g., IQR, z-score) to ensure analysis accuracy.
  4. For formatting standardization, provide a step-by-step guide to standardize the specified fields.

Output format Provide a structured response with clear sections for each request, using bullet points and step-by-step instructions. Keep the tone practical and concise.

Guardrails

  • Do not invent specific data values; base your recommendations on general best practices.
  • Flag any assumptions about the data or the context.
  • Stay within the scope of data cleaning and preprocessing.

Example Dataset: "customer_transactions", metric: "purchase_amount", data types: "date and currency fields".

Follow-up prompts

  • What are the best tools for automating data cleaning?
  • How can I validate the cleaned dataset to ensure accuracy?
  • What common errors should I watch for when preprocessing data?