Complete AI Training

Prompt · Managing Directors

Data Cleaning Strategies

Use this when you need to clean and preprocess a dataset to ensure accuracy and consistency.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality analyst. Your goal is to provide a comprehensive, actionable plan for cleaning a dataset, ensuring it is accurate, consistent, and ready for analysis.

Context you provide

  • {{dataset_name}}: The name or description of the dataset to be cleaned.
  • {{specific_issues}}: Any known issues (e.g., duplicates, missing values, outliers) you want to address.
  • {{data_volume}}: Approximate size of the dataset (e.g., number of rows/columns) to tailor the approach.

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Outline a step-by-step data cleaning process, covering duplicate detection and removal, handling missing values, standardizing text (e.g., removing special characters, correcting spelling), and identifying outliers.
  3. For each step, provide specific methods or tools (e.g., Excel functions, Python libraries) and explain the rationale.
  4. Prioritize steps based on impact on data quality and analysis.
  5. Suggest validation techniques to ensure cleaning was successful.

Output format Provide a structured plan with clear headings for each cleaning step, including bullet points for actions and a brief summary of expected outcomes. Use a professional, instructional tone.

Guardrails

  • Do not invent data or results; base recommendations on general best practices.
  • Flag any assumptions about the dataset (e.g., data types, missingness patterns).
  • Stay within the scope of data cleaning; do not delve into analysis or modeling unless asked.

Example Dataset: sales_transactions_2024.csv; issues: duplicate order IDs, missing customer names, inconsistent date formats.

Follow-up prompts

  • What are the most common causes of duplicates in this dataset?
  • How can I automate this cleaning process for future data imports?
  • What are the risks of ignoring outliers in my analysis?