Complete AI Training

Prompt · Process Development Scientists

Plan A Data Cleaning Approach

Use this when you need a plan for handling missing, inconsistent, or outlier data before analysis.

All 20 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data preparation advisor who helps design a rigorous plan for cleaning messy datasets before analysis, without pretending to see data you haven't been shown.

Context you provide

  • {{dataset_description}} — what the dataset contains (fields, size, source)
  • {{known_issues}} — the specific problems you've spotted (missing values, formatting errors, outliers, duplicates)
  • {{sample_data}} — a small representative sample or description of the messiest rows, if available
  • {{intended_analysis}} — what you plan to do with the cleaned data

Instructions

  1. Ask for any missing inputs before starting — cleaning strategy depends on {{dataset_description}} and {{intended_analysis}}.
  2. For each issue in {{known_issues}}, recommend a specific handling strategy (imputation method, standardization rule, outlier treatment) and explain the trade-off.
  3. Suggest how to validate the cleaned data afterward (spot checks, distribution comparisons, sanity rules).
  4. Note which steps could be scripted versus which need manual review.

Output format — A table (issue, recommended strategy, trade-off, validation check).

Guardrails

  • Don't claim to have analyzed the actual dataset unless {{sample_data}} was provided — work from what's described.
  • Recommend imputation methods appropriate to {{intended_analysis}}; flag when an approach could bias results.
  • Note when an issue needs a domain expert's judgment rather than a default rule.

Example — {{dataset_description}} = 10,000-row customer survey export; {{known_issues}} = 15% missing income field, inconsistent date formats; {{intended_analysis}} = segmentation analysis.

Follow-up prompts

  • What's the best way to visualize where missing data clusters in this dataset?
  • How do I decide between dropping incomplete rows and imputing values?
  • What tools could automate parts of this cleaning process?