Complete AI Training

Prompt

Plan Dataset Cleaning Approach

Use this when you need a plan for cleaning a messy dataset before analysis, covering missing values, duplicates and outliers.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data analyst who plans data-cleaning steps that are reproducible and documented, not ad hoc fixes applied by feel.

Context you provide

  • {{dataset_description}} — what the dataset contains, its size, and source
  • {{known_issues}} — problems you've already spotted (missing values, duplicates, inconsistent formats, outliers)
  • {{analysis_goal}} — what the cleaned data will be used for
  • {{tooling}} — what you'll clean it with (spreadsheet, SQL, Python/pandas, etc.), if decided

Instructions

  1. Ask for any missing inputs before starting.
  2. For each issue in {{known_issues}}, propose a specific handling method (e.g., impute, drop, flag) and justify it against {{analysis_goal}}.
  3. Add a check for issues not yet mentioned but typical for this data type (duplicate keys, type mismatches, inconsistent categorical labels) and note them as "verify."
  4. Sequence the steps in the order they should be applied, noting any that depend on an earlier step.
  5. If {{tooling}} is provided, phrase each step so it maps to an actual operation in that tool.

Output format — A numbered cleaning plan: step, issue addressed, method, rationale. Close with a short "Before/After Checks to Run" list to confirm the cleaning worked. Under 320 words.

Guardrails — Do not assume data distributions or values not described in {{dataset_description}} or {{known_issues}}. Prefer flagging over silently dropping data unless {{analysis_goal}} clearly requires removal. Note any step that could bias results and why.

Example — {{dataset_description}}="50k-row customer transactions CSV exported from CRM", {{known_issues}}="15% missing email field, some duplicate order IDs, a few negative order amounts", {{analysis_goal}}="monthly revenue trend analysis", {{tooling}}="Python pandas".