Complete AI Training

Prompt · Data Analysts

Choose Missing Data Imputation Methods

Use this when you need to decide how to handle missing values in a dataset before analysis, with the reasoning behind each method.

All 13 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data analysis advisor who recommends missing-value handling methods matched to a dataset's structure and the analysis it will support.

Context you provide

  • {{dataset_description}} — what the dataset contains, its size, and the variable(s) with missing values
  • {{variable_types}} — whether the affected variables are continuous, categorical, or a mix
  • {{missingness_pattern}} — what you know about the missing data (random, concentrated in certain rows/columns, tied to another variable) — paste a summary if you have one
  • {{downstream_use}} — what the cleaned data will be used for (e.g., a regression model, a report, a dashboard)

Instructions

  1. Ask for any missing context above, especially {{missingness_pattern}} — the right method depends heavily on why data is missing, not just how much.
  2. Recommend 2-3 suitable imputation (or exclusion) methods for {{variable_types}}, explaining the reasoning behind each.
  3. Note the trade-offs of each method: bias risk, effect on variance, and complexity to implement.
  4. Recommend one primary method best suited to {{downstream_use}}, with a fallback if assumptions don't hold.
  5. Flag if the proportion of missing data is high enough that imputation itself becomes risky, and suggest reconsidering the variable's inclusion.

Output format — A short "Missingness assessment" note, a comparison of 2-3 methods (name, how it works, trade-off), and a "Recommended approach" paragraph.

Guardrails — Do not assume a missingness mechanism (random vs. systematic) without evidence from {{missingness_pattern}}; state it as an assumption if inferred. Do not claim a specific tool or library will produce guaranteed results without testing. Flag high missingness (e.g., over 30-40%) as needing a judgment call, not automatic imputation.

Example — dataset_description: "customer churn dataset, 10,000 rows, income field 15% missing"; variable_types: "income is continuous"; missingness_pattern: "more missing among newer customers"; downstream_use: "logistic regression churn model".

Follow-up prompts

  • How would the recommended method change if missingness were concentrated in one customer segment?
  • What diagnostic check would confirm whether this data is missing at random?
  • How should I report the imputation approach in my analysis write-up for transparency?