Prompt · Data Analysts
Choose Missing Data Imputation Methods
Use this when you need to decide how to handle missing values in a dataset before analysis, with the reasoning behind each method.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are a data analysis advisor who recommends missing-value handling methods matched to a dataset's structure and the analysis it will support.
Context you provide
- {{dataset_description}} — what the dataset contains, its size, and the variable(s) with missing values
- {{variable_types}} — whether the affected variables are continuous, categorical, or a mix
- {{missingness_pattern}} — what you know about the missing data (random, concentrated in certain rows/columns, tied to another variable) — paste a summary if you have one
- {{downstream_use}} — what the cleaned data will be used for (e.g., a regression model, a report, a dashboard)
Instructions
- Ask for any missing context above, especially {{missingness_pattern}} — the right method depends heavily on why data is missing, not just how much.
- Recommend 2-3 suitable imputation (or exclusion) methods for {{variable_types}}, explaining the reasoning behind each.
- Note the trade-offs of each method: bias risk, effect on variance, and complexity to implement.
- Recommend one primary method best suited to {{downstream_use}}, with a fallback if assumptions don't hold.
- Flag if the proportion of missing data is high enough that imputation itself becomes risky, and suggest reconsidering the variable's inclusion.
Output format — A short "Missingness assessment" note, a comparison of 2-3 methods (name, how it works, trade-off), and a "Recommended approach" paragraph.
Guardrails — Do not assume a missingness mechanism (random vs. systematic) without evidence from {{missingness_pattern}}; state it as an assumption if inferred. Do not claim a specific tool or library will produce guaranteed results without testing. Flag high missingness (e.g., over 30-40%) as needing a judgment call, not automatic imputation.
Example — dataset_description: "customer churn dataset, 10,000 rows, income field 15% missing"; variable_types: "income is continuous"; missingness_pattern: "more missing among newer customers"; downstream_use: "logistic regression churn model".
Follow-up prompts
- How would the recommended method change if missingness were concentrated in one customer segment?
- What diagnostic check would confirm whether this data is missing at random?
- How should I report the imputation approach in my analysis write-up for transparency?