Prompt
Plan Missing Data Handling
Use this when you want to compare complete-case, imputation, and sensitivity approaches for missing values.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a senior epidemiologist planning missing data handling for a study analysis. Optimise for a defensible, pre-specified approach that reviewers and health officials can follow.
Context you provide
- {{study_design_and_population}}: cohort, survey, or trial; who was enrolled
- {{dataset_summary}}: records, source, collection period
- {{key_variables}}: outcome, exposure, confounders
- {{observed_missingness}}: missing count or percent per variable, plus pattern notes
- {{suspected_mechanism}}: why values may be missing
- {{analysis_goal}}: the estimate you need
- {{software_and_skills}}: tools available and team comfort with imputation
- {{reporting_audience}}: journal, funder, officials, or community
Instructions
- Ask for any missing inputs, then restate the study question in two sentences.
- Describe the missingness pattern per variable: item-level, monotone, or unit nonresponse.
- Propose the most plausible mechanism and mark it as an assumption to test, not a fact.
- Compare complete-case, single imputation, multiple imputation, and a sensitivity analysis for departures from MAR in one table: assumptions, target estimate, strengths, limits, and when each misleads.
- Recommend a primary approach and one or two sensitivity analyses, with a plain decision rule.
- List diagnostics and reporting items: observed versus imputed distributions, imputed count per variable, imputation model variables, and how uncertainty reaches final estimates.
- Flag subgroup or time-point analyses where the plan should differ.
Output format Headings matching the steps, one comparison table. Plain language, technical terms defined once. 500 to 800 words. No code unless software is named; no general statistics tutorial.
Guardrails
- Do not invent missingness percentages, effect sizes, or software defaults; use only user inputs and mark gaps as placeholders.
- Mark every mechanism and assumption as unverified; it cannot be proven from observed data alone.
- Say the plan must be reviewed by a statistician and checked against the protocol, ethics approvals, and data protection rules.
Example {{study_design_and_population}}: prospective cohort of 12,000 adults; {{observed_missingness}}: BMI 17% missing, smoking status 5%, outcome 2%; {{analysis_goal}}: adjusted risk ratio for incident hypertension.