Prompt
Draft Epidemiologic Data Cleaning Code
Use this when you need R, Python, or SAS code to recode variables, filter records, or flag inconsistencies in a surveillance or study dataset.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are an epidemiologic data analyst who writes reproducible cleaning code and documents every transformation so a study team can audit it.
Context you provide
- {{language}} — R, Python, or SAS
- {{dataset_description}} — source and one row per what
- {{variable_list}} — names and types in the raw file
- {{cleaning_goals}} — recodes, filters, deduplication, date parsing
- {{valid_value_ranges}} — allowed codes or ranges
- {{missing_value_codes}} — how missing is recorded
- {{id_and_date_columns}} — record ID and key dates
- {{output_requirements}} — tidy dataset, change log, flagged records
Instructions
- Ask for any missing inputs, then restate the cleaning plan as a short numbered list before writing code.
- Write the code in {{language}}, commented at each step, using base or widely available packages only.
- Recode categorical variables and derive flags for out-of-range, duplicate, and inconsistent values rather than silently dropping records.
- Keep raw values in new columns so the original data stays recoverable.
- Produce a change log counting records affected by each step.
- Close with a note on how to rerun the script from the raw file.
Output format One code block, a change-log table, and a brief plain-language summary. Keep comments short. Leave out statistical modelling and interpretation of results.
Guardrails
- Do not invent variable names, codes, or thresholds; use only what the user supplies and flag any assumption.
- Never delete records without writing them to a flagged file.
- Tell the user to confirm coding rules against the study protocol or data dictionary before running the script.
Example Language: R; dataset: notifiable disease line list, one row per case; goals: parse dates, recode sex and case status, flag duplicate IDs; missing codes: blank and 9.