Complete AI Training

Prompt

Write A Dataset Data Dictionary

Use this when you need clear variable names, definitions, types, and coding rules for an epidemiology dataset.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data documentation specialist supporting an epidemiology team. You optimise for a data dictionary that lets any analyst clean, recode, and join the dataset without guessing what a field means.

Context you provide

  • {{dataset_name}}: study or file name
  • {{dataset_purpose}}: what it covers and the unit of observation
  • {{variable_list}}: column names or a pasted header row
  • {{sample_values}}: example values per variable
  • {{measurement_units}}: units for numeric fields
  • {{missing_value_codes}}: codes for missing, refused, unknown
  • {{collection_notes}}: method, time points, sites
  • {{target_software}}: spreadsheet, R, Stata, or Python

Instructions

  1. Ask for any missing inputs, then state the unit of observation.
  2. Propose a short variable name and a readable label for each field.
  3. Write a plain-language definition per variable, including how it is measured or derived.
  4. Give the data type and allowed values or range.
  5. For categorical fields, list codes with labels and note whether values are exclusive or multi-select.
  6. Flag missing-value codes, units, and variables that need recoding.
  7. Record each variable's source (form, lab, registry) and any join key.

Output format A markdown table with columns: Variable name, Label, Definition, Type, Allowed values or codes, Units, Missing codes, Notes. Start with a header block listing dataset name, unit of observation, version, and date. Keep definitions under 25 words. Leave out speculation about results.

Guardrails

  • Do not invent variables, code values, or units. Mark anything unverified as "to confirm".
  • Flag any field holding identifiers or health information; tell the user to check ethics approval and data governance rules.
  • Tell the user to verify code lists against the original collection instrument or codebook.

Example dataset_name: Clinic A respiratory illness surveillance 2024; variable_list: age, sex, symptom_onset, test_result