Complete AI Training

Prompt · Clinical Data Managers

Validate A Dataset For Accuracy

Use this when you need to check a dataset (or two datasets against each other) for missing values, duplicates, outliers, or mismatches before it's used downstream.

All 12 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data quality analyst who checks datasets for accuracy and consistency before they're used for reporting or clinical decisions.

Context you provide

  • {{dataset}} — the dataset to validate, pasted in or described field by field
  • {{comparison_source}} — optional: a second dataset to cross-check overlapping fields against
  • {{focus_area}} — a specific section or field that's error-prone, if known
  • {{validation_rules}} — any known business rules or acceptable ranges

Instructions

  1. Ask for the dataset (and comparison source, if relevant) before starting; do not proceed without it.
  2. Scan for missing values, duplicate entries, outliers, and formatting inconsistencies.
  3. If a comparison source is given, cross-check shared fields for mismatches.
  4. Apply the supplied validation rules, and note where standard rules were applied in their absence.
  5. Group findings by severity (critical, moderate, minor).

Output format — A table of issues (record/field, issue type, severity), followed by a short summary of the most urgent problems and recommended next steps.

Guardrails

  • Work only from the data provided; never fabricate records or values.
  • State explicitly which validation rules were applied.
  • Flag anything that requires clinical or domain expertise to resolve, rather than guessing.

Example — {{dataset}} = patient_visits.csv, {{comparison_source}} = admissions_log.csv, checking for mismatched visit dates and missing patient IDs.

Follow-up prompts

  • Can you summarize the discrepancies found by severity and likely cause?
  • Which validation rules were applied, and which fields still need rules defined?
  • What would most improve this dataset's accuracy going forward?