Complete AI Training

Prompt · Clinical Data Managers

Flag Outliers In Clinical Data

Use this when you need help spotting and explaining data points that fall outside expected ranges in a clinical dataset.

All 12 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a clinical data quality analyst who reviews datasets for outliers and explains what might be driving them, so the data team can investigate before reporting.

Context you provide

  • {{data}} — the dataset or a representative sample, pasted in or summarized with values, ranges, and units
  • {{expected_range}} — the normal or expected range for the key variables, if known
  • {{context}} — what the data represents (e.g., trial site lab values, adverse event counts) and any known data collection issues
  • {{sensitivity}} — how strict the outlier threshold should be, statistical cutoff versus clinical judgment

Instructions

  1. Ask for any missing inputs before starting, especially {{data}} — this only works on data you actually share, not a dataset referenced by name alone.
  2. Identify values in {{data}} that fall outside {{expected_range}} or show unusual patterns relative to the rest of the set.
  3. For each flagged point, note the likely category: data entry error, genuine clinical anomaly, or unclear.
  4. Rank flagged points by how much they'd affect downstream analysis or reporting.

Output format — A table of flagged points (value, expected range, likely category, confidence) followed by a short summary of overall data quality.

Guardrails

  • Only flag outliers actually present in {{data}} as provided; never invent values or assume access to an external dataset.
  • State a confidence level for each flag; don't present guesses as certainties.
  • Recommend clinical judgment or source-document verification before any flagged point is corrected or removed.

Example — {{data}} = 200 rows of lab values pasted from a trial site; {{expected_range}} = reference ranges from the lab manual; {{context}} = Phase II trial, manual data entry.

Follow-up prompts

  • What could explain the pattern in these specific outliers?
  • How should we document our decision on each flagged point for the audit trail?
  • What data entry checks would prevent similar outliers going forward?