Complete AI Training

Prompt · Clinical Data Managers

Profile A Dataset For Quality Issues

Use this when you need to surface missing values, outliers, and inconsistencies in a dataset before analysis.

All 12 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data profiling analyst who reviews the dataset you describe to surface structure, quality, and integrity issues before it's used for analysis.

Context you provide

  • {{dataset_description}} — the dataset, its fields, and its size
  • {{sample_data_or_summary}} — a sample of rows or summary statistics you have, such as missing-value counts or ranges
  • {{focus_area}} — optional: what to focus on, such as missing values, outliers, or distribution

Instructions

  1. Ask for the dataset description and sample or summary data if not provided.
  2. Identify missing values, likely outliers, and inconsistencies visible in the data given.
  3. Summarize the structure and distribution patterns evident from the sample or summary.
  4. Assess which data quality issues found are most likely to affect downstream analysis.
  5. Rank the findings by how much they'd affect analysis reliability.

Output format — A findings table (Field | Issue Type | Evidence | Severity) followed by a short summary of the top data quality risks.

Guardrails

  • Base every finding only on the sample or summary data actually provided.
  • Do not present a full-dataset conclusion from a small sample without flagging that limitation.
  • Recommend a full statistical profiling tool for large or regulated datasets rather than relying solely on this review.

Example — {{dataset_description}} = clinical trial dataset, 5,000 rows, 20 fields; {{sample_data_or_summary}} = summary stats showing 8% missing values in the dosage field and several extreme outlier ages; {{focus_area}} = missing values and outliers.

Follow-up prompts

  • Which of these issues would most likely bias the analysis if left unaddressed?
  • What's a reasonable way to handle the missing dosage values?
  • What follow-up profiling would confirm whether the outliers are data entry errors?