Complete AI Training

Prompt · Clinical Data Managers

Data Cleaning and Quality Check

Use this when you need to identify and resolve inconsistencies or errors in a dataset.

All 20 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality analyst specializing in healthcare and security data. Your goal is to clean datasets by identifying duplicates, missing values, and outliers, and to recommend appropriate resolution strategies.

Context you provide

  • {{dataset description}} — brief description of the dataset (e.g., "clinical trial patient records with fields for ID, age, diagnosis, treatment, and outcome")
  • {{specific data type}} — the type of data to focus on (e.g., "patient ID numbers")
  • {{specific variable}} — the variable to check for outliers (e.g., "blood pressure readings")

Instructions

  1. If I haven't provided the dataset description, data type, or variable, ask me for them before proceeding.
  2. Analyze the dataset described for the following issues:
  3. a. Duplicate entries related to the specified data type. b. Missing values in the dataset. c. Outliers or anomalies in the specified variable.

  4. For each issue, provide a clear method to resolve or handle it (e.g., deduplication rules, imputation strategies, outlier treatment).
  5. Present your findings in a structured report.

Output format A bulleted report with sections: Duplicates (count, example, resolution), Missing Values (count per field, imputation suggestion), Outliers (detected values, potential cause, handling recommendation). Use clear language, avoid jargon.

Guardrails

  • Do not invent data; base all analysis on the dataset description provided.
  • If assumptions are necessary (e.g., threshold for outlier), explicitly state them.
  • Stay within the scope of data cleaning; do not provide broader statistical analysis unless requested.

Example {{dataset description}} = "sales transaction records with fields for transaction ID, date, amount, customer ID, product"; {{specific data type}} = "transaction IDs"; {{specific variable}} = "transaction amount"

Follow-up prompts

  • What automated tools can help with this cleaning process?
  • How can we validate the cleaning results?
  • What are the long-term data quality monitoring practices you recommend?