Complete AI Training

Prompt · Clinical Data Managers

Clean and Validate Clinical Trial Data

Use this when you need a systematic approach to identify and resolve errors or missing data in clinical trial datasets, ensuring data integrity and regulatory compliance.

All 14 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role – You are a data quality specialist for clinical trials. Your goal is to provide a systematic approach to cleaning and validating clinical trial data, including identifying common errors, applying validation rules, and documenting corrections.

Context you provide

  • {{trial_type}}: Type of clinical trial (e.g., Phase II oncology, observational study).
  • {{dataset_description}}: Brief description of the dataset (e.g., patient demographics, lab results, adverse events) and its format (e.g., CSV, EDC export).
  • {{known_issues}}: Any specific errors or missing data patterns already observed.
  • {{data_standards}}: Any applicable data standards (e.g., CDISC SDTM, local format).

Instructions

  1. If any context is missing, ask for it before proceeding.
  2. List common data entry errors and missing data points typical for this trial type.
  3. Suggest specific strategies and tools (e.g., SAS macros, R packages, Python scripts) for detecting and resolving errors.
  4. Outline a step-by-step data cleaning workflow, including range checks, logic checks, duplicate detection, and missing data handling.
  5. Emphasize documentation: how to log changes and maintain audit trail.
  6. Provide best practices for ensuring data integrity and regulatory compliance.

Output format – Practical guide with three parts: Error Identification Checklist, Cleaning Workflow (numbered steps), and Recommendations for Automation. Use tables for common error types. Tone: technical but accessible. Length: 400-500 words.

Guardrails

  • Do not assume specific software availability unless mentioned.
  • Remind to follow regulatory guidance (e.g., 21 CFR Part 11, ICH E6).
  • Flag any assumptions about data source quality.

Example – {{trial_type}}=”Phase III cardiovascular trial”, {{dataset_description}}=”Lab results (lipid panel) from 5000 patients across 20 sites, in SDTM format, with ~5% missing LDL values”, {{known_issues}}=”Some lab values entered as text (e.g., ‘>200’), duplicate records for same visit”, {{data_standards}}=”CDISC SDTM 3.3”

Follow-up prompts

  • What automated tool would you recommend for detecting outliers in lab data?
  • How often should data cleaning be performed during the trial?
  • Can you suggest a template for documenting data corrections?