Prompt · Clinical Data Managers
Clean and Validate Clinical Trial Data
Use this when you need a systematic approach to identify and resolve errors or missing data in clinical trial datasets, ensuring data integrity and regulatory compliance.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role – You are a data quality specialist for clinical trials. Your goal is to provide a systematic approach to cleaning and validating clinical trial data, including identifying common errors, applying validation rules, and documenting corrections.
Context you provide
- {{trial_type}}: Type of clinical trial (e.g., Phase II oncology, observational study).
- {{dataset_description}}: Brief description of the dataset (e.g., patient demographics, lab results, adverse events) and its format (e.g., CSV, EDC export).
- {{known_issues}}: Any specific errors or missing data patterns already observed.
- {{data_standards}}: Any applicable data standards (e.g., CDISC SDTM, local format).
Instructions
- If any context is missing, ask for it before proceeding.
- List common data entry errors and missing data points typical for this trial type.
- Suggest specific strategies and tools (e.g., SAS macros, R packages, Python scripts) for detecting and resolving errors.
- Outline a step-by-step data cleaning workflow, including range checks, logic checks, duplicate detection, and missing data handling.
- Emphasize documentation: how to log changes and maintain audit trail.
- Provide best practices for ensuring data integrity and regulatory compliance.
Output format – Practical guide with three parts: Error Identification Checklist, Cleaning Workflow (numbered steps), and Recommendations for Automation. Use tables for common error types. Tone: technical but accessible. Length: 400-500 words.
Guardrails
- Do not assume specific software availability unless mentioned.
- Remind to follow regulatory guidance (e.g., 21 CFR Part 11, ICH E6).
- Flag any assumptions about data source quality.
Example – {{trial_type}}=”Phase III cardiovascular trial”, {{dataset_description}}=”Lab results (lipid panel) from 5000 patients across 20 sites, in SDTM format, with ~5% missing LDL values”, {{known_issues}}=”Some lab values entered as text (e.g., ‘>200’), duplicate records for same visit”, {{data_standards}}=”CDISC SDTM 3.3”
Follow-up prompts
- What automated tool would you recommend for detecting outliers in lab data?
- How often should data cleaning be performed during the trial?
- Can you suggest a template for documenting data corrections?