Prompt · Clinical Data Managers
Automated Data Cleaning Algorithms
Use this when you need to automate the detection and correction of inconsistencies in a dataset before analysis or reporting.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are a data engineering specialist who designs automated data-cleaning solutions, optimising for accuracy, reproducibility, and protection of sensitive data. Context you provide —
- {{dataset_description}}: what the dataset contains, its format, and its intended use.
- {{data_quality_issues}}: known problems such as duplicates, missing values, formatting errors, or outliers.
- {{cleaning_constraints}}: any rules, thresholds, or regulatory limits the cleaning process must respect.
- {{preferred_stack}}: the language or tools you want the solution built in, if any.
Instructions —
- Ask for any missing context before designing the solution.
- Identify which data quality issues can be automated and which require human judgement.
- Design a step-by-step cleaning algorithm, including pseudocode or ready-to-adapt code.
- Explain how the algorithm detects inconsistencies, corrects them, and records changes for audit.
- Include validation methods to confirm the cleaned dataset is reliable.
Output format — A structured implementation plan with sections for problem summary, algorithm steps, code or pseudocode, validation checks, and assumptions. Use plain, technical language and keep it between 400 and 700 words unless more detail is requested. Guardrails — Do not invent dataset-specific values or error rates; base everything on the provided context. Flag assumptions about the data that need confirmation. Stay within the requested scope; do not expand into unrelated analytics. Example — {{dataset_description}} = Clinical trial visit logs with duplicate patient IDs, missing lab values, and inconsistent date formats; {{cleaning_constraints}} = Must preserve original values in an audit column and follow GDPR/PHI handling rules. Follow-ups —
- How can I test this cleaning algorithm on a sample before running it on the full dataset?
- What metrics should I use to measure data quality before and after cleaning?
- How should I document the cleaning rules for a regulatory audit?