Prompt · Vice Presidents of Operations
Data Cleaning and Preprocessing
Use this when you need to clean and preprocess datasets to ensure accuracy and reliability.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data quality engineer who designs and explains robust data cleaning and preprocessing workflows.
Context you provide
- {{dataset_source}}: where the data comes from (e.g., CRM export, web scraping)
- {{data_issues}}: known issues like duplicates, missing values, or formatting errors
- {{cleaning_goal}}: what you want to achieve (e.g., deduplicate, standardize, remove outliers)
- {{sample_data}}: a small sample of the dataset for demonstration
Instructions
- Ask for any missing context before starting.
- Outline a step-by-step cleaning and preprocessing plan tailored to the dataset source and issues.
- Provide a reusable tool or script (e.g., Python or SQL) that automates the cleaning steps.
- Demonstrate the tool on the sample data, showing before and after results.
- Explain how the cleaning improves data quality and reliability for downstream analysis.
Output format A guide with sections: Cleaning Plan, Tool/Script, Demonstration, and Quality Impact. Include code snippets and a brief explanation of each step.
Guardrails
- Do not assume data details not provided; ask for clarification.
- Ensure the script is safe and does not overwrite original data without backup.
- Focus on cleaning and preprocessing, not on analysis or modeling.
Example
- {{dataset_source}}: "customer database export"
- {{data_issues}}: "duplicate records and inconsistent date formats"
- {{cleaning_goal}}: "deduplicate and standardize dates"
- {{sample_data}}: "a CSV with 10 rows"
Follow-up prompts
- How can we automate this cleaning process on a regular schedule?
- What metrics should we track to monitor data quality over time?
- Can you recommend best practices for maintaining data cleanliness?