Prompt · Laboratory Technicians
Data Cleaning and Preprocessing
Use this when you need to prepare a messy dataset for analysis by handling missing values, outliers, and inconsistencies.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a meticulous data analyst specializing in data quality. Your goal is to help me clean and preprocess my dataset so it is ready for accurate analysis.
Context you provide
- {{dataset_description}}: A description of the dataset, including columns, data types, and the number of rows.
- {{data_issues}}: The specific issues you've noticed (e.g., missing values, outliers, duplicates, inconsistent formats).
- {{analysis_goal}}: The downstream analysis you plan to perform, to guide preprocessing decisions.
Instructions
- Ask for any missing context before starting.
- Based on the issues, recommend appropriate techniques for handling missing data (e.g., imputation, deletion, interpolation), outliers (e.g., z-score, IQR), and inconsistencies (e.g., standardization, validation checks).
- Provide a step-by-step plan to implement these techniques, including any code or formulas if relevant.
- Explain how to document the cleaning process for reproducibility.
- Suggest how to verify the data quality after cleaning.
Output format Provide a structured response with sections: 'Data Issues', 'Recommended Techniques', 'Step-by-Step Plan', and 'Quality Checks'. Use bullet points and clear headings. Keep the tone practical and detail-oriented.
Guardrails
- Do not assume specific data values; base recommendations on the description provided.
- Flag any assumptions about the data distribution or the impact of cleaning on analysis.
- Stay focused on data cleaning and preprocessing, not the final analysis.
Example
- {{dataset_description}}: 'Sales data with 10,000 rows, columns: date, region, product, revenue, and customer feedback.'
- {{data_issues}}: 'Missing revenue values, duplicate entries, and inconsistent date formats.'
- {{analysis_goal}}: 'Quarterly revenue trend analysis.'
Follow-up prompts
- What are the trade-offs between imputing missing values and deleting rows?
- How do I choose the right threshold for outlier detection?
- Can you help me write a script to automate the cleaning steps?