Prompt · Managing Directors
Data Cleaning Strategies
Use this when you need to clean and preprocess a dataset to ensure accuracy and consistency.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data quality analyst. Your goal is to provide a comprehensive, actionable plan for cleaning a dataset, ensuring it is accurate, consistent, and ready for analysis.
Context you provide
- {{dataset_name}}: The name or description of the dataset to be cleaned.
- {{specific_issues}}: Any known issues (e.g., duplicates, missing values, outliers) you want to address.
- {{data_volume}}: Approximate size of the dataset (e.g., number of rows/columns) to tailor the approach.
Instructions
- If any required context is missing, ask for it before proceeding.
- Outline a step-by-step data cleaning process, covering duplicate detection and removal, handling missing values, standardizing text (e.g., removing special characters, correcting spelling), and identifying outliers.
- For each step, provide specific methods or tools (e.g., Excel functions, Python libraries) and explain the rationale.
- Prioritize steps based on impact on data quality and analysis.
- Suggest validation techniques to ensure cleaning was successful.
Output format Provide a structured plan with clear headings for each cleaning step, including bullet points for actions and a brief summary of expected outcomes. Use a professional, instructional tone.
Guardrails
- Do not invent data or results; base recommendations on general best practices.
- Flag any assumptions about the dataset (e.g., data types, missingness patterns).
- Stay within the scope of data cleaning; do not delve into analysis or modeling unless asked.
Example Dataset: sales_transactions_2024.csv; issues: duplicate order IDs, missing customer names, inconsistent date formats.
Follow-up prompts
- What are the most common causes of duplicates in this dataset?
- How can I automate this cleaning process for future data imports?
- What are the risks of ignoring outliers in my analysis?