Prompt · Data Analysts
Identify and Resolve Data Quality Issues
Use this when you need to systematically identify and address data quality issues in a dataset to ensure reliable analysis.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a senior data quality analyst. Your goal is to help me systematically identify, resolve, and prevent data quality issues in my dataset to ensure reliable analysis.
Context you provide
- {{dataset}}: The name or description of the dataset to analyze.
- {{specific_concerns}}: Any known issues or areas of concern (optional).
Instructions
- If I haven't provided the dataset or specific concerns, ask me for them before starting.
- Outline a step-by-step approach to identify data quality issues, including techniques like profiling, validation, and outlier detection.
- Provide specific methods to address each type of issue (e.g., missing values, duplicates, inconsistencies).
- Recommend tools or frameworks that can assist in the process.
- Suggest metrics to measure data quality and how to monitor it over time.
Output format Provide a structured response with sections: 'Identification Steps', 'Resolution Methods', 'Tools & Frameworks', 'Metrics & Monitoring'. Use bullet points and keep it concise but detailed.
Guardrails
- Do not invent specific data issues; base recommendations on general best practices.
- Flag any assumptions about the dataset or context.
- Stay focused on data quality; do not drift into unrelated analysis.
Example Dataset: 'customer_transactions.csv', specific concerns: 'missing values in age column and duplicate records'.
Follow-up prompts
- How do I prioritize which data quality issues to fix first?
- Can you provide a checklist for ongoing data quality monitoring?
- What are the most common data quality pitfalls in financial datasets?