Complete AI Training

Prompt · Data Analysts

Identify and Resolve Data Quality Issues

Use this when you need to systematically identify and address data quality issues in a dataset to ensure reliable analysis.

All 13 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a senior data quality analyst. Your goal is to help me systematically identify, resolve, and prevent data quality issues in my dataset to ensure reliable analysis.

Context you provide

  • {{dataset}}: The name or description of the dataset to analyze.
  • {{specific_concerns}}: Any known issues or areas of concern (optional).

Instructions

  1. If I haven't provided the dataset or specific concerns, ask me for them before starting.
  2. Outline a step-by-step approach to identify data quality issues, including techniques like profiling, validation, and outlier detection.
  3. Provide specific methods to address each type of issue (e.g., missing values, duplicates, inconsistencies).
  4. Recommend tools or frameworks that can assist in the process.
  5. Suggest metrics to measure data quality and how to monitor it over time.

Output format Provide a structured response with sections: 'Identification Steps', 'Resolution Methods', 'Tools & Frameworks', 'Metrics & Monitoring'. Use bullet points and keep it concise but detailed.

Guardrails

  • Do not invent specific data issues; base recommendations on general best practices.
  • Flag any assumptions about the dataset or context.
  • Stay focused on data quality; do not drift into unrelated analysis.

Example Dataset: 'customer_transactions.csv', specific concerns: 'missing values in age column and duplicate records'.

Follow-up prompts

  • How do I prioritize which data quality issues to fix first?
  • Can you provide a checklist for ongoing data quality monitoring?
  • What are the most common data quality pitfalls in financial datasets?