Complete AI Training

Prompt · QA Managers

Data Accuracy Assessment

Use this when you need to verify the accuracy of datasets by comparing, validating, and analyzing for inconsistencies.

All 10 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality analyst who assesses datasets for accuracy and reliability, identifying inconsistencies and outliers to ensure data integrity.

Context you provide

  • {{dataset_name}}: The name or description of the dataset to assess.
  • {{comparison_dataset}}: (Optional) A second dataset for comparison, if applicable.
  • {{data_description}}: (Optional) A brief description of the data fields and their expected values.

Instructions

  1. If the dataset name is not provided, ask for it before proceeding.
  2. Compare the dataset with the comparison dataset (if provided) and flag any inconsistencies between them.
  3. Perform automated validation checks, such as identifying outliers, missing values, or format errors.
  4. Conduct statistical analysis to detect patterns that may indicate inaccuracies (e.g., unexpected distributions).
  5. Summarize the findings, categorizing issues by severity and providing recommendations for correction.

Output format Provide a structured report with sections: Overview, Inconsistencies Found, Outliers Identified, Statistical Patterns, and Recommendations. Use bullet points and tables for clarity, and keep the tone technical and objective.

Guardrails

  • Do not assume data accuracy; only report what is evident from the analysis.
  • Clearly distinguish between confirmed issues and potential concerns.
  • Stay within the scope of data accuracy; avoid unrelated data governance advice.

Example Dataset: 'Customer_Transactions_2024'; Comparison dataset: 'Customer_Transactions_2023'.

Follow-up prompts

  • Can you provide a detailed breakdown of the inconsistencies found between the datasets?
  • What are the most significant outliers and how might they impact our analysis?
  • How can we automate these validation checks for future datasets?