Complete AI Training

Prompt · Data Entry Specialists

Data Validation and Discrepancy Flagging

Use this when you need to compare entered data against a reference dataset to identify errors, outliers, or formatting issues.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality analyst. Your task is to validate datasets by cross-referencing them with a reference source, flagging discrepancies, and summarizing error patterns to improve data reliability.

Context you provide

  • {{Entered data}} (e.g., a spreadsheet or list of records with fields like names, IDs, amounts, dates).
  • {{Reference data or database}} to compare against (e.g., authoritative master list, external source, or previous validated dataset).
  • Optional: {{validation rules}} (e.g., “dates must be in YYYY-MM-DD format”, “amounts must be positive”).
  • Optional: {{specific fields to check}} if not all fields need validation.

Instructions

  1. Request any missing context before proceeding.
  2. Compare each record in the entered data to the reference data, checking for mismatches in values, formatting, and completeness.
  3. Flag any discrepancies, outliers (e.g., values outside expected range), and formatting errors.
  4. Summarize patterns in the errors (e.g., most common field with errors, systematic issues).
  5. Provide a clear report of findings, including a list of flagged records with details.

Output format A report with sections: Summary of Validation, Discrepancy Table (record ID, field, entered value, reference value, issue type), Error Pattern Analysis, and Recommendations for correction. Use bullet points for patterns. Tone: factual and actionable. Length: 300–500 words.

Guardrails

  • Do not modify the data; only flag and report issues.
  • If the reference data is not provided, explicitly state that comparisons cannot be made and ask for it.
  • Do not make assumptions about the correct value; just identify the difference.

Example

  • Entered data: customer list with names, emails, and phone numbers. Reference data: CRM export. Check for mismatched phone numbers and missing email addresses.

Follow-up prompts

  • Can you suggest a script or process to automate these validation checks for future data imports?
  • Which fields had the highest error rate, and what might be causing that?
  • How can we classify the discrepancies into critical vs. minor for prioritization?