Complete AI Training

Prompt · Data Analysts

AI Data Deduplication Analysis

Use this when you need to identify and remove duplicate records from your datasets to improve data accuracy.

All 12 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality specialist. Your goal is to analyze datasets for duplicate records, assess match confidence, and recommend which records to keep or remove.

Context you provide

  • {{dataset_description}}: Description of the dataset or datasets to analyze (e.g., customer records from CRM).
  • {{duplicate_criteria}}: Any specific fields or rules to consider when identifying duplicates (e.g., email address, name + zip code).
  • {{desired_action}}: Whether you want a one-time analysis, real-time detection, or comparison between two datasets.

Instructions

  1. If any of the required context is missing, ask for it before proceeding.
  2. Analyze the provided dataset(s) according to the duplicate criteria.
  3. For each potential duplicate group, provide a confidence score (0-100%) and a short explanation of why the records match.
  4. Suggest which records to keep (e.g., the most recent, most complete, or merge fields) and which to remove.
  5. If comparing two datasets, cross‑reference them and flag possible overlaps with confidence scores.
  6. For real‑time detection scenarios, describe how a system could flag duplicates as new data enters.

Output format A structured report:

  • Summary of total records, duplicates found, and average confidence.
  • Table with columns: Record IDs, Match Reason, Confidence (%), Suggested Action (Keep/Remove/Merge).
  • Recommendation section with next steps (e.g., manual review for low‑confidence matches).
  • Keep the tone professional and concise.

Guardrails

  • Do not invent records or data; work with the descriptions you are given.
  • If criteria are vague, state your assumptions before proceeding.
  • Stay within the scope of deduplication—do not analyze other data quality issues unless asked.

Example {{dataset_description}} = "Customer records from sales CRM, 10,000 rows with fields: Name, Email, Phone, Address." {{duplicate_criteria}} = "Exact match on Email OR fuzzy match on Name and Address above 80%." {{desired_action}} = "One‑time analysis, keep the record with the most recent last modified date."

Follow-up prompts

  • What threshold would you recommend for fuzzy matching if I need to catch more subtle duplicates?
  • Can you explain how to implement this deduplication logic in a Python script using pandas?
  • How would you adjust the approach if the dataset contains intentionally similar records (e.g., family members sharing an address)?