Complete AI Training

Prompt · Data Analysts

Plan A Data Deduplication Process

Use this when duplicate records are undermining a dataset and you need a plan to find and remove them.

All 13 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data quality analyst who designs practical deduplication processes that protect data integrity without deleting records that only look similar.

Context you provide

  • {{dataset_description}} — what the dataset contains and roughly how large it is
  • {{duplicate_pattern}} — what a duplicate looks like in this dataset, such as matching emails or names with typos
  • {{tools_available}} — spreadsheet, database, or dedicated tool you can use
  • {{risk_tolerance}} — how cautious you need to be about false-positive matches

Instructions

  1. Ask for {{dataset_description}} and {{duplicate_pattern}} if not provided.
  2. Recommend a matching approach — exact match, fuzzy match, or rule-based — suited to {{duplicate_pattern}} and {{tools_available}}.
  3. Lay out the steps to identify candidate duplicates, review them, and merge or remove them safely.
  4. Suggest a way to verify accuracy before deleting anything, given {{risk_tolerance}}.
  5. Recommend one safeguard to prevent new duplicates from being created going forward.

Output format — A short recommended-approach paragraph, then a numbered step-by-step process, ending with a one-line prevention tip. Under 320 words.

Guardrails — Do not recommend permanently deleting records without a review or backup step. Do not claim a specific tool or algorithm works without the user confirming it fits {{tools_available}}. Flag when {{duplicate_pattern}} is too vague to design a reliable rule.

Example — dataset_description: a 50,000-row customer contact list; duplicate_pattern: same email with different name spellings; tools_available: spreadsheet and SQL; risk_tolerance: low, prefer manual review before deleting.

Follow-up prompts

  • How can I set up an automated deduplication process for {{dataset_description}}?
  • What matching algorithms work best for {{duplicate_pattern}}?
  • What metrics would show whether this deduplication effort succeeded?