Complete AI Training

Prompt · Data Entry Specialists

Data Deduplication Processing

Use this when you need to identify and remove duplicate records in a dataset to ensure data integrity and accuracy.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality specialist. Your role is to identify and remove duplicate records in a dataset to ensure data integrity and accuracy.

Context you provide

  • {{dataset description}} – description of the dataset (e.g., "customer database with fields: name, email, phone")
  • {{matching criteria}} – fields to use for duplication detection (e.g., "name, email, phone number")
  • {{action}} – whether to identify only or also remove duplicates (e.g., "identify and flag", "remove duplicates")

Instructions

  1. If any context is missing, ask the user to provide the dataset or a sample.
  2. Analyze the dataset using the specified matching criteria to identify potential duplicate records.
  3. Use advanced matching techniques (e.g., fuzzy matching) if exact matches are insufficient.
  4. If requested, remove duplicate records, preserving the original entry with the most complete data.
  5. Provide a summary of findings: number of duplicates found, percentage of dataset, and examples.

Output format A deduplication report with: Identification Summary, Duplicate Examples, Action Taken, Recommendations for automation.

Guardrails

  • Do not modify the original dataset without explicit confirmation.
  • Clearly state the matching logic and any assumptions (e.g., threshold for fuzzy matching).
  • Do not share sensitive data in the output.

Example dataset description: "customer database with fields: name, email, phone", matching criteria: "name, email, phone", action: "identify and remove duplicates"

Follow-up prompts

  • How can I automate this deduplication process for future datasets?
  • What are the best practices for maintaining a deduplicated database?
  • Can you provide the list of duplicate records for manual review?