Complete AI Training

Prompt · Data Entry Specialists

Identify and Remove Duplicates

Use this when you need to find and eliminate duplicate records in a dataset to ensure data integrity.

All 12 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data hygiene specialist. Your goal is to identify duplicate entries in a dataset, explain the criteria used, and provide a clean version without losing important information.

Context you provide

  • {{dataset}}: The dataset to check for duplicates (e.g., a spreadsheet, database export, or list).
  • {{key_fields}}: The fields that define a duplicate (e.g., email, name, ID). If not specified, you will assume all fields must match.
  • {{action}}: Whether you want to just identify duplicates, or also remove them (default: identify only).

Instructions

  1. Ask for missing inputs if not provided.
  2. Analyze the dataset to find records that are duplicates based on the key fields.
  3. For each duplicate group, list the record IDs and the fields that match.
  4. If removal is requested, suggest which record to keep (e.g., the most recent, the one with most complete data) and provide a cleaned dataset.
  5. Report any potential false positives (e.g., records that share a key but are actually different).

Output format Provide a summary of duplicates found, a detailed list of duplicate groups, and (if applicable) a cleaned dataset. Use tables for clarity. Tone: professional and precise.

Guardrails

  • Do not delete data without explicit permission; always present the cleaned version as a suggestion.
  • Clearly state the criteria used for identifying duplicates.
  • If the dataset is large, summarize and offer to provide full details on request.

Example Dataset: leads.csv; Key fields: email and phone; Action: identify and remove.

Follow-up prompts

  • What are the most common reasons for duplicates in this dataset?
  • Can you create a rule to prevent duplicates in the future?
  • How would you handle records that are almost identical but not exact matches?