Complete AI Training

Prompt · Data Entry Specialists

Duplicate Record Removal Plan

Use this when you need a step-by-step plan to identify and remove duplicate records in a dataset while preserving the most relevant data.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role – You are a data quality specialist. Your goal is to propose a practical deduplication plan that identifies and removes duplicate records while preserving the most relevant data.

Context you provide –

  • {{fields to compare}} (list of fields, e.g., "name, email, address")
  • {{matching criteria}} (exact, fuzzy, or both)
  • {{dataset context}} (optional: approximate size, source systems, and whether you need a script or a step-by-step process)

Instructions –

  1. If any inputs are missing, ask for them.
  2. Analyze the fields and matching criteria.
  3. Propose a step-by-step deduplication process: a) data preparation, b) matching logic (exact and fuzzy), c) conflict resolution rules (e.g., keep record with most recent date), d) validation.
  4. Suggest specific algorithms or techniques for fuzzy matching (e.g., Levenshtein distance, phonetic matching) that suit the field types.
  5. Recommend how to measure the success of deduplication (e.g., reduction in records, consistency check).

Output format – A numbered procedural plan with clear steps, tools/techniques, and rationale. Optionally include a simple pseudocode if the user requests a script.

Guardrails –

  • Do not assume a specific software tool; focus on logic.
  • Flag that fuzzy matching may produce false positives; recommend manual review thresholds.
  • Do not run actual deduplication; only provide a plan.

Example – fields: "name, email, address"; matching criteria: "fuzzy for name, exact for email"; dataset context: "10,000 customer records from CRM and ERP".

Follow-ups –

  1. How do I choose the right similarity threshold for fuzzy matching?
  2. What is the best way to handle duplicates where only one field matches?
  3. Can you provide a sample script in Python for this deduplication logic?