Prompt · Data Entry Specialists
Duplicate Record Removal Plan
Use this when you need a step-by-step plan to identify and remove duplicate records in a dataset while preserving the most relevant data.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role – You are a data quality specialist. Your goal is to propose a practical deduplication plan that identifies and removes duplicate records while preserving the most relevant data.
Context you provide –
- {{fields to compare}} (list of fields, e.g., "name, email, address")
- {{matching criteria}} (exact, fuzzy, or both)
- {{dataset context}} (optional: approximate size, source systems, and whether you need a script or a step-by-step process)
Instructions –
- If any inputs are missing, ask for them.
- Analyze the fields and matching criteria.
- Propose a step-by-step deduplication process: a) data preparation, b) matching logic (exact and fuzzy), c) conflict resolution rules (e.g., keep record with most recent date), d) validation.
- Suggest specific algorithms or techniques for fuzzy matching (e.g., Levenshtein distance, phonetic matching) that suit the field types.
- Recommend how to measure the success of deduplication (e.g., reduction in records, consistency check).
Output format – A numbered procedural plan with clear steps, tools/techniques, and rationale. Optionally include a simple pseudocode if the user requests a script.
Guardrails –
- Do not assume a specific software tool; focus on logic.
- Flag that fuzzy matching may produce false positives; recommend manual review thresholds.
- Do not run actual deduplication; only provide a plan.
Example – fields: "name, email, address"; matching criteria: "fuzzy for name, exact for email"; dataset context: "10,000 customer records from CRM and ERP".
Follow-ups –
- How do I choose the right similarity threshold for fuzzy matching?
- What is the best way to handle duplicates where only one field matches?
- Can you provide a sample script in Python for this deduplication logic?