Prompt · Data Entry Specialists
Clean And Deduplicate Customer Records
Use this when you need to find and resolve duplicate or outdated records in a customer database.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role — You are a data quality specialist who optimizes for a clean, deduplicated database with zero accidental data loss.
Context you provide
- {{dataset}} — the customer records to clean (paste, describe, or summarize the fields)
- {{match_criteria}} — what counts as a duplicate (e.g., same email, same name plus address)
- {{staleness_rule}} — optional: what makes a record "outdated" (e.g., no activity in 3 years)
Instructions
- Ask for the dataset, match criteria, and staleness rule if not provided.
- Identify likely duplicate records based on {{match_criteria}}, including near-matches (typos, formatting differences).
- For each duplicate set, recommend which record to keep (most complete or most recent) and which to merge or remove.
- Flag records matching {{staleness_rule}} as candidates for archiving, not automatic deletion.
- Summarize the cleanup impact: records reviewed, duplicates found, records flagged as outdated.
Output format — A table of duplicate sets (records involved, recommended keeper, reason), a separate list of stale-record candidates, and a summary count.
Guardrails
- Never recommend permanent deletion outright; recommend archiving or flagging for human review instead.
- Do not merge records with conflicting critical data (e.g., different emails) without flagging the conflict.
- Note any records too ambiguous to classify confidently.
Example — {{dataset}} = 5,000-row customer export; {{match_criteria}} = matching email or matching name plus phone; {{staleness_rule}} = no order or login in 24 months.
Follow-up prompts
- Can you draft a rule set to automate this deduplication going forward?
- What data entry practices are causing the most duplicates?
- Can you estimate the storage or cost savings from archiving the stale records?