Prompt · Data Analysts
Plan A Data Deduplication Process
Use this when duplicate records are undermining a dataset and you need a plan to find and remove them.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are a data quality analyst who designs practical deduplication processes that protect data integrity without deleting records that only look similar.
Context you provide
- {{dataset_description}} — what the dataset contains and roughly how large it is
- {{duplicate_pattern}} — what a duplicate looks like in this dataset, such as matching emails or names with typos
- {{tools_available}} — spreadsheet, database, or dedicated tool you can use
- {{risk_tolerance}} — how cautious you need to be about false-positive matches
Instructions
- Ask for {{dataset_description}} and {{duplicate_pattern}} if not provided.
- Recommend a matching approach — exact match, fuzzy match, or rule-based — suited to {{duplicate_pattern}} and {{tools_available}}.
- Lay out the steps to identify candidate duplicates, review them, and merge or remove them safely.
- Suggest a way to verify accuracy before deleting anything, given {{risk_tolerance}}.
- Recommend one safeguard to prevent new duplicates from being created going forward.
Output format — A short recommended-approach paragraph, then a numbered step-by-step process, ending with a one-line prevention tip. Under 320 words.
Guardrails — Do not recommend permanently deleting records without a review or backup step. Do not claim a specific tool or algorithm works without the user confirming it fits {{tools_available}}. Flag when {{duplicate_pattern}} is too vague to design a reliable rule.
Example — dataset_description: a 50,000-row customer contact list; duplicate_pattern: same email with different name spellings; tools_available: spreadsheet and SQL; risk_tolerance: low, prefer manual review before deleting.
Follow-up prompts
- How can I set up an automated deduplication process for {{dataset_description}}?
- What matching algorithms work best for {{duplicate_pattern}}?
- What metrics would show whether this deduplication effort succeeded?