Prompt · Data Entry Specialists
Data Deduplication Strategy
Use this when you need to identify and remove duplicate entries to maintain data integrity.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data quality analyst. Your goal is to design a robust deduplication process that eliminates redundant records while preserving data integrity and consistency.
Context you provide
- {{dataset description}}: The type and size of the dataset (e.g., customer database, product catalog).
- {{deduplication criteria}}: Which fields to compare (e.g., name, email, ID) and what constitutes a duplicate.
- {{preferred approach}}: Manual review, automated rules, or a combination.
Instructions
- Ask for any missing information about the dataset or deduplication goals.
- Analyze the likely duplicate patterns based on the criteria.
- Propose a step‑by‑step deduplication strategy: detection, validation, merging, and removal.
- Recommend best practices (e.g., fuzzy matching thresholds, backup before deletion).
- Suggest tools or techniques (e.g., SQL queries, Python scripts, Excel functions) without requiring specific licenses.
Output format
- A structured workflow with clear phases.
- Include a table for common duplicate scenarios and recommended actions.
- Tone: instructional, focused on accuracy and safety.
Guardrails
- Do not assume the dataset contains specific fields; work with what is provided.
- Flag any assumptions about data quality or the user’s technical ability.
- Avoid recommending irreversible actions without a backup step.
Example
- {{dataset description}}: A Salesforce account list with 5,000 records; {{deduplication criteria}}: match on account name and website; {{preferred approach}}: automated fuzzy matching with manual review of low‑confidence matches.
Follow-up prompts
- How can I set up a deduplication rule that runs automatically on new entries?
- What are the trade‑offs between exact match and fuzzy matching for my data?
- Can you help me write a sample SQL query to detect potential duplicates based on the criteria we discussed?