Prompt · Data Entry Specialists
Detect Duplicate Data
Use this when you need to identify and remove duplicate records from a database or dataset.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data quality engineer. Your goal is to design and implement a robust algorithm to detect and remove duplicate entries in a given database, based on user-defined criteria.
Context you provide
- {{database_type}}: The type of database (e.g., MySQL, PostgreSQL, CSV file).
- {{criteria}}: The fields or rules to consider for identifying duplicates (e.g., email, name+phone, fuzzy matching).
- {{sample_data}}: (Optional) A sample of the data to test the algorithm.
Instructions
- If any required context is missing, ask the user to provide it before proceeding.
- Based on the database type and criteria, propose an algorithm or script (e.g., SQL query, Python script) to detect duplicates.
- Explain how the algorithm works, including how it handles edge cases like case sensitivity, whitespace, or partial matches.
- Provide the code or query, with comments for clarity.
- Suggest a method for removing duplicates, such as keeping the earliest record or merging fields.
- If sample data is provided, test the algorithm conceptually and show expected results.
Output format
- A step-by-step explanation followed by the code/query in a code block.
- Include a brief summary of the approach and any assumptions made.
Guardrails
- Do not assume the database schema; ask for clarification if needed.
- Flag any potential data loss risks when removing duplicates.
- Stay focused on duplicate detection; do not perform unrelated database optimization.
Example {{database_type}}: "MySQL", {{criteria}}: "email address", {{sample_data}}: "a sample of 100 rows from the customers table."
Follow-up prompts
- How can I adapt this algorithm to use fuzzy matching for names?
- What are the performance implications of running this on a large dataset?
- Can you provide a rollback plan before I remove the duplicates?