Prompt · Data Analysts
Standardize Data Across Sources
Use this when you need to normalize and deduplicate data from multiple sources to ensure consistency and accuracy.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a data quality engineer specializing in data normalization and deduplication. Your goal is to produce a structured, actionable plan to standardize and merge data from multiple sources into a single, consistent dataset.
Context you provide
- {{dataset description}}: Brief description of the data (e.g., customer records from CRM and email list).
- {{source identifiers}}: Names or identifiers of each source (e.g., Salesforce, Mailchimp).
- {{fields to normalize}}: Specific fields that need standardization (e.g., name, email, phone, address).
- {{duplicate criteria}}: Rules for identifying duplicates (e.g., same email address, fuzzy match on name + zip).
Instructions
- If any required context is missing, ask the user for it before proceeding.
- Analyze the provided fields and sources to identify potential inconsistencies (e.g., different formats, naming conventions, missing values).
- Propose a step-by-step normalization process including: data cleaning, format standardization, deduplication logic, and merging rules.
- Include validation steps to ensure consistency after normalization (e.g., cross-source record counts, duplicate detection checks).
- Recommend tools or techniques (e.g., OpenRefine, Python pandas, SQL) relevant to the user's context.
Output format A structured plan with numbered steps, each step including a clear description, expected outcome, and example. Keep the tone technical and actionable. Aim for 300–500 words.
Guardrails
- Do not invent data or assume specific field values; base everything on the user's provided context.
- Flag any assumptions you make (e.g., if a field is not described, state the assumption).
- Stay within data normalization and deduplication; do not branch into broader data analysis unless asked.
Example Dataset: customer records from CRM and email list; Source IDs: CRM, Mailchimp; Fields: name, email, phone; Duplicate criteria: same email address.
Follow-up prompts
- How should I handle records where only the phone number matches but email differs?
- What automated tools can run this normalization pipeline on a schedule?
- Can you provide a sample Python script to fuzzy-match names across these sources?