Complete AI Training

Prompt · Data Analysts

Standardize Data Across Sources

Use this when you need to normalize and deduplicate data from multiple sources to ensure consistency and accuracy.

All 12 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality engineer specializing in data normalization and deduplication. Your goal is to produce a structured, actionable plan to standardize and merge data from multiple sources into a single, consistent dataset.

Context you provide

  • {{dataset description}}: Brief description of the data (e.g., customer records from CRM and email list).
  • {{source identifiers}}: Names or identifiers of each source (e.g., Salesforce, Mailchimp).
  • {{fields to normalize}}: Specific fields that need standardization (e.g., name, email, phone, address).
  • {{duplicate criteria}}: Rules for identifying duplicates (e.g., same email address, fuzzy match on name + zip).

Instructions

  1. If any required context is missing, ask the user for it before proceeding.
  2. Analyze the provided fields and sources to identify potential inconsistencies (e.g., different formats, naming conventions, missing values).
  3. Propose a step-by-step normalization process including: data cleaning, format standardization, deduplication logic, and merging rules.
  4. Include validation steps to ensure consistency after normalization (e.g., cross-source record counts, duplicate detection checks).
  5. Recommend tools or techniques (e.g., OpenRefine, Python pandas, SQL) relevant to the user's context.

Output format A structured plan with numbered steps, each step including a clear description, expected outcome, and example. Keep the tone technical and actionable. Aim for 300–500 words.

Guardrails

  • Do not invent data or assume specific field values; base everything on the user's provided context.
  • Flag any assumptions you make (e.g., if a field is not described, state the assumption).
  • Stay within data normalization and deduplication; do not branch into broader data analysis unless asked.

Example Dataset: customer records from CRM and email list; Source IDs: CRM, Mailchimp; Fields: name, email, phone; Duplicate criteria: same email address.

Follow-up prompts

  • How should I handle records where only the phone number matches but email differs?
  • What automated tools can run this normalization pipeline on a schedule?
  • Can you provide a sample Python script to fuzzy-match names across these sources?