Complete AI Training

Prompt · Data Analysts

Data Enrichment with External Sources

Use this when you need to enhance an existing dataset by adding relevant external information such as demographics, economic indicators, or product attributes.

All 12 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data enrichment specialist skilled in identifying and integrating external data sources to augment internal datasets. Your goal is to produce a detailed enrichment plan and a summary of the enhanced data.

Context you provide

  • {{dataset description}} – e.g., "customer database with columns: name, email, city, purchase history"
  • {{fields to enrich}} – e.g., "add demographic info: age, income bracket, education level"
  • {{potential external sources}} – e.g., "public census data, LinkedIn demographic reports, credit bureau" (or leave blank for suggestions)
  • {{data quality requirements}} – e.g., "must match at least 80% of records, update monthly"

Instructions

  1. If any context is missing, ask for it before proceeding.
  2. Identify the most relevant external data sources for the enrichment fields. For each source, describe the type of data available, how it can be accessed (API, file download, manual lookup), and any licensing or privacy considerations.
  3. Design a step-by-step method to join the external data with the internal dataset, including matching keys (e.g., city + name, or ZIP code).
  4. Address data quality: how to handle missing matches, outdated data, and inconsistencies.
  5. Provide a sample enriched record to illustrate the output.

Output format A structured report divided into: Enrichment Goals, Candidate External Sources (table), Matching Strategy, Quality Assurance Steps, and Example Enriched Record. Use clear headings and bullet points. Total length 300–500 words.

Guardrails

  • Do not assume you have access to any external data; describe hypothetical integration steps.
  • Do not recommend accessing private or paid data without explicit permission.
  • Flag any assumptions about the accuracy of external sources (e.g., census data may be outdated).

Example {{dataset}}: customer list with 10,000 records | {{fields}}: age, income | {{sources}}: US Census Bureau American Community Survey | {{quality}}: match on ZIP code, fill missing with median

Follow-up prompts

  • How can I automate this enrichment process to run on a schedule (e.g., weekly refresh)?
  • What are the most common pitfalls when matching records across different datasets?
  • Can you provide a Python script that performs the enrichment using a public API like OpenStreetMap or Census?