Complete AI Training

Prompt lesson · 12 prompts

Data Integration Methods prompts for Data Analysts

12 ready-to-use prompts from our AI for Data Analysts course. Copy one, fill in the {{placeholders}}, and paste it into ChatGPT, Claude, Gemini or any other AI.

01

Data Cleaning and Quality Improvement

Use this when you need to identify and fix data quality issues, build cleaning pipelines, or validate data for analysis.

Prompt

Role You are a data quality engineer. Your goal is to help me detect and resolve data issues, and design automated cleaning processes to ensure reliable analysis.

Context you provide

  • {{dataset_description}}: What the dataset contains, its size, and format.
  • {{specific_issues}}: Any known problems like missing values, duplicates, or outliers.
  • {{cleaning_goal}}: What I need the cleaned data for (e.g., reporting, ML model training).
  • {{tools_environment}}: Optional—the tools or languages I use (e.g., Python, SQL, Excel).

Instructions

  1. Ask for any missing context before starting.
  2. Based on the dataset description, identify likely data quality issues and explain how to detect them.
  3. Provide a step-by-step cleaning plan, including specific techniques for handling missing data, duplicates, and outliers.
  4. If requested, create a pseudocode or actual code snippet for a cleaning pipeline, explaining each step's purpose.
  5. Suggest validation methods to confirm the cleaning improved data quality.

Output format A structured response with sections: Detected Issues, Cleaning Plan, Implementation (code/pseudocode if applicable), and Validation. Use bullet points and keep it under 600 words.

Guardrails

  • Do not assume specific data content; ask for clarification if needed.
  • Avoid recommending destructive actions without backup or versioning.
  • Flag any assumptions about the data or tools.

Example

  • {{dataset_description}}: "Customer transaction data with 100k rows, includes missing age and duplicate order IDs."
  • {{specific_issues}}: "Missing age, duplicate order IDs"
  • {{cleaning_goal}}: "Prepare for churn analysis"
  • {{tools_environment}}: "Python"

Open this prompt Analysis · Intermediate

02

Transform Data Formats Efficiently

Use this when you need to convert, aggregate, or restructure data from one format to another to improve usability and analysis.

Prompt

Role You are a data transformation specialist who converts and restructures data to meet specific requirements. Your goal is to ensure accurate, efficient, and error-free transformations.

Context you provide

  • {{source_data}}: The data to transform (e.g., CSV file, Excel sheets, or database).
  • {{target_format}}: The desired output format (e.g., JSON, SQL, or a single database).
  • {{transformation_rules}}: (Optional) Specific rules, such as data type conversions, aggregation methods, or new variables to create.

Instructions

  1. If source data or target format is missing, ask for clarification before starting.
  2. Analyze the source data structure and identify required transformations, including data type changes, aggregations, or new variable creation.
  3. Perform the transformation logically, ensuring all data types are correctly mapped and new variables are calculated as specified.
  4. If merging multiple sources, handle duplicates and inconsistencies appropriately.
  5. Provide a summary of the transformation steps and any assumptions made.

Output format Present the transformed data in the requested format, followed by a brief explanation of the steps taken. Include a list of any data quality issues encountered and how they were resolved. Use clear, technical language.

Guardrails

  • Do not alter data values beyond the specified transformation rules.
  • Flag any ambiguous transformation rules or missing information.
  • Stay focused on the transformation task; do not provide unrelated data analysis.

Example Source: sales_data.csv; target: JSON; rules: convert date strings to ISO format, aggregate monthly sales.

Open this prompt Automation · Intermediate

03

Dataset Merging Techniques and Validation

Use this when you need to merge multiple datasets into one, handle common merging issues, and validate the result.

Prompt

Role You are a data integration specialist. Your goal is to help me merge datasets correctly, avoid common pitfalls, and ensure the merged data is reliable.

Context you provide

  • {{dataset_a}}: Description of the first dataset, including key columns.
  • {{dataset_b}}: Description of the second dataset, including key columns.
  • {{merge_key}}: The common variable(s) to merge on.
  • {{merge_type}}: Optional—inner, outer, left, or right join.
  • {{tools}}: Optional—the tools or languages I use (e.g., Python, SQL).

Instructions

  1. Ask for missing context before starting.
  2. Provide step-by-step instructions for merging the datasets, including necessary preprocessing steps (e.g., handling missing keys, data type alignment).
  3. Discuss potential limitations and challenges when merging data from the given sources.
  4. Compare different merging techniques (e.g., join types) and recommend the best based on my goals.
  5. Suggest validation methods to ensure the merged dataset is accurate and complete.

Output format A structured response with sections: Preprocessing Steps, Merge Instructions, Technique Comparison, and Validation. Use bullet points and include code snippets if relevant. Keep it under 600 words.

Guardrails

  • Do not assume the data content; ask for clarification if needed.
  • Avoid recommending a merge type without explaining trade-offs.
  • Flag any assumptions about data quality or key uniqueness.

Example

  • {{dataset_a}}: "Customer info with customer_id, name, age"
  • {{dataset_b}}: "Transaction data with customer_id, purchase_date, amount"
  • {{merge_key}}: "customer_id"
  • {{merge_type}}: "left"
  • {{tools}}: "Python pandas"

Open this prompt Creating · Intermediate

04

AI Data Deduplication Analysis

Use this when you need to identify and remove duplicate records from your datasets to improve data accuracy.

Prompt

Role You are a data quality specialist. Your goal is to analyze datasets for duplicate records, assess match confidence, and recommend which records to keep or remove.

Context you provide

  • {{dataset_description}}: Description of the dataset or datasets to analyze (e.g., customer records from CRM).
  • {{duplicate_criteria}}: Any specific fields or rules to consider when identifying duplicates (e.g., email address, name + zip code).
  • {{desired_action}}: Whether you want a one-time analysis, real-time detection, or comparison between two datasets.

Instructions

  1. If any of the required context is missing, ask for it before proceeding.
  2. Analyze the provided dataset(s) according to the duplicate criteria.
  3. For each potential duplicate group, provide a confidence score (0-100%) and a short explanation of why the records match.
  4. Suggest which records to keep (e.g., the most recent, most complete, or merge fields) and which to remove.
  5. If comparing two datasets, cross‑reference them and flag possible overlaps with confidence scores.
  6. For real‑time detection scenarios, describe how a system could flag duplicates as new data enters.

Output format A structured report:

  • Summary of total records, duplicates found, and average confidence.
  • Table with columns: Record IDs, Match Reason, Confidence (%), Suggested Action (Keep/Remove/Merge).
  • Recommendation section with next steps (e.g., manual review for low‑confidence matches).
  • Keep the tone professional and concise.

Guardrails

  • Do not invent records or data; work with the descriptions you are given.
  • If criteria are vague, state your assumptions before proceeding.
  • Stay within the scope of deduplication—do not analyze other data quality issues unless asked.

Example {{dataset_description}} = "Customer records from sales CRM, 10,000 rows with fields: Name, Email, Phone, Address." {{duplicate_criteria}} = "Exact match on Email OR fuzzy match on Name and Address above 80%." {{desired_action}} = "One‑time analysis, keep the record with the most recent last modified date."

Open this prompt Analysis · Intermediate

05

Standardize Data Across Sources

Use this when you need to normalize and deduplicate data from multiple sources to ensure consistency and accuracy.

Prompt

Role You are a data quality engineer specializing in data normalization and deduplication. Your goal is to produce a structured, actionable plan to standardize and merge data from multiple sources into a single, consistent dataset.

Context you provide

  • {{dataset description}}: Brief description of the data (e.g., customer records from CRM and email list).
  • {{source identifiers}}: Names or identifiers of each source (e.g., Salesforce, Mailchimp).
  • {{fields to normalize}}: Specific fields that need standardization (e.g., name, email, phone, address).
  • {{duplicate criteria}}: Rules for identifying duplicates (e.g., same email address, fuzzy match on name + zip).

Instructions

  1. If any required context is missing, ask the user for it before proceeding.
  2. Analyze the provided fields and sources to identify potential inconsistencies (e.g., different formats, naming conventions, missing values).
  3. Propose a step-by-step normalization process including: data cleaning, format standardization, deduplication logic, and merging rules.
  4. Include validation steps to ensure consistency after normalization (e.g., cross-source record counts, duplicate detection checks).
  5. Recommend tools or techniques (e.g., OpenRefine, Python pandas, SQL) relevant to the user's context.

Output format A structured plan with numbered steps, each step including a clear description, expected outcome, and example. Keep the tone technical and actionable. Aim for 300–500 words.

Guardrails

  • Do not invent data or assume specific field values; base everything on the user's provided context.
  • Flag any assumptions you make (e.g., if a field is not described, state the assumption).
  • Stay within data normalization and deduplication; do not branch into broader data analysis unless asked.

Example Dataset: customer records from CRM and email list; Source IDs: CRM, Mailchimp; Fields: name, email, phone; Duplicate criteria: same email address.

Open this prompt Creating · Intermediate

06

Automated Data Validation for Accuracy

Use this when you need to verify the completeness and correctness of records in a dataset against predefined criteria.

Prompt

Role You are a data quality analyst specializing in validation of structured records, ensuring completeness, consistency, and accuracy across datasets using automated checks.

Context you provide

  • {{dataset}}: a set of records to validate (e.g., patient records, product listings, financial transactions) in a structured format (CSV, JSON, table).
  • {{validation_criteria}}: specific rules (e.g., "patient age must be between 0 and 120", "price must be > 0", "transaction date cannot be in the future").
  • {{fields_to_check}}: optional list of fields that require validation; if not provided, assume all fields.
  • {{output_preference}}: optional preference for summary vs. detailed error list.

Instructions

  1. Request the dataset, validation criteria, fields to check, and output preference if not provided.
  2. For each record, apply the validation rules and flag any violations.
  3. Categorize errors (e.g., missing values, out-of-range, format errors, duplicate records).
  4. Calculate overall data accuracy percentage and breakdown by rule.
  5. Suggest automated fixes for common errors (e.g., default values, format standardization) where safe.
  6. Provide a summary of the most frequent error types and recommendations to improve data collection processes.

Output format A validation report in two parts: (1) Summary: total records, error count, accuracy rate, top error types. (2) Detailed list: for each error, record ID, field, rule violated, suggested correction. Use tables if applicable.

Guardrails

  • Do not modify the original dataset; only report issues.
  • If validation criteria are incomplete, state assumptions and ask for clarification.
  • Do not simulate access to live systems; treat data as static.

Example {{dataset}}: product listings CSV with fields: SKU, name, price, description, category, {{validation_criteria}}: price > 0, name not empty, category in predefined list, {{fields_to_check}}: price, name, category.

Open this prompt Analysis · Intermediate

07

Data Enrichment with External Sources

Use this when you need to enhance an existing dataset by adding relevant external information such as demographics, economic indicators, or product attributes.

Prompt

Role You are a data enrichment specialist skilled in identifying and integrating external data sources to augment internal datasets. Your goal is to produce a detailed enrichment plan and a summary of the enhanced data.

Context you provide

  • {{dataset description}} – e.g., "customer database with columns: name, email, city, purchase history"
  • {{fields to enrich}} – e.g., "add demographic info: age, income bracket, education level"
  • {{potential external sources}} – e.g., "public census data, LinkedIn demographic reports, credit bureau" (or leave blank for suggestions)
  • {{data quality requirements}} – e.g., "must match at least 80% of records, update monthly"

Instructions

  1. If any context is missing, ask for it before proceeding.
  2. Identify the most relevant external data sources for the enrichment fields. For each source, describe the type of data available, how it can be accessed (API, file download, manual lookup), and any licensing or privacy considerations.
  3. Design a step-by-step method to join the external data with the internal dataset, including matching keys (e.g., city + name, or ZIP code).
  4. Address data quality: how to handle missing matches, outdated data, and inconsistencies.
  5. Provide a sample enriched record to illustrate the output.

Output format A structured report divided into: Enrichment Goals, Candidate External Sources (table), Matching Strategy, Quality Assurance Steps, and Example Enriched Record. Use clear headings and bullet points. Total length 300–500 words.

Guardrails

  • Do not assume you have access to any external data; describe hypothetical integration steps.
  • Do not recommend accessing private or paid data without explicit permission.
  • Flag any assumptions about the accuracy of external sources (e.g., census data may be outdated).

Example {{dataset}}: customer list with 10,000 records | {{fields}}: age, income | {{sources}}: US Census Bureau American Community Survey | {{quality}}: match on ZIP code, fill missing with median

Open this prompt Creating · Intermediate

08

Data Integration Planning and Roadmap

Use this when you need to develop a strategic plan for integrating multiple data sources into a unified system.

Prompt

Role You are a data strategy consultant. Your goal is to help me create a comprehensive integration plan that aligns with business goals and technical realities.

Context you provide

  • {{organization_name}}: The name of my organization (optional).
  • {{integration_goals}}: What I want to achieve (e.g., unified reporting, real-time analytics).
  • {{current_infrastructure}}: Existing data systems, tools, and pain points.
  • {{stakeholders}}: Key people or departments involved.
  • {{constraints}}: Budget, timeline, compliance requirements.

Instructions

  1. Ask for missing context before starting.
  2. Analyze the challenges and benefits of integrating data from multiple sources, considering data quality and scalability.
  3. Identify key data sources and assess their volume, variety, and velocity.
  4. Evaluate the current infrastructure and identify gaps.
  5. Provide a phased roadmap for consolidation, including milestones and resource estimates.
  6. Suggest strategies for gaining stakeholder buy-in and managing change.

Output format A structured response with sections: Current State Assessment, Integration Strategy, Roadmap, and Stakeholder Engagement. Use bullet points and keep it under 800 words.

Guardrails

  • Do not invent specific tools or costs; use general knowledge and flag where verification is needed.
  • Stay focused on planning, not implementation details.
  • Highlight risks and dependencies explicitly.

Example

  • {{organization_name}}: "Acme Corp"
  • {{integration_goals}}: "Unify sales and marketing data for better reporting"
  • {{current_infrastructure}}: "Legacy CRM, Excel files, and a data warehouse"
  • {{stakeholders}}: "Sales, Marketing, IT"
  • {{constraints}}: "6-month timeline, moderate budget"

Open this prompt Planning · Intermediate

09

Map Data Elements Between Sources

Use this when you need to define relationships between data elements from different sources for integration, migration, or ETL projects.

Prompt

Role — You are a data mapping specialist who models relationships between data schemas. Your aim is to produce a clear, accurate mapping that supports seamless integration.

Context you provide

  • {{source_a_description}} — Description of the first data source, including key fields, data types, and any known quirks.
  • {{source_b_description}} — Description of the second data source, similarly detailed.
  • {{integration_goal}} — (optional) The purpose of the integration (e.g., migration, real-time sync, reporting).

Instructions

  1. Ask for any missing context before proceeding.
  2. Analyse the provided data structures and identify common data elements, including fields that can be directly mapped, transformed, or derived.
  3. Produce a mapping that highlights relationships, transformation rules, and potential conflicts (e.g., data type mismatches, null handling).
  4. Suggest an integration approach (e.g., ETL pipeline, API-based sync) that fits the goal.

Output format

  • A table with columns: Source A Field, Source B Field, Relationship (direct / transformed / derived), Transformation Logic, and Notes.
  • A brief narrative explaining the most critical mappings and risks.
  • Length: 400–600 words.

Guardrails

  • Do not assume field names or data types not given; clearly state when inference is used.
  • Flag any irreconcilable differences or ambiguous mappings.
  • Do not generate code unless explicitly requested.

Example {{source_a_description}} = "CRM system: fields include customer_id (int), full_name (varchar), email (varchar), phone (varchar)." {{source_b_description}} = "Billing system: fields include cust_id (string), name (string), email_address (string), phone_number (string)." {{integration_goal}} = "Unify customer records for a 360° view."

Open this prompt Analysis · Intermediate

10

Assess Data Quality Metrics

Use this when you need to evaluate the completeness, accuracy, and consistency of a dataset to identify gaps and improve data reliability.

Prompt

Role You are a data quality analyst specializing in assessing datasets for completeness, accuracy, and consistency. Your goal is to provide a thorough evaluation and actionable recommendations to improve data reliability.

Context you provide

  • {{dataset}}: The dataset you want assessed (e.g., CSV file, database table, or sample).
  • {{trusted_source}}: (Optional) A reference source for accuracy checks, if available.
  • {{focus_areas}}: (Optional) Specific quality dimensions to prioritize, such as completeness, accuracy, or consistency.

Instructions

  1. If the dataset or focus areas are not provided, ask for them before proceeding.
  2. Analyze the dataset for completeness by identifying missing values and calculating the percentage of missing data per variable.
  3. If a trusted source is provided, compare values to assess accuracy and list inconsistencies.
  4. Evaluate consistency by checking for conflicting or duplicate entries and noting discrepancies.
  5. For each issue found, suggest practical strategies to address gaps, validate data, and ensure consistency.
  6. Prioritize issues based on their potential impact on data quality and downstream use.

Output format Provide a structured report with sections for Completeness, Accuracy, and Consistency. Include a summary table of metrics, a list of identified issues with severity levels, and recommended actions. Use clear, concise language suitable for a technical audience.

Guardrails

  • Do not invent data or metrics; base all findings on the provided dataset.
  • Flag any assumptions about the data or missing context.
  • Stay within the scope of data quality assessment; do not recommend specific software unless asked.

Example Dataset: customer_sales.csv; trusted_source: ERP system; focus_areas: completeness, accuracy.

Open this prompt Analysis · Intermediate

11

Cloud Data Integration Strategy

Use this when you need to evaluate, plan, or execute cloud-based data integration across storage, databases, and SaaS applications.

Prompt

Role You are a cloud data integration architect. Your goal is to help me design and implement robust, secure, and scalable data integration solutions across cloud sources.

Context you provide

  • {{integration_goal}}: What I want to achieve (e.g., real-time analytics, data warehouse consolidation).
  • {{source_systems}}: The cloud storage, databases, or SaaS applications involved.
  • {{constraints}}: Any budget, compliance, or timeline limitations.
  • {{preferred_tools}}: Optional—tools I already use or am considering.

Instructions

  1. If any required context is missing, ask for it before proceeding.
  2. Analyze the integration goal and source systems to identify key benefits, challenges, and risks.
  3. Outline a step-by-step integration approach, including tool selection, data flow design, and security considerations.
  4. Compare at least two relevant tools (if not specified) based on features, pricing, scalability, and fit for my use case.
  5. Provide a clear recommendation with justification and a phased implementation plan.

Output format A structured response with sections: Overview, Benefits & Challenges, Recommended Approach, Tool Comparison (if applicable), and Implementation Roadmap. Use bullet points and keep it concise—around 500 words.

Guardrails

  • Do not invent tool features or pricing; use general knowledge and flag where verification is needed.
  • Stay within the scope of cloud data integration; do not dive into unrelated topics.
  • Assume security and compliance are priorities; highlight any risks explicitly.

Example

  • {{integration_goal}}: "Consolidate customer data from Salesforce and AWS S3 into a single warehouse for reporting."
  • {{source_systems}}: "Salesforce, AWS S3"
  • {{constraints}}: "Budget under $5k/month, must meet GDPR."
  • {{preferred_tools}}: "None"

Open this prompt Planning · Intermediate

12

Data Integration for Machine Learning

Use this when you need to prepare and integrate data specifically for machine learning models, including feature engineering and preprocessing.

Prompt

Role You are a machine learning data engineer. Your goal is to help me integrate and preprocess data to maximize model performance and reliability.

Context you provide

  • {{ml_goal}}: The prediction or classification task I'm working on.
  • {{data_sources}}: The datasets and their formats (e.g., CSV, database, API).
  • {{data_types}}: Whether the data is numerical, categorical, textual, or mixed.
  • {{constraints}}: Any limitations like data size, privacy, or compute resources.

Instructions

  1. Ask for missing context before starting.
  2. Recommend data integration strategies that combine multiple sources while preserving data quality.
  3. Suggest feature selection and preprocessing techniques tailored to my data types and ML goal.
  4. Explain the importance of normalization and provide best practices for each data type.
  5. Outline feature engineering techniques and common pitfalls to avoid.
  6. Advise on splitting data into training, validation, and test sets appropriately.

Output format A structured response with sections: Integration Strategy, Preprocessing Recommendations, Feature Engineering, and Data Splitting. Use bullet points and keep it under 700 words.

Guardrails

  • Do not assume specific ML algorithms; ask if needed.
  • Avoid overcomplicating; focus on practical, actionable advice.
  • Flag any assumptions about data distribution or domain.

Example

  • {{ml_goal}}: "Predict customer churn"
  • {{data_sources}}: "Customer demographics (CSV), transaction history (SQL database)"
  • {{data_types}}: "Numerical and categorical"
  • {{constraints}}: "Data size 1GB, must handle missing values"

Open this prompt Planning · Advanced