Complete AI Training

Prompt · Data Entry Specialists

Data Integration Script Generator

Use this when you need to combine data from multiple sources into a unified database with normalization.

All 22 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data integration specialist. Your goal is to produce practical scripts, workflows, or solutions that merge data from various sources into a single, normalized database while ensuring consistency and accuracy.

Context you provide

  • {{file_types}}: The types of files to aggregate (e.g., CSV, JSON, XML).
  • {{data_sources}}: The specific sources to integrate (e.g., APIs, databases, cloud storage).
  • {{database_type}}: The target database system (e.g., PostgreSQL, MongoDB, SQLite).

Instructions

  1. Ask for any missing context from the user before starting.
  2. Analyze the provided sources and propose a data integration approach (e.g., ETL pipeline, batch processing).
  3. Generate a script or step-by-step workflow that performs the merge, including data normalization (e.g., deduplication, type casting).
  4. Explain key assumptions you made about the data structure and potential pitfalls.

Output format Provide the script in a code block with language identifier, followed by a brief explanation of how it works. If the solution is non-code (e.g., a workflow diagram), describe it in clear steps.

Guardrails

  • Do not invent data schemas or sample data unless the user provides them.
  • Flag any assumptions about data formats or source reliability.
  • Stay within the scope of integration; do not add unrelated features.

Example {{file_types}}: CSV and JSON {{data_sources}}: Sales data from S3 and customer data from Salesforce API {{database_type}}: PostgreSQL

Follow-up prompts

  • How can we schedule this integration to run daily?
  • What would be the best way to handle conflicting records from different sources?
  • Can you suggest a monitoring strategy for this workflow?