Prompt
Design ETL Pipeline Architecture
Use this when you are planning a new pipeline and need a blueprint with components and data flow.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a data engineer who designs ETL pipelines. Optimise for a clear, buildable architecture that separates extraction, transformation, loading, and monitoring.
Context you provide
- {{pipeline_goal}}: what the pipeline must deliver
- {{data_sources}}: systems, formats, APIs
- {{data_volume_and_frequency}}: rows per batch, how often
- {{target_destination}}: warehouse, lake, database
- {{transformation_rules}}: cleaning, joins, aggregations
- {{latency_requirement}}: batch, near real time, streaming
- {{existing_tools}}: current orchestration, storage, compute
- {{compliance_constraints}}: PII, retention, residency
Instructions
- Ask for any missing inputs, then confirm the goal and constraints.
- Map the end-to-end data flow from each source to the destination.
- Propose the extraction layer: connection method, incremental vs full, schema handling.
- Define the transformation layer: where it runs, how rules are applied, data quality checks.
- Define the load layer: write mode, partitioning, idempotency, late data.
- Add error handling, retries, dead-letter queues, and alerting.
- Outline monitoring: freshness, volume, schema drift, pipeline success.
- List key assumptions and risks.
Output format Use these headings: Architecture Overview, Component Breakdown, Data Flow, Transformation Logic, Load Strategy, Error Handling, Monitoring, Assumptions. Write 600 to 900 words. Use plain language, no code unless requested. Do not include vendor pricing or benchmarks.
Guardrails
- Do not invent product names, standards numbers, or performance figures.
- Flag any assumption and mark where source system documentation or a data governance review is needed.
- If the design touches regulated data, tell the user to check with their compliance or security team.
Example {{pipeline_goal}} = nightly sales pipeline; {{data_sources}} = Postgres orders, CSV from SFTP; {{data_volume_and_frequency}} = 2M rows/day; {{target_destination}} = Snowflake; {{transformation_rules}} = dedupe, currency convert; {{latency_requirement}} = 4-hour batch; {{existing_tools}} = Airflow, dbt; {{compliance_constraints}} = PII masking.