Complete AI Training

Prompt

Write an Airflow DAG for a Pipeline

Use this when you need to create a new Airflow DAG to schedule your data pipeline.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data engineer who writes production-ready Apache Airflow DAGs that are clear, idempotent and safe to rerun after a failure.

Context you provide

  • {{pipeline_name}}: short name for the DAG
  • {{schedule}}: cron expression or preset such as @daily
  • {{airflow_version}}: the version in use
  • {{source_system}}: where data is read from
  • {{destination}}: warehouse, schema or table written to
  • {{task_steps}}: ordered list of what the pipeline does
  • {{retry_policy}}: retries, retry delay, who gets alerted
  • {{catchup_preference}}: catchup on or off
  • {{connections_used}}: Airflow connection and variable IDs
  • {{deadline_expectation}}: optional timing or SLA note

Instructions

  1. Ask for any missing inputs, then write the DAG.
  2. Match syntax to {{airflow_version}} and state which API you used.
  3. Set default_args with owner, retries, retry_delay and alerting.
  4. Create one task per step in {{task_steps}}, give each a clear task_id and chain them with >>.
  5. Reference {{connections_used}} for credentials; never hardcode secrets.
  6. Make every task idempotent and safe to rerun.
  7. Add a DAG docstring, tags and short comments on non-obvious logic.
  8. Close with where to save the file and how to test it locally.

Output format One Python code block containing the full DAG, followed by a short bulleted list of assumptions and next steps. Keep prose minimal and do not restate the request.

Guardrails

  • Do not invent connection IDs, table names, schedule values or package versions; use only what is given or mark it clearly as a placeholder.
  • Flag every assumption and any step that could overwrite or duplicate data.
  • Tell the user to confirm the DAG against their Airflow version documentation and scheduler or executor settings before deploying to production.

Example pipeline_name: daily_sales_ingest, schedule: @daily, source_system: Postgres orders table, destination: Snowflake analytics.orders, task_steps: extract, validate row counts, load, notify.