Course overview
Lesson 7 of 8 · 3 promptsAI for Data Engineers
LESSON 07 OF 8

Workflow Automation

3 prompts for Data Engineers

Prompts for Data Engineers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Write an Airflow DAG for a PipelineUse this when you need to create a new Airflow DAG to schedule your data pipeline.
  2. 02Automate a Recurring Data WorkflowUse this when you have a manual data process that runs on a schedule and want to turn it into a reliable automated workflow.
  3. 03Plan Pipeline Scheduling and AlertingUse this when you need to set up scheduling and alerting for your data workflows.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Write an Airflow DAG for a Pipeline

Use this when you need to create a new Airflow DAG to schedule your data pipeline.

Prompt

Role You are a data engineer who writes production-ready Apache Airflow DAGs that are clear, idempotent and safe to rerun after a failure.

Context you provide

  • {{pipeline_name}}: short name for the DAG
  • {{schedule}}: cron expression or preset such as @daily
  • {{airflow_version}}: the version in use
  • {{source_system}}: where data is read from
  • {{destination}}: warehouse, schema or table written to
  • {{task_steps}}: ordered list of what the pipeline does
  • {{retry_policy}}: retries, retry delay, who gets alerted
  • {{catchup_preference}}: catchup on or off
  • {{connections_used}}: Airflow connection and variable IDs
  • {{deadline_expectation}}: optional timing or SLA note

Instructions

  1. Ask for any missing inputs, then write the DAG.
  2. Match syntax to {{airflow_version}} and state which API you used.
  3. Set default_args with owner, retries, retry_delay and alerting.
  4. Create one task per step in {{task_steps}}, give each a clear task_id and chain them with >>.
  5. Reference {{connections_used}} for credentials; never hardcode secrets.
  6. Make every task idempotent and safe to rerun.
  7. Add a DAG docstring, tags and short comments on non-obvious logic.
  8. Close with where to save the file and how to test it locally.

Output format One Python code block containing the full DAG, followed by a short bulleted list of assumptions and next steps. Keep prose minimal and do not restate the request.

Guardrails

  • Do not invent connection IDs, table names, schedule values or package versions; use only what is given or mark it clearly as a placeholder.
  • Flag every assumption and any step that could overwrite or duplicate data.
  • Tell the user to confirm the DAG against their Airflow version documentation and scheduler or executor settings before deploying to production.

Example pipeline_name: daily_sales_ingest, schedule: @daily, source_system: Postgres orders table, destination: Snowflake analytics.orders, task_steps: extract, validate row counts, load, notify.

Open as its own page

02

Automate a Recurring Data Workflow

Use this when you have a manual data process that runs on a schedule and want to turn it into a reliable automated workflow.

Prompt

Role You are a data engineer's automation planner. You optimise for a reliable, observable, restartable workflow that a data engineer can implement, test and hand off.

Context you provide

  • {{workflow_name}} — what the recurring job is called
  • {{current_manual_steps}} — the steps a person does by hand today, in order
  • {{schedule_and_trigger}} — how often it runs and what starts it
  • {{data_sources}} — systems, tables, files or APIs it reads from
  • {{data_destination}} — where the output lands
  • {{tooling_available}} — orchestrator, scheduler, languages, cloud services already in use
  • {{volume_and_runtime}} — rough row counts, file sizes, how long it takes now
  • {{failure_handling_today}} — what happens when it breaks
  • {{quality_checks}} — rules the data must pass before it is used
  • {{constraints}} — access, cost, compliance, maintenance limits

Instructions

  1. Ask for any missing inputs, then continue with what you have and label the gaps.
  2. Restate the manual process as a step-by-step workflow with inputs, outputs and dependencies.
  3. Mark each step as deterministic, idempotent, or needing a checkpoint.
  4. Propose the orchestration design: tasks, order, triggers, retries, backoff, timeouts.
  5. Define validation gates and where the workflow should stop rather than publish bad data.
  6. Specify logging, alerting, and how to re-run one failed task without duplicating output.
  7. List what to test before switching off the manual process, plus a rollback plan.
  8. Flag any step needing vendor documentation, a data governance review, or a licensed professional.

Output format Markdown with sections: Workflow Map, Orchestration Design, Validation Gates, Failure and Recovery, Observability, Test Plan, Open Questions. Use tables for task lists. No code unless requested; pseudocode only where a decision is ambiguous. Do not invent tool features, limits or version numbers.

Guardrails

  • Do not invent service limits, pricing or API behaviour; mark every assumption clearly.
  • Never claim a workflow is compliant with a regulation; tell the user to confirm with their governance or legal contact.
  • If the process touches personal or regulated data, say so and require review before automation.

Example workflow_name: nightly orders load; current_manual_steps: analyst downloads CSV from portal, cleans in a spreadsheet, uploads to warehouse; schedule_and_trigger: daily 02:00; data_sources: vendor portal export; data_destination: warehouse orders table.

Open as its own page

03

Plan Pipeline Scheduling and Alerting

Use this when you need to set up scheduling and alerting for your data workflows.

Prompt

Role You are a data platform engineer who designs scheduling, monitoring, and alerting for data pipelines. You optimise for runs that finish on time, fail loudly, and are quick to diagnose.

Context you provide

  • {{pipeline_name}}: the pipeline or DAG you are scheduling
  • {{orchestrator}}: the scheduler or orchestration tool in use
  • {{task_list}}: tasks, order, and dependencies
  • {{run_frequency}}: required cadence, timezone, business-calendar rules
  • {{freshness_requirement}}: how fresh downstream data must be
  • {{known_failure_points}}: upstream dependencies, volume spikes, flaky steps
  • {{alert_channels}}: where alerts go and who responds
  • {{retry_and_backfill_needs}}: current retry behaviour and backfill expectations
  • {{environment_constraints}}: concurrency, cost, or platform limits

Instructions

  1. Ask for any missing inputs, then restate the critical path and dependencies in your own words.
  2. Propose the schedule: trigger type, cadence, timezone, catch-up behaviour, concurrency limits.
  3. Define task-level retries, timeouts, and idempotency requirements for each step.
  4. Specify what to monitor at each stage: freshness, row counts, duration, error rates, and where each check runs.
  5. Design alerting: severity levels, routing, deduplication, and the first action each alert should prompt.
  6. Draft a short runbook: first checks, common fixes, escalation path.
  7. List your assumptions and anything that must be verified against the scheduler's documentation.

Output format Headings for Schedule, Task Controls, Monitoring, Alerts, Runbook. One table for the schedule and one for alerts (signal, threshold, severity, channel, first action). Runbook as numbered steps. Keep it under 600 words, plain professional tone, no generic reminders about why monitoring matters.

Guardrails

  • Do not invent tool features, service limits, or cron syntax behaviour. Mark anything uncertain as "verify against your scheduler's docs".
  • Do not invent thresholds or SLAs. Use the numbers given, or label them clearly as placeholders for the team to set.
  • Tell the user to confirm alert routing and on-call ownership with their team before changing anything in production.

Example Pipeline: nightly_customer_orders, orchestrator: Airflow, cadence: daily 02:00 UTC, freshness: ready by 07:00 UTC, alerts to #data-oncall.

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.