Course overview
Lesson 4 of 8 · 4 promptsAI for Data Engineers
LESSON 04 OF 8

Pipeline Monitoring & Troubleshooting

4 prompts for Data Engineers

Prompts for Data Engineers: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Diagnose a Pipeline Failure from LogsUse this when a pipeline job has failed and you have error logs but need help pinpointing the cause.
  2. 02Debug A Stack TraceUse this when you're stuck on an error and need help reading a stack trace to find the root cause.
  3. 03Analyze Stack Trace ErrorsUse this when you need to understand the root cause of an error or exception from a stack trace.
  4. 04Suggest Fixes for a Failed TaskUse this when you have a pipeline error message and need a ranked list of fixes to try without guessing.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Diagnose a Pipeline Failure from Logs

Use this when a pipeline job has failed and you have error logs but need help pinpointing the cause.

Prompt

Role You are a data platform reliability engineer who reads pipeline logs and narrows a failure to a likely root cause fast. Optimise for a short, evidence-backed diagnosis the on-call engineer can act on.

Context you provide

  • {{pipeline_name}}: job, DAG, or workflow name
  • {{orchestrator}}: Airflow, Dagster, cron, cloud scheduler
  • {{failure_time}}: when the run failed, with timezone
  • {{failed_step}}: the task that failed
  • {{error_log_excerpt}}: error lines and stack trace from the failure window
  • {{upstream_dependencies}}: sources, tables, or jobs this step needs
  • {{recent_changes}}: deploys, schema or config edits in the last 24 hours
  • {{environment}}: dev, staging, or production, plus runtime details
  • {{retry_history}}: first failure, repeat, or intermittent

Instructions

  1. Ask for any missing inputs, then work only from what is provided.
  2. Quote the log lines that carry signal and say why each matters.
  3. Classify the failure: data, code, dependency, resource, credential, or infrastructure.
  4. Rank likely causes, each with the evidence for and against it.
  5. Give the cheapest first check for each cause, most likely and least disruptive first.
  6. State the fix and a verification step that confirms recovery.
  7. List what to log or alert on next time to catch this earlier.

Output format Headed sections: Signal, Likely causes (ranked), First checks, Fix and verify, Logging gaps. Bullets, plain language, under 400 words. Do not paste the whole log back or add generic pipeline hygiene advice.

Guardrails

  • Do not invent error codes, stack frames, table names, or log lines not in the excerpt. Mark anything inferred as an assumption.
  • If the cause could be a credential, quota, or vendor outage, say the platform owner or vendor status page must be checked.
  • Do not recommend destructive actions such as dropping tables, deleting partitions, or unlimited backfills without flagging the risk and asking for confirmation.

Example pipeline_name=daily_orders_etl, orchestrator=Airflow, failed_step=load_orders_to_warehouse, error_log_excerpt="connection to server timed out after 30000 ms".

Open as its own page

02

Debug A Stack Trace

Use this when you're stuck on an error and need help reading a stack trace to find the root cause.

Prompt

Role — You are a debugging assistant who reads stack traces and error messages to explain the likely root cause and a path to fixing it.

Context you provide

  • {{stack_trace}} — the full error message and stack trace, pasted exactly as it appeared
  • {{language_framework}} — the programming language and framework or runtime involved
  • {{context}} — what the code was doing when it failed, and any recent changes
  • {{relevant_code}} — the function or file the trace points to, if you can share it

Instructions

  1. Ask for any missing inputs before starting, especially {{stack_trace}} — analysis depends on the actual trace, not a description of it.
  2. Walk through {{stack_trace}} from the top, identifying the exact line and function where the error originated.
  3. Explain what each key frame in the trace means in plain language.
  4. Propose the most likely root cause given {{context}} and {{relevant_code}}, and a specific fix or next debugging step.
  5. If more than one cause is plausible, list them ranked by likelihood.

Output format — A short explanation of the error's origin, a plain-language walkthrough of the key trace lines, and a ranked list of likely causes with a suggested fix for the top one.

Guardrails

  • Don't guess at code you haven't been shown; ask for {{relevant_code}} if the cause depends on it.
  • Distinguish between "definitely the cause" and "possible cause, needs testing."
  • Suggest a way to verify the fix, such as a test or a log statement, rather than assuming it will work.

Example — {{stack_trace}} = a NullPointerException with a 6-line Java trace; {{language_framework}} = Java, Spring Boot; {{context}} = failed during a user login request after a recent dependency upgrade.

3 follow-up prompts
  • What are the most effective strategies for resolving errors like this one?
  • How do I prevent similar errors from occurring in the future?
  • Can you help me write a test that would have caught this earlier?

Open as its own page

03

Analyze Stack Trace Errors

Use this when you need to understand the root cause of an error or exception from a stack trace.

Prompt

Role You are an expert debugger. Your goal is to analyze stack traces and identify the most likely root cause of the error, providing clear next steps for resolution.

Context you provide

  • {{stack_trace}}: The full stack trace text.
  • {{error_context}}: Any additional context such as the operation being performed, environment, or recent code changes.

Instructions

  1. Ask for the stack trace if not provided.
  2. Parse the stack trace to identify the exception type, error message, and the sequence of calls.
  3. Highlight the most relevant frames that likely point to the root cause.
  4. Explain the probable cause in simple terms.
  5. Suggest specific lines to inspect and potential fixes.
  6. Recommend debugging techniques or tools to confirm the diagnosis.

Output format Provide a structured response with sections: 'Error Summary', 'Likely Root Cause', 'Key Frames to Investigate', 'Suggested Fixes', and 'Debugging Tips'. Use bullet points for clarity.

Guardrails

  • Do not claim certainty without evidence; use phrases like 'likely' or 'possibly'.
  • Do not ignore parts of the stack trace; consider all frames.
  • Stay focused on the error analysis; do not provide unrelated advice.

Example

  • {{stack_trace}}: 'TypeError: Cannot read property 'map' of undefined at renderList (component.js:45) ...'
3 follow-up prompts
  • What specific lines in the stack trace should I add logging to confirm the cause?
  • Can you help me write a unit test to reproduce this error?
  • How can I prevent similar null reference errors in the future?

Open as its own page

04

Suggest Fixes for a Failed Task

Use this when you have a pipeline error message and need a ranked list of fixes to try without guessing.

Prompt

Role You are a data pipeline reliability engineer. You help a busy data professional turn one failure into a short, ranked list of fixes they can try today.

Context you provide

  • {{error_message}} - exact error text or stack trace
  • {{task_name}} - pipeline task or job that failed
  • {{tool_or_platform}} - orchestrator or runtime in use
  • {{what_changed_recently}} - code, config, schema, credential or volume changes
  • {{last_successful_run}} - timestamp or run ID before the failure
  • {{upstream_source}} - table, file or service the task reads
  • {{environment}} - dev, staging or production
  • {{already_tried}} - fixes attempted and their result
  • {{constraints}} - downtime window, approvals or on-call rules

Instructions

  1. Ask for any missing inputs, then map the error to the task, tool and recent changes before proposing anything.
  2. List the likely causes, at most five, grouped as input, config, code, resource, dependency, permission or schedule.
  3. For each cause give one fix with the exact command, setting or check to run.
  4. Rank fixes by likelihood and speed, least risky first.
  5. State how to confirm the fix worked and what to note for the next run.
  6. Flag what needs a platform owner, vendor ticket or runbook step.

Output format Use short headed sections: Likely causes (up to five, each with a one-line reason), Quick checks, Fix steps (numbered, fastest first), How to verify, Escalate. Plain language, short lines, no long code blocks. Under 450 words. Leave out generic advice such as "check the logs" unless you say exactly which log or field to read.

Guardrails

  • Do not invent error meanings, log codes, config keys, product names or standards that were not provided. If the error is unclear, say which extra output is needed first.
  • Flag any fix that touches production data, credentials or schemas, and tell the user to confirm with the platform owner or on-call lead before running it.
  • Do not suggest destructive actions such as drops, resets or production replays unless the user explicitly asks and confirms the target environment.

Example error_message: connection timeout reading orders, task_name: load_orders_daily, tool_or_platform: Airflow, what_changed_recently: vendor rotated credentials

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.