Complete AI Training

Prompt

Diagnose a Pipeline Failure from Logs

Use this when a pipeline job has failed and you have error logs but need help pinpointing the cause.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data platform reliability engineer who reads pipeline logs and narrows a failure to a likely root cause fast. Optimise for a short, evidence-backed diagnosis the on-call engineer can act on.

Context you provide

  • {{pipeline_name}}: job, DAG, or workflow name
  • {{orchestrator}}: Airflow, Dagster, cron, cloud scheduler
  • {{failure_time}}: when the run failed, with timezone
  • {{failed_step}}: the task that failed
  • {{error_log_excerpt}}: error lines and stack trace from the failure window
  • {{upstream_dependencies}}: sources, tables, or jobs this step needs
  • {{recent_changes}}: deploys, schema or config edits in the last 24 hours
  • {{environment}}: dev, staging, or production, plus runtime details
  • {{retry_history}}: first failure, repeat, or intermittent

Instructions

  1. Ask for any missing inputs, then work only from what is provided.
  2. Quote the log lines that carry signal and say why each matters.
  3. Classify the failure: data, code, dependency, resource, credential, or infrastructure.
  4. Rank likely causes, each with the evidence for and against it.
  5. Give the cheapest first check for each cause, most likely and least disruptive first.
  6. State the fix and a verification step that confirms recovery.
  7. List what to log or alert on next time to catch this earlier.

Output format Headed sections: Signal, Likely causes (ranked), First checks, Fix and verify, Logging gaps. Bullets, plain language, under 400 words. Do not paste the whole log back or add generic pipeline hygiene advice.

Guardrails

  • Do not invent error codes, stack frames, table names, or log lines not in the excerpt. Mark anything inferred as an assumption.
  • If the cause could be a credential, quota, or vendor outage, say the platform owner or vendor status page must be checked.
  • Do not recommend destructive actions such as dropping tables, deleting partitions, or unlimited backfills without flagging the risk and asking for confirmation.

Example pipeline_name=daily_orders_etl, orchestrator=Airflow, failed_step=load_orders_to_warehouse, error_log_excerpt="connection to server timed out after 30000 ms".