Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for bioinformaticians

Workflow Run Failure Triage Agent

Get a failed pipeline run to a valid output with the least repeated compute

Workflow Run Failure Triage Agent: what goes in, what the agent does and what you get

What it does

A pipeline dies on step 40 of 60 and the log is thousands of lines long. This agent reads the workflow logs, finds the first real failure, and classifies it: memory, disk, bad or missing input, tool version, timeout or a flaky node. For each class it applies the usual fix, such as raising memory for that task, fixing a path, or pinning a tool version. It resumes from the cache so finished steps are not repeated, then checks the output file for size, format and expected rows. If the same task fails twice, or a rerun would use a lot of compute, it stops and asks the scientist. Edge case: a task seems to need more memory but the real cause is a corrupted input, so more memory never helps.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueApprovedYes, continueNoNo 1 STARTS WHEN Workflow run failed 2 USES A TOOL Read logs and find the first failed task 3 DOES Classify the cause as memory, disk, input, version,timeout or node 4 DOES Pick the smallest fix for that class 5 CHECKS THE RESULT Is this the first attempt at this task and below thecompute limit? If not: stop and ask the scientist, with the log excerptand what was tried. Back to step 3. 6 YOU APPROVE Scientist approves any rerun above the compute limit 7 USES A TOOL Apply the fix and resume from cache 8 CHECKS THE RESULT Does the output pass format, size and row-countchecks? If not: reclassify, since the first diagnosis wasprobably wrong, and pick another fix. Back to step 3. 9 RESULT Run summary with cause, fix and output check
Read the steps as a list
  1. Workflow run failed
  2. Read logs and find the first failed task
  3. Classify the cause as memory, disk, input, version, timeout or node
  4. Pick the smallest fix for that class
  5. Is this the first attempt at this task and below the compute limit?If not: stop and ask the scientist, with the log excerpt and what was tried. Back to step 3.
  6. Scientist approves any rerun above the compute limitThe agent waits here for your OK.
  7. Apply the fix and resume from cache
  8. Does the output pass format, size and row-count checks?If not: reclassify, since the first diagnosis was probably wrong, and pick another fix. Back to step 3.
  9. Run summary with cause, fix and output check

How it decides

It matches the first failure message and resource use to known failure classes and picks the smallest fix that targets that cause.

  • Raise memory by 50% for out-of-memory errors, at most twice
  • Check input checksums before any memory change
  • Pin the last working tool version after a version mismatch
  • Stop after two failures of the same task

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Compute limit that needs approval (default 200 core hours)
  • Memory increase step (default 50%)
  • Number of attempts per task (default 2)
  • Output checks to run
  • Where failure summaries are sent

What keeps you in control

It always asks you first

  • Scientist approves reruns costing over the compute limit
  • Scientist approves any change to pipeline code

Hard limits

  • Never delete the cache or raw inputs
  • Never change pipeline code without approval

It stops when

  • Done: the output passes checks and the summary is saved
  • Stop: the same task fails twice or the cause is unknown

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensA variant calling run failed at task 41 with an exit 137. The agent classified it as memory and raised the limit from 16 to 24 GB. The rerun failed again, so the output check failed. Checking the inputs it found a truncated BAM from an interrupted copy. It requested a fresh copy, resumed from cache, and the output passed with 98.7% expected rows. The scientist saw the summary.

More agents for bioinformaticians