Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for quality assurance testers

Firmware Release Candidate Regression Agent

Show every regression in a release candidate and trace each to its cause

Firmware Release Candidate Regression Agent: what goes in, what the agent does and what you get

What it does

Before every release someone should run all unit, integration and bench tests, but the full list takes a day and gets trimmed. This agent runs the whole list on each release candidate build. It compares each result with the last released build and lists tests that now fail, run slower or use more memory. For each new failure it bisects the commits between the two builds to find the change that caused it, then tells the owner of that commit. After a fix lands it reruns the failed tests and a wider sample to make sure nothing else broke. The lead approves the release. Edge case: a bench test fails once and passes twice, so the agent labels it flaky and shows the pass rate.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueYes, continueApprovedNoNo 1 STARTS WHEN Release candidate tagged 2 USES A TOOL Build the candidate and run unit tests 3 USES A TOOL Run integration and bench tests 4 DOES Compare each result with the last release 5 CHECKS THE RESULT Does each new failure repeat on three reruns? If not: mark it flaky with its pass rate and check thebench state. Back to step 4. 6 USES A TOOL Bisect commits between releases for each realfailure 7 DOES Report the culprit commit and owner 8 USES A TOOL Rerun failed tests and a wider sample after the fix 9 CHECKS THE RESULT Is every previous pass still passing? If not: bisect the new regression and report it beforecontinuing. Back to step 6. 10 YOU APPROVE Lead approves the release 11 RESULT Regression report signed for the release
Read the steps as a list
  1. Release candidate tagged
  2. Build the candidate and run unit tests
  3. Run integration and bench tests
  4. Compare each result with the last release
  5. Does each new failure repeat on three reruns?If not: mark it flaky with its pass rate and check the bench state. Back to step 4.
  6. Bisect commits between releases for each real failure
  7. Report the culprit commit and owner
  8. Rerun failed tests and a wider sample after the fix
  9. Is every previous pass still passing?If not: bisect the new regression and report it before continuing. Back to step 6.
  10. Lead approves the releaseThe agent waits here for your OK.
  11. Regression report signed for the release

How it decides

It calls a test a regression when it passed on the last release and fails reliably on the candidate. It then bisects to find the commit.

  • Rerun a failing test three times before calling it a regression
  • Flag any test that is 20% slower or uses 5% more RAM
  • Block the release on any failing safety or boot test
  • Treat tests with under 90% pass rate as flaky, not regressions

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Test suites included
  • Rerun count (default 3)
  • Flaky pass rate (default 90%)
  • Performance regression thresholds
  • Who is notified of failures

What keeps you in control

It always asks you first

  • Lead approves the release
  • Lead approves shipping with a known flaky test

Hard limits

  • Never skip the safety and boot tests
  • Never tag a release, only recommend

It stops when

  • Done: all tests match the last release or have an approved exception
  • Stop: a blocking test fails and no fix is available

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensCandidate 4.2.0-rc2 ran 1,240 tests. Three failed that passed in 4.1.3. One passed on rerun and was marked flaky at 66%. The other two shared a cause: a bisect across 38 commits pointed to a timer change. After the fix the rerun was green, but the wider sample showed a UART test newly failing, so the check failed. The agent bisected that to the same commit's side effect and the lead approved after the second fix.

More agents for quality assurance testers