AI agent for quality assurance testers
Firmware Release Candidate Regression Agent
Show every regression in a release candidate and trace each to its cause
What it does
Before every release someone should run all unit, integration and bench tests, but the full list takes a day and gets trimmed. This agent runs the whole list on each release candidate build. It compares each result with the last released build and lists tests that now fail, run slower or use more memory. For each new failure it bisects the commits between the two builds to find the change that caused it, then tells the owner of that commit. After a fix lands it reruns the failed tests and a wider sample to make sure nothing else broke. The lead approves the release. Edge case: a bench test fails once and passes twice, so the agent labels it flaky and shows the pass rate.
How it works
Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.
Read the steps as a list
- Release candidate tagged
- Build the candidate and run unit tests
- Run integration and bench tests
- Compare each result with the last release
- Does each new failure repeat on three reruns?If not: mark it flaky with its pass rate and check the bench state. Back to step 4.
- Bisect commits between releases for each real failure
- Report the culprit commit and owner
- Rerun failed tests and a wider sample after the fix
- Is every previous pass still passing?If not: bisect the new regression and report it before continuing. Back to step 6.
- Lead approves the releaseThe agent waits here for your OK.
- Regression report signed for the release
How it decides
It calls a test a regression when it passed on the last release and fails reliably on the candidate. It then bisects to find the commit.
- Rerun a failing test three times before calling it a regression
- Flag any test that is 20% slower or uses 5% more RAM
- Block the release on any failing safety or boot test
- Treat tests with under 90% pass rate as flaky, not regressions
Make it yours
Every agent is a starting point. You choose these settings for your own situation.
- Test suites included
- Rerun count (default 3)
- Flaky pass rate (default 90%)
- Performance regression thresholds
- Who is notified of failures
What keeps you in control
It always asks you first
- Lead approves the release
- Lead approves shipping with a known flaky test
Hard limits
- Never skip the safety and boot tests
- Never tag a release, only recommend
It stops when
- Done: all tests match the last release or have an approved exception
- Stop: a blocking test fails and no fix is available
Set it up
We guide you through the set-up, step by step
Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.
- One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
- The agent then walks you through connecting your own data, one source at a time
- A downloadable copy with the flow chart, the rules and the full guide