AI agent for risk managers
Credit Scorecard Backtest Agent
Show whether the scorecard still ranks risk correctly and what cutoff changes would do
What it does
A scorecard built three years ago still approves loans, but defaults are creeping up in some segments. This agent scores a past cohort with the current scorecard and compares the predicted and actual default rates. It checks stability across segments such as region, product and score band, and measures how far each has drifted. It flags segments where predictions are off by more than a set amount. It then tests alternative cutoffs and reruns the comparison to see their effect on approval rate and loss rate. It documents each run so results are reproducible. The risk lead approves any change to the scorecard. Edge case: a segment with too few loans is reported as insufficient data rather than drift.
How it works
Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.
Read the steps as a list
- Quarterly backtest
- Load the cohort and the current scorecard
- Score the cohort and compare predicted with actual defaults
- Does each segment have enough loans to judge?If not: Merge small segments or mark them insufficient data. Back to step 3.
- Measure stability and drift by segment
- Is drift within 20 percent in every segment?If not: Flag the segments and test alternative cutoffs. Back to step 5.
- Rerun the comparison with each alternative cutoff
- Draft the report with approval and loss effects
- Risk lead approves any scorecard changeThe agent waits here for your OK.
- Backtest report and run log
How it decides
It flags a segment when predicted default differs from actual by more than 20 percent with enough loans and picks cutoffs that keep losses within target.
- Require at least 200 loans per segment
- Flag drift above 20 percent
- Pick cutoffs that keep losses within target
- Log every run with its inputs
Make it yours
Every agent is a starting point. You choose these settings for your own situation.
- Minimum loans per segment (default 200)
- Drift tolerance (default 20 percent)
- Cutoffs tested (default plus and minus 15 points)
- Run frequency (default quarterly)
What keeps you in control
It always asks you first
- Risk lead approves any scorecard or cutoff change
Hard limits
- Never change the live scorecard
- Never use data that is not in the cohort
- Keep results reproducible
It stops when
- Done: Report approved and decisions recorded
- Stop: Outcome data is incomplete, so hand to the risk lead
Set it up
We guide you through the set-up, step by step
Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.
- One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
- The agent then walks you through connecting your own data, one source at a time
- A downloadable copy with the flow chart, the rules and the full guide