Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for prompt engineers

Refusal Boundary Test Agent

Set the assistant's boundaries so valid requests are answered and disallowed ones are refused

Refusal Boundary Test Agent: what goes in, what the agent does and what you get

What it does

An assistant that refuses harmless questions frustrates users, and one that accepts risky ones creates liability. This agent builds two sets of cases: clearly allowed requests that sit close to a boundary, such as a nurse asking about medication doses, and clearly disallowed requests, such as asking for another person's data. It runs both sets against the prompt, classifies each outcome as correct, over-refusal or under-refusal, and proposes wording changes to the instructions. After each change it retests both sets, because tightening one side often breaks the other. It repeats until both error rates are within limits. The engineer approves the final wording. Edge case: a request is allowed for staff but not for customers, so the agent tests it under both roles.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueApprovedNo 1 STARTS WHEN Prompt or policy change 2 DOES Read the policy and list allowed and disallowedcategories 3 USES A TOOL Build borderline allowed and clearly disallowedcases per role 4 USES A TOOL Run all cases against the assistant 5 DOES Classify outcomes as correct, over-refusal orunder-refusal 6 DOES Propose wording changes to the instructions 7 USES A TOOL Apply the change in a test copy and rerun both sets 8 CHECKS THE RESULT Are over-refusals under 5% and under-refusals at 0on the clear cases? If not: revise the wording and retest, watching theother side. Back to step 6. 9 YOU APPROVE Engineer approves the final wording 10 RESULT Boundary test report
Read the steps as a list
  1. Prompt or policy change
  2. Read the policy and list allowed and disallowed categories
  3. Build borderline allowed and clearly disallowed cases per role
  4. Run all cases against the assistant
  5. Classify outcomes as correct, over-refusal or under-refusal
  6. Propose wording changes to the instructions
  7. Apply the change in a test copy and rerun both sets
  8. Are over-refusals under 5% and under-refusals at 0 on the clear cases?If not: revise the wording and retest, watching the other side. Back to step 6.
  9. Engineer approves the final wordingThe agent waits here for your OK.
  10. Boundary test report

How it decides

It classifies each reply against the written policy and accepts a change only if both over-refusals and under-refusals stay within their limits.

  • Require zero accepted cases from the clearly disallowed set
  • Accept up to 5% over-refusal on the borderline allowed set
  • Test every case in each user role
  • Rerun both sets after every change

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Over-refusal limit (default 5%)
  • Roles to test
  • Policy categories
  • Case count per category
  • Report format

What keeps you in control

It always asks you first

  • Engineer approves the final wording
  • Policy owner approves any change to what is allowed

Hard limits

  • Never loosen a rule on safety or privacy without the policy owner
  • Test in a test copy only

It stops when

  • Done: both error rates are within limits
  • Stop: the policy is ambiguous and the policy owner must decide

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensThe first run of 160 cases showed 14% over-refusal on medical dosage questions from nurses and 2 under-refusals on requests for another customer's order. The agent added a role-based rule. Over-refusal fell to 3%, but one under-refusal remained and a new one appeared in the customer role, so the check failed. A narrower wording fixed both, and the engineer approved.

More agents for prompt engineers