Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI agent for prompt engineers

Prompt Injection Red-Team Agent

Find the ways your prompt can be attacked and fix them without hurting normal use

Prompt Injection Red-Team Agent: what goes in, what the agent does and what you get

What it does

An assistant that reads emails, web pages or uploaded files can be told by that content to ignore its rules or reveal its instructions. This agent attacks your own prompt before anyone else does. It generates a set of attack inputs: direct requests to reveal the system prompt, hidden commands inside a document, role-play tricks and multi-turn pressure. It runs each against the prompt and the tools, scores every response for leaks and rule breaks, and writes stronger variants of the attacks that succeeded. It then suggests prompt or tool changes, applies them in a test copy, and reruns the whole set plus a set of normal requests to check that normal use still works. The engineer approves the final prompt. Edge case: a fix blocks the attack but also blocks a legitimate summary request, so the agent loosens it.

How it works

Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.

Start and resultWhat it doesA check on its own workWaits for your OKGoes back and retries
Yes, continueApprovedNo 1 STARTS WHEN Prompt or tool change ready 2 USES A TOOL Generate attack inputs across leak, hidden commandand role-play types 3 USES A TOOL Run each attack against the assistant in a testsetup 4 DOES Score each response for leaks and rule breaks 5 DOES Write stronger variants of the attacks thatsucceeded 6 USES A TOOL Run the stronger variants 7 DOES Suggest prompt or tool permission changes 8 USES A TOOL Apply the changes in a test copy and rerun allattacks and normal requests 9 CHECKS THE RESULT Do all attacks fail and do normal requests stillpass? If not: revise the fix, loosening it if normal requestsfail, and retest. Back to step 4. 10 YOU APPROVE Engineer approves the final prompt 11 RESULT Red-team report with attacks, results and the fix
Read the steps as a list
  1. Prompt or tool change ready
  2. Generate attack inputs across leak, hidden command and role-play types
  3. Run each attack against the assistant in a test setup
  4. Score each response for leaks and rule breaks
  5. Write stronger variants of the attacks that succeeded
  6. Run the stronger variants
  7. Suggest prompt or tool permission changes
  8. Apply the changes in a test copy and rerun all attacks and normal requests
  9. Do all attacks fail and do normal requests still pass?If not: revise the fix, loosening it if normal requests fail, and retest. Back to step 4.
  10. Engineer approves the final promptThe agent waits here for your OK.
  11. Red-team report with attacks, results and the fix

How it decides

It scores a response as failed when it reveals protected text, follows an instruction from untrusted content or uses a tool outside the allowed scope, and escalates the attacks that succeed.

  • Count any reveal of protected text as a failure
  • Count any action driven by text from a document or page as a failure
  • Run at least 100 attacks across 5 types
  • Require 100% of normal test requests to pass

Make it yours

Every agent is a starting point. You choose these settings for your own situation.

  • Attack types and count (default 100)
  • Scoring rules
  • Normal test set
  • Tools included
  • Report format

What keeps you in control

It always asks you first

  • Engineer approves the final prompt
  • Security lead approves releasing with a known residual risk

Hard limits

  • Test only your own assistant in a test setup
  • Never use real customer data in attacks

It stops when

  • Done: all attacks fail and normal requests pass
  • Stop: the attack succeeds and needs a design change such as removing a tool

Set it up

We guide you through the set-up, step by step

Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.

10 minto set it up in your AI
5 AIsChatGPT, Claude, Copilot, Gemini, Grok
  • One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
  • The agent then walks you through connecting your own data, one source at a time
  • A downloadable copy with the flow chart, the rules and the full guide
Get access to this agent

An example run

What happensAgainst a support assistant, the agent ran 120 attacks. Nine succeeded, mostly hidden commands in pasted emails that made the assistant offer a refund. Stronger variants raised successes to 14. The agent added a rule that text from emails is never an instruction, and the retest cut it to 1, but a normal request to summarize an email failed, so the check failed. A narrower rule fixed both, and the engineer approved.

More agents for prompt engineers