AI agent for journalists
Document Dump Triage Agent
Give the reporter a ranked reading list from a large document set with notes and entities.
What it does
A leak or records release can contain thousands of files, and the reporter cannot read them all. This agent sorts the set by type, date and source, extracts names, dates, organizations and amounts, and flags documents likely to answer the story question the journalist writes. It reads the flagged ones and checks that the match is real, not just a keyword hit, and notes what each contains. It then refines its criteria from what the reporter confirms and rejects, and reruns on the full set to find missed documents. It produces a reading list ordered by value, with a short note on each. It never decides what is newsworthy and does not contact anyone. The journalist approves the reading list. Edge case: a scanned document with poor text is flagged for manual review.
How it works
Follow the arrows from top to bottom. The orange dashed arrow is the loop: when a check fails, the agent goes back and tries again.
Read the steps as a list
- Document set added
- Sort files by type, date and source and extract text
- Is the text readable in each file?If not: mark the file for manual review and retry extraction with another method. Back to step 2.
- Extract names, dates, organizations and amounts
- Score documents against the story question
- Read the top documents and confirm each match
- Build the reading list with notes
- Journalist marks useful and not useful documentsThe agent waits here for your OK.
- Do the confirmed documents show patterns the criteria missed?If not: refine the criteria and rerun on the full set. Back to step 5.
- Ranked reading list and entity index
How it decides
It scores a document by how many entities and facts in it relate to the story question, and treats a keyword-only match as weak until its content is checked.
- Rank higher when a document matches two or more key entities
- Treat a keyword-only match as weak until checked
- Mark poor scans for manual review
- Rerun after every batch of feedback
Make it yours
Every agent is a starting point. You choose these settings for your own situation.
- Story question
- Entities of interest
- Date range
- Reading list size (default 50)
- Extraction methods
What keeps you in control
It always asks you first
- The reading list
- Any contact with sources
Hard limits
- Never publish or share the documents
- Keep the original files unchanged
It stops when
- Done: reporter has what is needed
- Stop: set is unreadable
Set it up
We guide you through the set-up, step by step
Members get the full set-up guide for this agent. No technical skills needed: you copy, paste and upload.
- One set of instructions to paste into your AI, with the clicks for ChatGPT, Claude, Microsoft 365 Copilot, Gemini and Grok
- The agent then walks you through connecting your own data, one source at a time
- A downloadable copy with the flow chart, the rules and the full guide