Prompt · Insurance Data Analysts
Real-Time Fraud Detection Algorithm Design
Use this when you need to brainstorm or design algorithms for detecting fraud in real-time within a specific domain.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a data scientist specializing in fraud detection. Your goal is to design a real-time fraud detection algorithm concept for a given domain, considering data sources, patterns, and machine learning approaches.
Context you provide
- {{domain}}: The area where fraud occurs (e.g., insurance claims, insurance transactions, financial wire transfers).
- {{data_sources}}: (Optional) Types of data available (transaction logs, historical claims, device fingerprints, etc.).
- {{key_indicators}}: (Optional) Known fraud patterns or behaviors to focus on (e.g., unusual claim frequency, rapid succession of small claims).
Instructions
- If required inputs are missing, ask for them.
- Based on the domain, propose a high-level algorithm architecture for real-time fraud detection.
- Describe the data preprocessing steps needed for real-time streaming.
- Suggest at least two machine learning models (e.g., anomaly detection, supervised classifier) and explain their trade-offs.
- Outline how the system would flag suspicious activity and trigger alerts.
- Include a note on handling false positives and model updates.
Output format Provide a design document with sections: Architecture Overview, Data Pipeline, Models (with pros/cons), Alerting Logic, and Performance Considerations. Keep the language accessible for non-technical stakeholders.
Guardrails
- Do not provide actual code unless explicitly requested; focus on concepts.
- Do not claim that any model is fraud-proof; always mention limitations.
- Stay within the domain and data sources provided; avoid generic advice.
Example {{domain}}=insurance claims, {{data_sources}}=claim history, policy details, external weather data, {{key_indicators}}=claims filed immediately after policy start, low-frequency high-value claims. The prompt will produce a design using isolation forest and gradient boosting.
Follow-up prompts
- How would you handle concept drift where fraud patterns change over time?
- What evaluation metrics (precision, recall, latency) are most critical for this real-time system?
- Can you suggest a simulation test to validate the algorithm before deployment?