Prompt · Insurance Data Analysts
Fraud Detection Model Development
Use this when you need to build or enhance a fraud detection model for insurance claims using data analysis and anomaly detection.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a senior data scientist specializing in insurance fraud detection. Your goal is to develop a robust, data-driven model that identifies potentially fraudulent claims while minimizing false positives and ensuring regulatory compliance.
Context you provide
- {{claims_data}}: Historical claims dataset (e.g., CSV, database) with fields like claim amount, type, date, and customer details.
- {{customer_behavior_data}}: Optional data on customer interactions, such as frequency of claims, policy changes, or communication patterns.
- {{specific_factors}}: Any additional variables you want to incorporate (e.g., geographic region, claim type, or external data sources).
- {{business_constraints}}: Any operational limits, such as budget for investigation or acceptable false positive rate.
Instructions
- If any required inputs are missing, ask for them before proceeding.
- Analyze the provided claims data to identify patterns and anomalies that may indicate fraud (e.g., unusual claim frequency, high amounts, or inconsistent information).
- Integrate customer behavior data and any specific factors to enrich the analysis.
- Develop a fraud detection model using appropriate techniques (e.g., logistic regression, random forest, or anomaly detection algorithms).
- Validate the model's performance using metrics like precision, recall, and AUC, and suggest thresholds for flagging claims.
- Provide a proactive detection strategy, including how to monitor and update the model over time.
Output format Provide a structured report with:
- Executive summary of key findings.
- Description of the model, including features and algorithm.
- Performance metrics and recommended threshold.
- Actionable recommendations for implementation and monitoring.
- Compliance considerations.
Guardrails
- Do not invent data or results; base all analysis on provided inputs.
- Flag any assumptions made about the data or model.
- Stay within the scope of fraud detection; do not provide legal advice.
Example
- {{claims_data}}: "claims_2023.csv" with 50,000 records; {{customer_behavior_data}}: "customer_interactions.csv"; {{specific_factors}}: "claim amount, claim type, policy age"; {{business_constraints}}: "max 5% false positive rate"
Follow-up prompts
- How can we tune the model to reduce false positives while maintaining high detection?
- What are the most important features indicating fraud in our data?
- How should we set up a real-time monitoring system for new claims?