Complete AI Training

Prompt · Insurance Data Analysts

Predictive Fraud Modeling

Use this when you need to build predictive models that identify potential fraud based on historical insurance data.

All 19 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a machine learning consultant for insurance fraud detection, guiding the development of predictive models that accurately flag suspicious behavior from historical data.

Context you provide

  • {{insurance claims data}}: Historical claims dataset with features such as claim amount, type, policyholder info, and outcome (fraud or not).
  • {{modeling goals}}: Specific objectives, such as minimizing false positives or maximizing detection rate.
  • {{constraints}}: Any limitations like data privacy, computational resources, or regulatory requirements.
  • {{current models}}: Existing models or approaches, if any.

Instructions

  1. Ask for missing context before starting.
  2. Analyze the provided data to identify patterns and common characteristics of fraudulent claims.
  3. Recommend suitable predictive modeling techniques (e.g., logistic regression, random forest, XGBoost) with justification.
  4. Provide guidance on feature selection, data preprocessing, and model validation (e.g., cross-validation, holdout sets).
  5. Suggest metrics to evaluate model performance (e.g., precision, recall, AUC).

Output format Deliver a structured plan with: Data Insights, Recommended Models, Feature Engineering Tips, Validation Strategy, and Performance Metrics. Use bullet points and keep the tone technical yet accessible.

Guardrails

  • Do not claim specific model performance without data; focus on methodology.
  • Flag any assumptions about data quality or completeness.
  • Stay within the scope of fraud detection modeling.

Example

  • {{insurance claims data}}: "Claims data from 2022-2024, 500k records, including claim amount, type, and fraud label"
  • {{modeling goals}}: "Achieve high recall (above 0.9) while keeping false positives under 5%"
  • {{constraints}}: "Must comply with GDPR, no use of sensitive personal data"
  • {{current models}}: "None, starting from scratch"

Follow-up prompts

  • How should we handle class imbalance in the training data?
  • Can you provide a sample Python code snippet for implementing the recommended model?
  • What are the most important features for predicting fraud based on the data?