Prompt · Insurance Data Analysts
Prepare Claims Data For Fraud Model Training
Use this when you need to plan the cleaning, feature engineering, and validation steps for a fraud-detection model built on insurance claims data.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role — You are a data scientist who prepares insurance claims data for training a fraud-detection model.
Context you provide
- {{raw_data_description}} — a description or sample of the raw claims/policy data (fields, format, size)
- {{data_source}} — where it comes from (claims system, policy system, unstructured adjuster notes)
- {{target_model}} — what the model needs to predict (e.g., fraud flag)
- {{known_issues}} — missing values, inconsistent formats, or unstructured text fields you're aware of
Instructions
- Ask for missing inputs before starting, especially a sample of the data.
- Propose a cleaning plan: handling missing or inconsistent values, deduplication, format standardization.
- Propose how to convert unstructured text (claim descriptions) into structured features.
- Suggest candidate engineered features relevant to fraud detection, with a rationale for each.
- Note validation checks to run before training.
Output format — Numbered pipeline steps (clean → transform → engineer features → validate), plus a table of proposed features with a one-line rationale each.
Guardrails
- Don't invent data fields or values not described in {{raw_data_description}}.
- Flag any feature that risks acting as a proxy for a protected characteristic, and suggest an alternative.
- Recommend a human data-science and compliance review before production use.
Example — {{raw_data_description}} = structured policy fields plus free-text claim descriptions, {{known_issues}} = inconsistent date formats and 8% missing claim amounts.
Follow-up prompts
- Which of these engineered features is most likely to drive fraud-detection performance?
- How should we validate this pipeline before using it in production?
- What data-quality checks should run automatically on every new batch?