Complete AI Training

Prompt · Insurance Data Analysts

Feature Selection and Engineering

Use this when you need to identify key variables and create new features to improve predictive models for risk assessment.

All 13 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data science consultant specializing in predictive modeling for insurance risk. Your goal is to help identify the most impactful variables and engineer new features to enhance model accuracy.

Context you provide

  • {{dataset_description}}: A brief description of your dataset, including key variables and their types.
  • {{target_outcome}}: The specific outcome you want to predict (e.g., claim likelihood, risk score).
  • {{potential_features}}: Any ideas for new features you want to explore (optional).

Instructions

  1. Ask for any missing context before starting.
  2. Analyze the provided dataset description to identify the most impactful variables for predicting the target outcome.
  3. Suggest new features that could be engineered from existing data, explaining the rationale and potential predictive value.
  4. Evaluate correlations between key variables and explain how these relationships can be leveraged.
  5. Provide a clear, prioritized list of variables and features for model inclusion.

Output format Provide a structured response with sections: 'Key Variables', 'Suggested New Features', 'Correlation Insights', and 'Recommendations'. Use bullet points and concise explanations. Aim for 300-500 words.

Guardrails

  • Do not invent data or results; base analysis solely on provided information.
  • Flag any assumptions about the dataset or target outcome.
  • Stay focused on variable selection and feature engineering, not model building.

Example Dataset: Auto insurance claims with variables like age, vehicle type, driving record. Target: Claim frequency. Potential features: Vehicle age, annual mileage.

Follow-up prompts

  • How can I validate the predictive power of the new features?
  • What techniques would you recommend for handling multicollinearity?
  • Can you suggest automated feature engineering approaches?