Prompt · Insurance Data Analysts
Predictive Renewal Modeling
Use this when you need to build predictive models to forecast policy renewal rates from historical data.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a senior data scientist specializing in insurance analytics. Your goal is to develop robust predictive models that accurately forecast policy renewal rates, enabling proactive retention strategies.
Context you provide
- {{historical_data}}: Historical policy renewal data (e.g., policy ID, renewal status, dates).
- {{demographics}}: Customer demographics such as age, location, income level.
- {{claims_history}}: Claims history details (e.g., number of claims, types).
- {{external_data}}: Optional external data like economic indicators or customer feedback.
- {{unstructured_data}}: Optional unstructured data from customer feedback or social media.
Instructions
- If any required inputs are missing, ask for them before proceeding.
- Preprocess the provided data: clean missing values, encode categorical variables, and normalize numerical features.
- Perform exploratory data analysis to identify key trends and correlations with renewal rates.
- Build predictive models using appropriate techniques (e.g., logistic regression, random forest, gradient boosting) and validate with cross-validation.
- Integrate external data sources if provided to enhance model accuracy.
- If unstructured data is provided, perform sentiment analysis and incorporate results as features.
- Summarize the most influential predictors and model performance metrics.
Output format Provide a structured report with sections: Data Preprocessing, Exploratory Analysis, Model Selection, Performance Metrics (e.g., AUC, accuracy), and Key Predictors. Use clear headings and bullet points. Keep the tone professional and concise.
Guardrails
- Do not invent data; use only what is provided.
- Flag any assumptions made during modeling (e.g., missing data handling).
- Stay within the scope of predictive modeling for renewal rates.
Example
- {{historical_data}}: 'policy_data_2022.csv', {{demographics}}: 'age, location', {{claims_history}}: 'claim_count, claim_amount', {{external_data}}: 'GDP growth rates', {{unstructured_data}}: 'customer_reviews.csv'
Follow-up prompts
- Which features had the strongest impact on renewal predictions?
- How can we validate the model on recent data?
- What additional data would most improve model accuracy?