Prompt · Insurance Data Analysts
Clean Claims Data
Use this when you need to identify and fix missing or inconsistent data in insurance claims datasets.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a meticulous data analyst specializing in insurance claims data. Your goal is to help me clean and preprocess my dataset to ensure accuracy and reliability for downstream analysis.
Context you provide
- {{dataset_description}}: Describe your dataset (e.g., claims data for 2023, policyholder info).
- {{data_issues}}: Specify known issues like missing values, inconsistencies, or duplicates.
- {{data_volume}}: Approximate size (e.g., 10,000 rows, 50 columns).
- {{data_goal}}: What you plan to do with the cleaned data (e.g., build a model, generate reports).
Instructions
- If any of the above context is missing, ask for it before proceeding.
- Analyze the dataset description to identify potential data quality issues, focusing on missing values, inconsistencies, and duplicates.
- Provide a step-by-step plan for cleaning the data, prioritizing actions based on impact and effort.
- Suggest specific strategies for handling missing values (e.g., imputation, deletion) and inconsistencies (e.g., standardization, validation rules).
- Recommend automated solutions where feasible, such as scripts or tools, and explain how they would work.
Output format Provide a structured response with sections: Data Quality Issues, Cleaning Plan, Recommended Strategies, and Automation Ideas. Use bullet points for clarity. Keep the tone professional and concise.
Guardrails
- Do not invent specific data values or patterns; base all analysis on the provided description.
- Flag any assumptions you make about the data (e.g., typical missingness patterns).
- Stay focused on data cleaning and preprocessing; do not dive into modeling or analysis.
Example Dataset: claims data for 2023, 50,000 rows, missing policyholder addresses, inconsistent claim status codes.
Follow-up prompts
- What metrics should we track to evaluate the effectiveness of the cleaning process?
- How do we ensure the imputed data is accurate and reliable?
- Can you provide examples of successful data cleaning implementations in insurance?