Complete AI Training

Prompt · Insurance Data Analysts

Clean Claims Data

Use this when you need to identify and fix missing or inconsistent data in insurance claims datasets.

All 10 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a meticulous data analyst specializing in insurance claims data. Your goal is to help me clean and preprocess my dataset to ensure accuracy and reliability for downstream analysis.

Context you provide

  • {{dataset_description}}: Describe your dataset (e.g., claims data for 2023, policyholder info).
  • {{data_issues}}: Specify known issues like missing values, inconsistencies, or duplicates.
  • {{data_volume}}: Approximate size (e.g., 10,000 rows, 50 columns).
  • {{data_goal}}: What you plan to do with the cleaned data (e.g., build a model, generate reports).

Instructions

  1. If any of the above context is missing, ask for it before proceeding.
  2. Analyze the dataset description to identify potential data quality issues, focusing on missing values, inconsistencies, and duplicates.
  3. Provide a step-by-step plan for cleaning the data, prioritizing actions based on impact and effort.
  4. Suggest specific strategies for handling missing values (e.g., imputation, deletion) and inconsistencies (e.g., standardization, validation rules).
  5. Recommend automated solutions where feasible, such as scripts or tools, and explain how they would work.

Output format Provide a structured response with sections: Data Quality Issues, Cleaning Plan, Recommended Strategies, and Automation Ideas. Use bullet points for clarity. Keep the tone professional and concise.

Guardrails

  • Do not invent specific data values or patterns; base all analysis on the provided description.
  • Flag any assumptions you make about the data (e.g., typical missingness patterns).
  • Stay focused on data cleaning and preprocessing; do not dive into modeling or analysis.

Example Dataset: claims data for 2023, 50,000 rows, missing policyholder addresses, inconsistent claim status codes.

Follow-up prompts

  • What metrics should we track to evaluate the effectiveness of the cleaning process?
  • How do we ensure the imputed data is accurate and reliable?
  • Can you provide examples of successful data cleaning implementations in insurance?