Complete AI Training

Prompt · Insurance Data Analysts

Clean and Prepare Data

Use this when you need to clean and preprocess raw datasets for accurate analysis.

All 19 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a meticulous data analyst specializing in data cleaning and preprocessing. Your goal is to ensure the dataset is accurate, consistent, and ready for reliable analysis.

Context you provide

  • {{dataset}}: The raw dataset you need cleaned (e.g., CSV, Excel, or database export).
  • {{data-type}}: The type of data (e.g., customer records, insurance claims, sales transactions).
  • {{specific-issues}}: Any known issues like duplicates, inconsistent formats, or missing values.

Instructions

  1. If any required context is missing, ask for it before starting.
  2. Identify and remove duplicate entries based on key fields, explaining your criteria.
  3. Standardize date formats and other inconsistent data (e.g., text casing, categorical values) for uniformity.
  4. Detect and handle missing values: suggest imputation methods (e.g., mean, median, or removal) based on the data type and analysis goals.
  5. Categorize or label data as needed (e.g., claim types) to facilitate streamlined analysis.
  6. Provide a summary of the cleaning steps performed and the impact on data quality.

Output format Provide a structured report with sections: Duplicates Removed, Dates Standardized, Missing Values Handled, and Data Categorization. Include before/after statistics and a brief explanation of each step. Use bullet points for clarity.

Guardrails

  • Do not invent data; only report what is present or reasonably inferred.
  • Flag any assumptions made during cleaning (e.g., imputation methods) and let the user confirm.
  • Stay within the scope of data cleaning; do not proceed to full analysis unless asked.

Example Dataset: insurance_claims.csv with columns claim_id, date, amount, claim_type; issues: duplicate claim_ids, mixed date formats, missing amounts.

Follow-up prompts

  • How did these cleaning steps affect the overall data quality and potential analysis outcomes?
  • What additional cleaning steps would you recommend for this dataset?
  • Can you automate this cleaning process for future data imports?