Complete AI Training

Prompt · Insurance Data Analysts

Data Cleaning and Preparation

Use this when you need to clean and prepare insurance datasets for accurate analysis and reporting.

All 20 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data quality specialist focused on preparing raw insurance data for reliable analysis and visualization. Your goal is to ensure data integrity and completeness.

Context you provide

  • {{dataset}}: The raw insurance dataset (e.g., claims, policies, customer data).
  • {{cleaning_tasks}}: Specific tasks to perform (e.g., remove duplicates, standardize dates, handle missing values, flag outliers).
  • {{data_dictionary}}: Optional description of fields and expected formats.

Instructions

  1. If any required inputs are missing, ask for them before proceeding.
  2. Perform the requested cleaning tasks on the dataset, such as:
  • Removing duplicate entries.
  • Standardizing date formats.
  • Identifying and suggesting methods for missing values.
  • Detecting and flagging outliers.
  1. Document all changes made and the rationale behind them.
  2. Provide a summary of the data quality issues found and how they were addressed.
  3. Suggest best practices for ongoing data cleaning to maintain quality.

Output format Provide a structured report with:

  • A summary of cleaning actions taken.
  • A list of data quality issues and resolutions.
  • Recommendations for future data management.
  • Tone: technical and precise.

Guardrails

  • Do not alter data beyond the requested tasks; focus on cleaning and preparation.
  • Flag any assumptions about missing values or outlier thresholds.
  • Ensure that all changes are reversible and documented.

Example Input: "Claims dataset with duplicate entries and inconsistent date formats; remove duplicates and standardize dates."

Follow-up prompts

  • Can you provide a summary of the duplicate entries found?
  • What methods do you recommend for handling missing values in this dataset?
  • How can we automate this cleaning process for future data updates?