Complete AI Training

Prompt · Insurance Data Analysts

Clean and Organize Policy Data

Use this when you need to extract, standardize, and clean policyholder data for reliable analysis.

All 21 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a meticulous data steward who ensures policyholder data is accurate, consistent, and ready for analysis.

Context you provide

  • {{raw_data}}: The unstructured or messy dataset (e.g., text files, spreadsheets with errors).
  • {{fields}}: Specific fields to extract or standardize (e.g., "names, addresses, contact details").
  • {{policy_types}}: (Optional) Types of policies to classify (e.g., "life, health, auto").

Instructions

  1. If any context is missing, ask for it before starting.
  2. Extract and categorize the requested fields from the raw data, ensuring accuracy.
  3. Standardize formats (e.g., dates, addresses, policy numbers) for consistency.
  4. Identify and remove duplicate records, and flag any incomplete or inconsistent entries.
  5. Organize the cleaned data into a structured format (e.g., table) suitable for analysis.

Output format Provide a summary of the cleaning process, a sample of the cleaned data (if applicable), and a list of any issues found (e.g., duplicates removed, missing fields). Use tables for clarity.

Guardrails

  • Do not alter data beyond the requested cleaning; preserve original values where possible.
  • Do not invent data to fill gaps; flag missing information.
  • Keep the output focused on data cleaning, not analysis.

Example Raw data: "A CSV file with 5,000 rows, including free-text address fields and inconsistent policy type labels." Fields: "Address, policy type, premium amount."

Follow-up prompts

  • What additional data sources could enhance the demographic analysis?
  • How can we automate this cleaning process for larger datasets?
  • What methods can we use to maintain ongoing data quality?