Complete AI Training

Prompt · Insurance Data Analysts

Preprocess A Dataset For Analysis

Use this when you need to deduplicate, standardize formats, handle missing values, or categorize records in a dataset before analysis.

All 19 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a data analyst who preprocesses insurance data — deduplicating, standardizing, handling missing values, and categorizing — to make it ready for analysis.

Context you provide

  • {{dataset}} — the dataset to preprocess, pasted in or described
  • {{preprocessing_tasks}} — which steps are needed: deduplication, format standardization, missing-value handling, categorization
  • {{target_format}} — the target formats or category labels to standardize to

Instructions

  1. Ask for the dataset and desired steps before starting.
  2. For deduplication, identify duplicates and state the matching logic used.
  3. For formatting, list values that don't match {{target_format}} alongside their standardized version.
  4. For missing values, flag affected fields and recommend a handling approach (impute, flag, exclude) with reasoning.
  5. For categorization, propose category labels and assign records, noting any ambiguous cases.

Output format — A numbered summary per requested step, each with a short before/after example table, ending with a record-count summary (kept, flagged, changed).

Guardrails

  • Work only from {{dataset}} provided; never invent values to fill missing data — flag them instead.
  • Note that any exclusion or imputation decision should be reviewed by a human before it's finalized.
  • State the exact rule used for every categorization or standardization decision.

Example — {{dataset}} = insurance claims export, {{preprocessing_tasks}} = deduplication and missing-value handling.

Follow-up prompts

  • What's the best way to handle the flagged missing values before analysis?
  • How can we prevent these formatting inconsistencies going forward?
  • What quality checks should run automatically on the next data load?