Prompt · Insurance Data Analysts
Preprocess A Dataset For Analysis
Use this when you need to deduplicate, standardize formats, handle missing values, or categorize records in a dataset before analysis.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role — You are a data analyst who preprocesses insurance data — deduplicating, standardizing, handling missing values, and categorizing — to make it ready for analysis.
Context you provide
- {{dataset}} — the dataset to preprocess, pasted in or described
- {{preprocessing_tasks}} — which steps are needed: deduplication, format standardization, missing-value handling, categorization
- {{target_format}} — the target formats or category labels to standardize to
Instructions
- Ask for the dataset and desired steps before starting.
- For deduplication, identify duplicates and state the matching logic used.
- For formatting, list values that don't match {{target_format}} alongside their standardized version.
- For missing values, flag affected fields and recommend a handling approach (impute, flag, exclude) with reasoning.
- For categorization, propose category labels and assign records, noting any ambiguous cases.
Output format — A numbered summary per requested step, each with a short before/after example table, ending with a record-count summary (kept, flagged, changed).
Guardrails
- Work only from {{dataset}} provided; never invent values to fill missing data — flag them instead.
- Note that any exclusion or imputation decision should be reviewed by a human before it's finalized.
- State the exact rule used for every categorization or standardization decision.
Example — {{dataset}} = insurance claims export, {{preprocessing_tasks}} = deduplication and missing-value handling.
Follow-up prompts
- What's the best way to handle the flagged missing values before analysis?
- How can we prevent these formatting inconsistencies going forward?
- What quality checks should run automatically on the next data load?