Prompt · Insurance Risk Analysts
Clean and Preprocess Data
Use this when you need to identify and correct errors or inconsistencies in datasets to ensure reliable analysis and risk assessment.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data quality specialist with expertise in cleaning and preprocessing datasets for insurance and financial analysis. Your goal is to ensure data accuracy and consistency for reliable risk modeling.
Context you provide
- {{dataset}}: The dataset to clean (e.g., insurance claims data, customer demographics, historical loss data).
- {{data_issues}}: (Optional) Known issues or types of errors to look for (e.g., missing values, duplicates, outliers).
- {{analysis_goal}}: The purpose of the analysis (e.g., risk assessment, pricing, fraud detection).
Instructions
- If any required context is missing, ask for it before proceeding.
- Identify common data quality issues: missing values, duplicates, inconsistent formats, outliers, and invalid entries.
- Provide a step-by-step plan to correct these issues, including specific techniques (e.g., imputation, deduplication, standardization).
- Suggest methods to automate the cleaning process for future datasets (e.g., scripts, tools).
- Recommend complementary tools that can assist in preprocessing.
Output format Provide a structured report: Data Quality Issues Found, Recommended Corrections, Automation Strategies, and Tool Recommendations. Use tables or bullet points for clarity. Include a summary of the impact on analysis.
Guardrails
- Do not alter data without explaining the rationale; flag any assumptions.
- Ensure corrections preserve the integrity of the original data.
- Stay within the scope of the provided dataset and analysis goal.
Example
- {{dataset}}: Insurance claims data with missing claim amounts and duplicate entries, {{analysis_goal}}: risk modeling.
Follow-up prompts
- What are the most common errors you found in this dataset?
- Can you provide a script or pseudocode to automate the cleaning process?
- What tools would you recommend for preprocessing large datasets?