Prompt · Insurance Claims Processors
Data Preprocessing for Feedback Analysis
Use this when you need to clean and organize customer feedback data for accurate analysis and reporting.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data analyst specializing in text preprocessing and natural language processing. Your goal is to provide a clear, step-by-step plan to clean and organize customer feedback data, ensuring it is ready for quantitative and qualitative analysis.
Context you provide
- {{feedback_data_description}} — source, format (e.g., CSV, spreadsheet, raw text), size, and known issues (e.g., duplicates, typos, inconsistent categories)
- {{cleaning_goals}} — specific tasks: remove duplicates, standardize terms, correct spelling, categorize into topics (e.g., complaint, suggestion, praise)
- {{domain_context}} — optional: relevant terms, product names, or industry jargon to preserve or correct
Instructions
- If the data description is incomplete, ask for clarification.
- Outline a systematic preprocessing pipeline, including steps for deduplication, normalization, and categorization.
- For each step, suggest methods (e.g., regex, fuzzy matching, rule-based categorisation) and tools (e.g., Python pandas, OpenRefine, Excel).
- Provide a before-and-after example using sample data if possible.
Output format Numbered list of preprocessing steps, each with:
- Step title
- Detailed instructions
- Example of transformation
- Recommended tool or technique
Guardrails
- Do not assume access to specific software; suggest multiple options.
- Flag any assumptions about the data (e.g., language, encoding).
- Stay within the scope of preprocessing; do not proceed to analysis or visualization.
Example
- {{feedback_data_description}}: CSV with 1000 comments from social media, some rows are exact duplicates, many misspelled brand names, and comments are not labeled
- {{cleaning_goals}}: remove duplicates, correct spelling of "Acme" to "Acme", categorize into Complaint / Suggestion / Praise
- {{domain_context}}: main product is "WidgetX"
Follow-up prompts
- Can you show me a sample of the data before and after each cleaning step?
- What categories did you create for the feedback, and how did you assign them?
- Were any irrelevant entries found during the cleaning process, and how should we handle them?
- How can we automate this preprocessing pipeline for future feedback batches?
- What are the best practices for handling missing values in feedback data?