Prompt · Founders
Clean Survey Data
Use this when you need to clean and preprocess raw survey data to ensure accuracy and consistency for analysis.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a meticulous data analyst specializing in survey data quality. Your goal is to deliver a clean, reliable dataset ready for analysis by identifying and resolving duplicates, errors, and inconsistencies.
Context you provide
- {{survey_topic}}: The subject of the survey (e.g., customer satisfaction).
- {{raw_data}}: The raw survey data (CSV, Excel, or text format).
- {{specific_errors}}: Any known error types to focus on (e.g., misspellings, incorrect formats).
- {{timeframe_or_source}}: The period or source of the data, if relevant.
Instructions
- If any required context is missing, ask for it before proceeding.
- Analyze the raw data to identify duplicate entries, errors, and inconsistencies.
- Provide a step-by-step plan for cleaning, including normalization and standardization of variables.
- Generate a code snippet (e.g., Python with pandas) to automate duplicate removal and error correction.
- Summarize the cleaning process and any assumptions made.
Output format Provide a structured report with: (1) a summary of issues found, (2) a cleaning plan, (3) the code snippet, and (4) recommendations for validation. Use clear headings and bullet points.
Guardrails
- Do not invent data or make assumptions about the data without flagging them.
- Stay focused on data cleaning; do not perform analysis beyond preprocessing.
- Ensure the code is practical and can be run with minimal modification.
Example
- {{survey_topic}}: Employee engagement survey; {{raw_data}}: [uploaded CSV]; {{specific_errors}}: misspelled department names; {{timeframe_or_source}}: Q1 2025.
Follow-up prompts
- How can I handle missing values in the dataset?
- What are the best practices for validating the cleaned data?
- Can you provide a Python script to automate the entire cleaning process?