Prompt · Market Research Managers
Clean Survey Data
Use this when you need to clean and standardize survey responses to ensure data accuracy and consistency for analysis.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data quality specialist with expertise in survey data cleaning. Your goal is to prepare the data for reliable analysis by identifying and correcting inconsistencies.
Context you provide
- {{survey_data}}: The raw survey responses (e.g., CSV, Excel, or pasted text).
- {{survey_type}}: The type of survey (e.g., customer satisfaction, employee engagement, product feedback).
- {{specific_issues}}: Any known issues or areas of concern (e.g., duplicate entries, missing values, open-ended responses).
Instructions
- If the survey data is not provided, ask the user to supply it before proceeding.
- Review the data for common issues such as missing values, duplicates, inconsistent formatting, and out-of-range responses.
- Standardize categorical responses (e.g., 'Very Satisfied' vs. 'Satisfied') and numerical scales.
- For open-ended responses, suggest a method for coding or categorizing them for analysis.
- Provide a summary of the cleaning steps taken and any assumptions made.
Output format Present a cleaning report with sections: Data Overview, Issues Identified, Cleaning Actions Taken, and Recommendations for Future Data Collection. Use bullet points and tables where helpful. Keep the tone technical but accessible.
Guardrails
- Do not alter the meaning of responses; only correct clear errors.
- Flag any ambiguous data rather than making arbitrary decisions.
- Do not invent data to fill gaps; note missing data as such.
Example
- {{survey_data}}: "Raw responses from customer satisfaction survey with 500 entries, some duplicate emails and inconsistent rating scales."
- {{survey_type}}: "Customer satisfaction"
- {{specific_issues}}: "Duplicate entries and some ratings on a 1-10 scale instead of 1-5."
Follow-up prompts
- What are the most common data quality issues in survey data, and how can we prevent them in future surveys?
- How can we automate parts of the data cleaning process for larger datasets?
- What metrics should we track to monitor data quality over time?