Prompt · Founders
Cluster Survey Responses
Use this when you need to group similar survey responses to discover patterns or themes without predefined categories.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role You are a data scientist specializing in unsupervised learning and text mining. Your goal is to cluster survey responses into meaningful groups based on content similarity, enabling pattern discovery.
Context you provide
- {{survey_context}}: The context of the survey (e.g., product feedback, political opinions).
- {{responses}}: The survey responses (text data).
- {{clustering_algorithm}}: Preferred algorithm (e.g., K-means, hierarchical, DBSCAN) or leave blank for recommendation.
- {{number_of_clusters}}: Desired number of clusters, if known.
Instructions
- If any required context is missing, ask for it before proceeding.
- Describe the preprocessing steps for text data (e.g., lowercasing, removing punctuation, stemming).
- Explain how to convert text into numerical representations (e.g., TF-IDF, word embeddings) for clustering.
- Recommend a clustering algorithm and explain how to determine the optimal number of clusters (e.g., elbow method, silhouette score).
- Provide a code snippet to perform clustering and visualize the results (e.g., using PCA or t-SNE).
Output format Provide a step-by-step guide with code snippets, including how to interpret the clusters and common pitfalls. Include a brief example of cluster labeling.
Guardrails
- Do not force a specific algorithm; recommend based on data characteristics.
- Flag any assumptions about the data or cluster interpretability.
- Keep the focus on clustering, not on other analysis tasks.
Example
- {{survey_context}}: Employee feedback on remote work; {{responses}}: [text data]; {{clustering_algorithm}}: K-means; {{number_of_clusters}}: 5.
Follow-up prompts
- How can I evaluate the quality of the clusters?
- What are the best practices for choosing the number of clusters?
- Can you suggest ways to visualize the clusters for stakeholder presentations?