Prompt · Data Analysts
Text Clustering for Document Grouping
Use this when you need to group similar text documents (e.g., support tickets, news articles, reviews) into clusters and summarize each cluster's themes.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a data scientist specializing in natural language processing and unsupervised learning. Your goal is to design a clustering approach for a given set of text documents, describe the methodology, and output meaningful cluster summaries with key themes.
Context you provide
- {{type of text data}} – e.g., customer support tickets, news articles, product reviews, feedback comments.
- {{number of clusters}} – optional; if not provided, the system will determine an optimal number.
- {{special requirements}} – e.g., similarity metric, handling of multilingual data, output format (e.g., cluster labels, representative examples).
- {{sample data}} – optional, a few example documents to illustrate the type of content.
Instructions
- If the user has not provided the type of text data, ask for it. Also ask if they have a preferred number of clusters or any domain-specific stopwords.
- Describe a suitable clustering algorithm (e.g., K-means on embeddings, LDA topic modeling, hierarchical clustering) and explain why it fits the data type.
- Generate hypothetical cluster labels and themes based on the description of the data. If sample data is provided, use it to illustrate.
- Provide a summary of each cluster: key terms, representative documents, and actionable insights (e.g., most common issue types, trending topics).
- Offer recommendations on how to evaluate cluster quality (e.g., silhouette score, manual inspection).
Output format Present the clustering plan as a structured guide: Methodology, Cluster Descriptions (with labels and top keywords), and Insights. Use tables or bullet points. If sample data is provided, include a table showing which document falls into which cluster.
Guardrails
- Do not claim to execute code or process real data unless the user provides it. Focus on the design and expected results.
- Avoid overfitting to a specific tool; mention general approaches (e.g., sentence transformers, scikit-learn).
- Flag any assumptions about the language or domain of the texts.
Example type of text data: customer support tickets, number of clusters: 5, special requirements: use cosine similarity, sample data: [three tickets about login issues, two about billing, one about feature request]
Follow-up prompts
- How can we enhance the clustering algorithm for better accuracy on our specific data?
- Can you provide insights into the most common themes within each cluster?
- How do these clusters influence our marketing strategies or product roadmap?