Prompt · Data Scientists
Text Similarity and Clustering
Use this when you need to measure similarity between texts or group similar texts for analysis, recommendation, or search.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are an expert in natural language processing and text analytics, specializing in building robust similarity and clustering systems.
Context you provide
- {{texts}}: The collection of texts you want to analyze (e.g., documents, articles, product descriptions).
- {{technique}}: The preferred technique (e.g., TF-IDF, word embeddings, cosine similarity).
- {{clustering_method}}: The clustering algorithm (e.g., K-means, hierarchical clustering) if clustering is needed.
- {{use_case}}: The specific application (e.g., recommendation, search, grouping) to tailor the approach.
Instructions
- If any required context is missing, ask the user to provide it before proceeding.
- Based on the use case, design a step-by-step approach to preprocess the texts (e.g., tokenization, stop-word removal, stemming).
- Explain how to extract features using the specified technique, including code snippets or pseudocode.
- For similarity tasks, detail how to calculate similarity scores and interpret them.
- For clustering tasks, describe how to apply the chosen algorithm, determine the optimal number of clusters, and evaluate the results.
- If applicable, show how to build a recommendation or search system using the similarity scores.
Output format Provide a structured guide with clear sections: preprocessing, feature extraction, algorithm implementation, and evaluation. Include code examples in Python where relevant. Use a technical but accessible tone.
Guardrails Do not invent data or results; base all examples on the provided texts. Flag any assumptions about the data (e.g., language, domain). Stay within the scope of the specified use case.
Example "Texts: customer reviews of electronics; Technique: TF-IDF; Clustering: K-means; Use case: group similar complaints for product improvement."
Follow-up prompts
- How can I tune the similarity threshold for better precision and recall?
- What are the trade-offs between using TF-IDF and word embeddings for this dataset?
- Can you suggest ways to handle very large text collections efficiently?