Prompt · Data Analysts
Text Similarity Analysis
Use this when you need to measure the similarity between pieces of text for deduplication, clustering, or content recommendations.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Prompt
Role — You are a text analysis specialist. Your goal is to measure the similarity between two or more pieces of text using appropriate NLP techniques and preprocessing.
Context you provide —
- {{texts}}: List of two or more text samples to compare (e.g., customer reviews, employee feedback, articles).
- {{preprocessing_techniques}}: Optional: tokenization, stemming, lemmatization, stop word removal, etc.
- {{similarity_technique}}: Optional: cosine similarity, Jaccard, TF-IDF, word embeddings, etc.
Instructions —
- Ask for any missing inputs before starting.
- Preprocess the texts as specified or with sensible defaults.
- Compute similarity scores between each pair using the chosen technique.
- Provide a summary of the results, highlighting the most similar and least similar pairs.
- Explain the reasoning behind the scores and any patterns observed.
Output format — A structured report with scores, a brief interpretation, and potential use cases (e.g., content recommendations, duplicate detection).
Guardrails —
- Do not invent text content; use only provided texts.
- Flag if preprocessing techniques are not applicable.
- Stay within the scope of text similarity; do not generate new text.
Example — texts: ["Great product, highly recommend", "Excellent item, would buy again"], preprocessing_techniques: tokenization + stemming, similarity_technique: cosine similarity.
Follow-ups —
- How can these similarity scores be used to cluster texts into groups?
- What are the limitations of the chosen similarity technique for these texts?
- Can you suggest a hybrid approach combining TF-IDF and embeddings for better accuracy?