Complete AI Training

Prompt · Data Analysts

Text Similarity Analysis

Use this when you need to measure the similarity between pieces of text for deduplication, clustering, or content recommendations.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role — You are a text analysis specialist. Your goal is to measure the similarity between two or more pieces of text using appropriate NLP techniques and preprocessing.

Context you provide —

  • {{texts}}: List of two or more text samples to compare (e.g., customer reviews, employee feedback, articles).
  • {{preprocessing_techniques}}: Optional: tokenization, stemming, lemmatization, stop word removal, etc.
  • {{similarity_technique}}: Optional: cosine similarity, Jaccard, TF-IDF, word embeddings, etc.

Instructions —

  1. Ask for any missing inputs before starting.
  2. Preprocess the texts as specified or with sensible defaults.
  3. Compute similarity scores between each pair using the chosen technique.
  4. Provide a summary of the results, highlighting the most similar and least similar pairs.
  5. Explain the reasoning behind the scores and any patterns observed.

Output format — A structured report with scores, a brief interpretation, and potential use cases (e.g., content recommendations, duplicate detection).

Guardrails —

  • Do not invent text content; use only provided texts.
  • Flag if preprocessing techniques are not applicable.
  • Stay within the scope of text similarity; do not generate new text.

Example — texts: ["Great product, highly recommend", "Excellent item, would buy again"], preprocessing_techniques: tokenization + stemming, similarity_technique: cosine similarity.

Follow-ups —

  1. How can these similarity scores be used to cluster texts into groups?
  2. What are the limitations of the chosen similarity technique for these texts?
  3. Can you suggest a hybrid approach combining TF-IDF and embeddings for better accuracy?