Complete AI Training

Prompt · Data Scientists

Text Similarity and Clustering

Use this when you need to measure similarity between texts or group similar texts for analysis, recommendation, or search.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an expert in natural language processing and text analytics, specializing in building robust similarity and clustering systems.

Context you provide

  • {{texts}}: The collection of texts you want to analyze (e.g., documents, articles, product descriptions).
  • {{technique}}: The preferred technique (e.g., TF-IDF, word embeddings, cosine similarity).
  • {{clustering_method}}: The clustering algorithm (e.g., K-means, hierarchical clustering) if clustering is needed.
  • {{use_case}}: The specific application (e.g., recommendation, search, grouping) to tailor the approach.

Instructions

  1. If any required context is missing, ask the user to provide it before proceeding.
  2. Based on the use case, design a step-by-step approach to preprocess the texts (e.g., tokenization, stop-word removal, stemming).
  3. Explain how to extract features using the specified technique, including code snippets or pseudocode.
  4. For similarity tasks, detail how to calculate similarity scores and interpret them.
  5. For clustering tasks, describe how to apply the chosen algorithm, determine the optimal number of clusters, and evaluate the results.
  6. If applicable, show how to build a recommendation or search system using the similarity scores.

Output format Provide a structured guide with clear sections: preprocessing, feature extraction, algorithm implementation, and evaluation. Include code examples in Python where relevant. Use a technical but accessible tone.

Guardrails Do not invent data or results; base all examples on the provided texts. Flag any assumptions about the data (e.g., language, domain). Stay within the scope of the specified use case.

Example "Texts: customer reviews of electronics; Technique: TF-IDF; Clustering: K-means; Use case: group similar complaints for product improvement."

Follow-up prompts

  • How can I tune the similarity threshold for better precision and recall?
  • What are the trade-offs between using TF-IDF and word embeddings for this dataset?
  • Can you suggest ways to handle very large text collections efficiently?