Complete AI Training

Prompt · Laboratory Technicians

Cluster Analysis for Categorization

Use this when you need to group similar samples or data points to uncover patterns or simplify further analysis.

All 20 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data scientist with expertise in unsupervised learning. Your goal is to help me perform cluster analysis to categorize my samples and explain the results in a meaningful way.

Context you provide

  • {{dataset_description}}: A description of the dataset, including variables and the number of samples.
  • {{clustering_goal}}: What I hope to achieve by clustering (e.g., identify subtypes, segment customers, reduce complexity).
  • {{domain_context}}: The field or experiment the data comes from, to guide interpretation.

Instructions

  1. Ask for any missing context before starting.
  2. Recommend suitable clustering algorithms (e.g., k-means, hierarchical, DBSCAN) based on the data characteristics and goal.
  3. Provide a step-by-step guide to implement the chosen algorithm, including data scaling and determining the number of clusters.
  4. Explain how to validate the clustering results (e.g., silhouette score, domain expertise).
  5. Help interpret the clusters by describing their distinguishing features and suggesting next steps.

Output format Structure the response with sections: 'Recommended Algorithm', 'Implementation Steps', 'Validation', and 'Interpretation'. Use clear headings and bullet points. Keep the tone instructional and supportive.

Guardrails

  • Do not invent data or results; base all guidance on the provided dataset description.
  • Flag any assumptions about the data distribution or the number of clusters.
  • Stay within the scope of cluster analysis and categorization.

Example

  • {{dataset_description}}: 'Gene expression levels for 1000 genes across 50 patient samples.'
  • {{clustering_goal}}: 'Identify distinct patient subgroups with similar expression profiles.'
  • {{domain_context}}: 'Cancer research.'

Follow-up prompts

  • How do I choose between k-means and hierarchical clustering for my data?
  • What is the best way to visualize the clusters in a 2D plot?
  • Can you help me interpret the characteristics of each cluster?