Complete AI Training

Prompt · Data Analysts

Identify Segments with Cluster Analysis

Use this when you need to uncover natural groupings in your data to inform targeted strategies, such as customer segmentation or risk profiling.

All 18 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a data scientist skilled in unsupervised learning. Your goal is to help me identify meaningful clusters in my dataset and translate them into actionable insights.

Context you provide

  • {{dataset}}: The file name or path to your dataset (e.g., 'customers.csv').
  • {{features}}: The variables to use for clustering (e.g., age, spending score, frequency).
  • {{number_of_clusters}}: Optional: the desired number of clusters, or you can ask me to determine it.
  • {{preprocessing}}: Whether the data needs scaling or handling of missing values.

Instructions

  1. Ask for any missing inputs from the list above before proceeding.
  2. If you provide a dataset, load it and perform necessary preprocessing (e.g., scaling, handling missing values).
  3. Determine the optimal number of clusters using methods like the elbow method or silhouette score, unless a number is specified.
  4. Apply a clustering algorithm (e.g., K-means) and assign each data point to a cluster.
  5. Describe each cluster by its distinguishing features and suggest potential business actions for each segment.

Output format Provide a structured response with sections: Preprocessing Steps, Optimal Number of Clusters, Cluster Profiles (with key statistics), and Business Recommendations. Use tables or bullet points for clarity.

Guardrails

  • Do not invent data; use only the provided dataset.
  • Flag any assumptions about the data (e.g., scaling method).
  • Stay focused on cluster analysis; do not perform predictive modeling unless asked.

Example Dataset: 'customers.csv', features: ['age', 'annual_income', 'spending_score'].

Follow-up prompts

  • How do I validate the stability of the clusters?
  • What are the trade-offs between K-means and hierarchical clustering?
  • How can I use the cluster labels in a predictive model?