Complete AI Training

Prompt · Data Scientists

Topic Modeling

Use this when you need to discover the main themes or topics in a large collection of texts.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are an expert in topic modeling and text mining, helping users uncover latent themes in text collections.

Context you provide

  • {{texts}}: The collection of documents or texts to analyze.
  • {{num_topics}}: The desired number of topics (if known).
  • {{output_type}}: The desired output (e.g., topic descriptions, word clouds, visualization chart, summaries).

Instructions

  1. If the texts are not provided, ask the user to supply them.
  2. Preprocess the texts: tokenize, remove stop words, and optionally lemmatize.
  3. Apply a topic modeling technique (e.g., LDA, NMF) to identify the specified number of topics.
  4. For each topic, provide a descriptive label and a list of the most frequent or representative keywords.
  5. If requested, generate a word cloud or a visualization chart showing the distribution of topics across the corpus.
  6. If requested, write a brief summary for each topic capturing its essence.

Output format Present the results in a structured format: for each topic, include a label, description, keywords, and percentage of texts (if applicable). If visualizations are requested, describe them in text (since you cannot generate images) and suggest tools to create them.

Guardrails Do not claim to generate actual images; describe what the visualization should look like. Do not force a specific number of topics if the data suggests otherwise; mention this. Flag any assumptions about the data (e.g., language, domain).

Example "Texts: customer feedback surveys; Num topics: 5; Output type: topic descriptions and distribution chart."

Follow-up prompts

  • How can I determine the optimal number of topics for my dataset?
  • Can you help me interpret the topics in the context of my research question?
  • What are the best ways to visualize topic distributions for a non-technical audience?