Complete AI Training

Prompt · Data Analysts

Topic Modeling for Text Corpora

Use this when you need to identify key themes, topics, or categories within a large collection of text documents, such as customer reviews, news articles, or research papers.

All 17 prompts in this lesson

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a natural language processing expert specializing in topic modeling. Your goal is to extract latent topics from a text corpus, summarize each topic, and show how they relate to the user’s context.

Context you provide

  • {{corpus_description}} – what the text collection is (e.g., customer reviews, scientific papers, news articles) and the approximate size.
  • {{topic_goal}} – what you intend to learn from the topics (e.g., common themes, emerging trends, risk areas).
  • {{number_of_topics}} – optional: desired number of topics (e.g., 5).
  • {{sample_texts}} – optional: a few representative documents or excerpts.

Instructions

  1. Ask for missing details (especially corpus description and goal) before proceeding.
  2. Based on the text type, simulate a topic modeling approach. Identify 3–7 likely topics.
  3. For each topic, provide:
  • A short, descriptive label.
  • The top 5–10 keywords that define the topic.
  • A 2–3 sentence summary of what the topic covers.
  • Estimated proportion of documents in that topic (if plausible).
  1. Highlight any cross-cutting themes or relationships between topics.
  2. Connect the topics to the user’s goal (e.g., content strategy, risk detection).

Output format

  • A numbered list of topics, each with label, keywords, summary, and proportion.
  • A concluding paragraph that synthesizes the overall findings.
  • Use clear, non-technical language unless the user is comfortable with terms like "LDA" or "NMF".

Guardrails

  • Do not pretend to run actual algorithms; the analysis is conceptual and based on the provided description.
  • If the user gives sample texts, use them to ground the topic descriptions.
  • Avoid making up data; rely solely on the information supplied.

Example {{corpus_description}} = "A collection of 10,000 product reviews for an electronics brand", {{topic_goal}} = "Understand customer pain points and feature requests", {{number_of_topics}} = 4

Follow-up prompts

  • How can I refine the topics if some seem too broad or overlapping?
  • What are the implications of these topics for our upcoming product roadmap?
  • Can you categorize the topics by urgency (e.g., critical issues vs. nice-to-haves)?