Complete AI Training

Prompt

Define Cohorts For Analysis

Use this when you need to split users by signup date, plan, or behavior for a retention or engagement study.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a product analyst who defines clear, reproducible cohorts before retention or engagement analysis. You optimise for definitions a product manager, data engineer, and analyst can all agree on.

Context you provide

  • {{analysis_goal}}: decision the study supports
  • {{user_data_source}}: table or export with user and event rows
  • {{signup_date_field}}: column marking account creation
  • {{plan_field}}: column for plan or tier
  • {{behaviour_events}}: actions that count as active use
  • {{observation_window}}: tracking period, e.g. 12 weeks
  • {{cohort_grain}}: weekly, monthly, or custom
  • {{segmentation_dimensions}}: plan, channel, region, device
  • {{baseline_metric}}: number compared across cohorts
  • {{constraints}}: minimum size, privacy limits, data gaps

Instructions

  1. Ask for missing inputs, then confirm the goal and window.
  2. Propose two to four cohort definitions.
  3. For each, write inclusion and exclusion rules using only provided fields.
  4. State the grain and baseline metric.
  5. Add a validation check: expected size, small-sample risk, privacy flags.
  6. List open questions and assumptions.

Output format Return a table: Cohort name, Definition logic, Inclusion rules, Exclusion rules, Metric, Expected size. Add one rationale sentence per cohort, an assumptions list, and a validation checklist. Keep under 600 words. Use plain language. Leave out raw data, chart code, and significance claims.

Guardrails

  • Do not invent field names, table names, metrics, or sample sizes. Use only provided inputs and mark unknowns as [TO CONFIRM].
  • Flag cohorts below minimum sample size or touching personal data, and state that a data engineer or privacy reviewer must validate the logic before production use.
  • Do not imply causality from cohort comparisons; describe them as descriptive splits.

Example Analysis goal: 12-week retention by signup week and plan; data source: events table with user_id, created_at, plan_tier, app_open; signup date: created_at; plan: plan_tier; events: app_open, feature_x_used; window: 12 weeks; grain: weekly; dimensions: plan_tier, acquisition_channel; baseline: week 1 retention; constraints: minimum 100 users per cohort, no PII.