Course overview
Lesson 4 of 9 · 3 promptsAI for Product Analysts
LESSON 04 OF 9

A/B Test Planning

3 prompts for Product Analysts

Prompts for Product Analysts: copy one, fill it in, paste it into your AI.

Track progress as a member

In this lesson

  1. 01Draft A/B Test HypothesisUse this when you want to turn a product change idea into a testable experiment statement.
  2. 02Estimate A/B Test Sample SizeUse this when you need a quick estimate of how many users and days an A/B test requires.
  3. 03Define A/B Test Success CriteriaUse this when you must decide what result will make you ship, iterate, or stop a test.
1Copy the promptClick Copy on the prompt you need.
2Paste it into your AIChatGPT, Claude, Gemini or Copilot.
3Fill in the {{brackets}}Your own details, or let the AI ask you.
4Follow up and checkUse the follow-ups, then check the facts.
01

Draft A/B Test Hypothesis

Use this when you want to turn a product change idea into a testable experiment statement.

Prompt

Role - You turn rough product ideas into clear, testable A/B test hypotheses that a product team can review and run. Optimise for precision and shared understanding.

Context you provide

  • {{product_area}} - page, feature, or flow
  • {{change_idea}} - the specific change
  • {{primary_metric}} - success metric
  • {{baseline_rate}} - current value, if known
  • {{expected_direction}} - increase, decrease, or no change
  • {{target_segment}} - who sees the change
  • {{reasoning}} - why it should move the metric

Instructions

  1. Ask for any missing inputs, then draft the hypothesis.
  2. Write one if/then/because sentence naming the change, primary metric, expected direction, segment, and reason.
  3. State the primary metric and how it is measured.
  4. List one guardrail metric to watch for unintended harm.
  5. Note the target segment and any exclusions.
  6. Add a one-sentence success criterion.
  7. Flag assumptions and any inputs needing confirmation.

Output format Return a short markdown block: Hypothesis (one sentence), Primary metric, Guardrails (bullet list), Segment, Success criterion, Assumptions. Keep under 120 words. Use plain language. Leave out statistical formulas, sample size maths, tool setup, and implementation steps.

Guardrails

  • Do not invent baseline rates, lift percentages, or sample sizes. If a number is missing, ask.
  • Flag every assumption clearly.
  • Tell the user to check with a data scientist or experimentation platform before finalising sample size or test duration, and confirm privacy or consent requirements with a qualified professional.

Example Product area: Checkout page. Change idea: Add express pay button. Primary metric: Checkout completion rate. Baseline rate: 62%. Expected direction: Increase. Target segment: Returning mobile users. Reasoning: Fewer form fields reduce friction.

Open as its own page

02

Estimate A/B Test Sample Size

Use this when you need a quick estimate of how many users and days an A/B test requires.

Prompt

Role — You are a product analytics partner who turns a test hypothesis into a defensible sample size and duration estimate, optimising for a plan the team can schedule and trust.

Context you provide

  • Metric and its type (conversion rate or continuous mean): {{metric_and_type}}
  • Baseline value for that metric: {{baseline_value}}
  • Smallest effect worth detecting: {{minimum_detectable_effect}}
  • Number of variants including control: {{variant_count}}
  • Significance level and power targets: {{alpha_and_power}}
  • Eligible daily traffic (users or sessions): {{daily_eligible_traffic}}
  • Split, ramp and holdout constraints: {{traffic_split_and_ramp}}
  • Known seasonality or a fixed deadline: {{seasonality_or_deadline}}

Instructions

  1. Ask for any missing inputs above before calculating. Do not guess values.
  2. Confirm the metric type, whether the test is one or two sided, and whether the effect is absolute or relative.
  3. Compute the required sample size per variant and in total using the standard formula for that metric type. Show the formula and every substituted value.
  4. Convert to calendar days using eligible daily traffic and the split. Round up to whole users and whole days.
  5. Note the minimum runtime needed to cover one full business cycle if the traffic pattern looks weekly.
  6. Add a short sensitivity view showing how the day count moves if the detectable effect is larger or smaller.
  7. List what would invalidate the estimate, such as overlapping tests or a mid-test change.

Output format — A short table of inputs, then the per-variant and total sample size, then the day estimate, then the sensitivity view. Plain language, no jargon without a one-line definition. Under 400 words. Leave out significance-testing theory and tool-specific code unless asked.

Guardrails

  • Use only the user's figures. Label every assumption and never invent baselines, traffic volumes or benchmark numbers.
  • State that the estimate assumes random assignment and independent users, and that real results may need a longer run.
  • Tell the user to confirm the significance level, power and effect size with their product manager or data science lead before committing to a launch date.

Example — Metric: signup conversion, baseline 4.2%, MDE 0.5 percentage points, 2 variants, 80% power, 12,000 eligible users per day, 50/50 split, no deadline.

Open as its own page

03

Define A/B Test Success Criteria

Use this when you must decide what result will make you ship, iterate, or stop a test.

Prompt

Role You are a product analytics partner who turns experiment goals into pre-registered success criteria for a ship, iterate, or stop decision.

Context you provide

  • {{experiment_name}}: test name.
  • {{primary_metric}}: single success metric.
  • {{baseline_value}}: current metric value.
  • {{minimum_detectable_effect}}: smallest change worth acting on.
  • {{traffic_split}}: control and treatment split.
  • {{planned_duration}}: test length.
  • {{guardrail_metrics}}: metrics that must not worsen.
  • {{business_objective}}: desired business outcome.
  • {{decision_makers}}: who approves the decision.

Instructions

  1. Ask for any missing inputs, then restate the test in one sentence.
  2. Confirm how the primary metric is calculated.
  3. Ask what minimum detectable effect is meaningful.
  4. Define ship criteria: primary metric improves by at least the MDE and guardrails stay acceptable.
  5. Define iterate criteria: movement is positive but misses the MDE, or a guardrail issue is fixable.
  6. Define stop criteria: primary metric is flat or negative, or a guardrail breaches its limit.
  7. Note segments or qualitative data that would explain a borderline result.
  8. Produce a one-page decision brief.

Output format A one-page decision brief with sections: Test summary, Primary metric, Ship if, Iterate if, Stop if, Guardrails, Open questions. Use plain language and bullets. Maximum 300 words. Leave out implementation details and formulas unless requested.

Guardrails Do not invent baseline rates, MDEs, or thresholds. If an input is missing, state the assumption and ask the user to confirm. If the user cannot provide an MDE or guardrail limits, tell them to consult a data scientist or analytics lead before launch. Remind the user that experiment consent and data privacy rules must be checked with a legal or privacy advisor.

Example Experiment name: Onboarding checklist v2; primary metric: 7-day activation rate; baseline: 28%; MDE: 2 percentage points; traffic split: 50/50; duration: 2 weeks; guardrails: support ticket rate, refund rate; business objective: lift paid conversion; decision makers: PM, growth lead.

Open as its own page

Skills for these tasks

Give your AI these skills and it does these tasks the expert way. Connect your AI once and it picks them up by itself.