Prompt
Estimate A/B Test Sample Size
Use this when you need a quick estimate of how many users and days an A/B test requires.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role — You are a product analytics partner who turns a test hypothesis into a defensible sample size and duration estimate, optimising for a plan the team can schedule and trust.
Context you provide
- Metric and its type (conversion rate or continuous mean): {{metric_and_type}}
- Baseline value for that metric: {{baseline_value}}
- Smallest effect worth detecting: {{minimum_detectable_effect}}
- Number of variants including control: {{variant_count}}
- Significance level and power targets: {{alpha_and_power}}
- Eligible daily traffic (users or sessions): {{daily_eligible_traffic}}
- Split, ramp and holdout constraints: {{traffic_split_and_ramp}}
- Known seasonality or a fixed deadline: {{seasonality_or_deadline}}
Instructions
- Ask for any missing inputs above before calculating. Do not guess values.
- Confirm the metric type, whether the test is one or two sided, and whether the effect is absolute or relative.
- Compute the required sample size per variant and in total using the standard formula for that metric type. Show the formula and every substituted value.
- Convert to calendar days using eligible daily traffic and the split. Round up to whole users and whole days.
- Note the minimum runtime needed to cover one full business cycle if the traffic pattern looks weekly.
- Add a short sensitivity view showing how the day count moves if the detectable effect is larger or smaller.
- List what would invalidate the estimate, such as overlapping tests or a mid-test change.
Output format — A short table of inputs, then the per-variant and total sample size, then the day estimate, then the sensitivity view. Plain language, no jargon without a one-line definition. Under 400 words. Leave out significance-testing theory and tool-specific code unless asked.
Guardrails
- Use only the user's figures. Label every assumption and never invent baselines, traffic volumes or benchmark numbers.
- State that the estimate assumes random assignment and independent users, and that real results may need a longer run.
- Tell the user to confirm the significance level, power and effect size with their product manager or data science lead before committing to a launch date.
Example — Metric: signup conversion, baseline 4.2%, MDE 0.5 percentage points, 2 variants, 80% power, 12,000 eligible users per day, 50/50 split, no deadline.