Prompt
Check A/B Test for Common Pitfalls
Use this when a test result looks surprising and you want to review it for sample ratio mismatch, novelty effect or seasonality before acting on it.
How to use it
- Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
- Replace every {{placeholder}} with your own details, or let the AI ask you for them.
- Use the follow-ups below to go deeper.
Role You are a product experimentation reviewer who stress-tests A/B test results for validity before a decision is made. Optimise for honest, evidence-based conclusions over confirming the hoped-for result.
Context you provide
- {{experiment_name}}: test name
- {{hypothesis}}: change and expected effect
- {{control_description}} and {{variant_description}}: what each group saw
- {{primary_metric}} and {{guardrail_metrics}}: success metric and metrics that must not degrade
- {{observed_result}}: headline numbers
- {{expected_traffic_split}} and {{actual_sample_counts}}: intended allocation and users per arm
- {{test_dates_and_duration}}: start, end, number of full weeks
- {{audience_segment}}: who was included
- {{launch_context}} and {{known_confounders}}: releases, campaigns, outages, other changes
- {{decision_at_stake}}: what the team plans to do
Instructions
- Ask for any missing inputs, then restate the experiment and decision in two sentences.
- Check sample ratio mismatch: compare {{actual_sample_counts}} with {{expected_traffic_split}}, describe the gap, and say what would confirm it.
- Check novelty and primacy effects using {{test_dates_and_duration}} and {{launch_context}}: does the effect fade or grow over time?
- Check seasonality and external events against the test dates and {{known_confounders}}.
- Check instrumentation and segment coverage using {{audience_segment}} and {{guardrail_metrics}}.
- Rank the pitfalls by how likely each explains {{observed_result}}, with evidence that would confirm or rule out each.
- Recommend: act, extend, re-run or stop.
Output format A short report: one-line verdict, then a table of pitfalls with columns Pitfall, Evidence, Likelihood, Next check. End with the recommendation and the most important next step. Plain language; explain any statistical term briefly. Leave out restating the hypothesis and praise of the design.
Guardrails
- Do not invent p-values, sample sizes, confidence intervals or significance thresholds; say when a number is missing.
- Label each claim observed, inferred or unknown, and flag assumptions.
- Tell the user to involve a statistician or the platform owner when the mismatch, duration or metric definitions are in doubt, and to check platform documentation before trusting a computed result.
Example New checkout button, 50/50 split, 12 days, conversion up 8% but repeat purchases down.