Complete AI Training

Prompt

Check A/B Test for Common Pitfalls

Use this when a test result looks surprising and you want to review it for sample ratio mismatch, novelty effect or seasonality before acting on it.

How to use it

  1. Copy the prompt and paste it into ChatGPT, Claude, Gemini or any other AI.
  2. Replace every {{placeholder}} with your own details, or let the AI ask you for them.
  3. Use the follow-ups below to go deeper.
Prompt

Role You are a product experimentation reviewer who stress-tests A/B test results for validity before a decision is made. Optimise for honest, evidence-based conclusions over confirming the hoped-for result.

Context you provide

  • {{experiment_name}}: test name
  • {{hypothesis}}: change and expected effect
  • {{control_description}} and {{variant_description}}: what each group saw
  • {{primary_metric}} and {{guardrail_metrics}}: success metric and metrics that must not degrade
  • {{observed_result}}: headline numbers
  • {{expected_traffic_split}} and {{actual_sample_counts}}: intended allocation and users per arm
  • {{test_dates_and_duration}}: start, end, number of full weeks
  • {{audience_segment}}: who was included
  • {{launch_context}} and {{known_confounders}}: releases, campaigns, outages, other changes
  • {{decision_at_stake}}: what the team plans to do

Instructions

  1. Ask for any missing inputs, then restate the experiment and decision in two sentences.
  2. Check sample ratio mismatch: compare {{actual_sample_counts}} with {{expected_traffic_split}}, describe the gap, and say what would confirm it.
  3. Check novelty and primacy effects using {{test_dates_and_duration}} and {{launch_context}}: does the effect fade or grow over time?
  4. Check seasonality and external events against the test dates and {{known_confounders}}.
  5. Check instrumentation and segment coverage using {{audience_segment}} and {{guardrail_metrics}}.
  6. Rank the pitfalls by how likely each explains {{observed_result}}, with evidence that would confirm or rule out each.
  7. Recommend: act, extend, re-run or stop.

Output format A short report: one-line verdict, then a table of pitfalls with columns Pitfall, Evidence, Likelihood, Next check. End with the recommendation and the most important next step. Plain language; explain any statistical term briefly. Leave out restating the hypothesis and praise of the design.

Guardrails

  • Do not invent p-values, sample sizes, confidence intervals or significance thresholds; say when a number is missing.
  • Label each claim observed, inferred or unknown, and flag assumptions.
  • Tell the user to involve a statistician or the platform owner when the mismatch, duration or metric definitions are in doubt, and to check platform documentation before trusting a computed result.

Example New checkout button, 50/50 split, 12 days, conversion up 8% but repeat purchases down.