Skip to main content
Statistical Analysis

Stop Peeking: How to Run Valid A/B Test Statistics

A practical walkthrough for analysts and PMs: how to avoid peeking, set sample size, and interpret p-values correctly in A/B tests.

You can stop an A/B test the moment your dashboard shows p

1. Know what you're actually testing

Every A/B test compares a control (the current version) against a variation (the new version) on a measurable outcome like conversion rate or revenue (Spotify Confidence). The null hypothesis assumes no real difference between the two, and any observed gap is just chance (Spotify Confidence). Your job is to gather enough evidence to reject that null. Hypothesis testing follows five steps: state the hypotheses, choose the significance level, compute the test statistic, find the p-value, and interpret the results (Stat Trek). Skipping steps is how you end up fooling yourself.

2. Commit to a sample size before you start

Decide on your sample size in advance. That's the core defense against peeking. Evan Miller's rule of thumb is n = 16 × (σ² / δ²), where δ is the minimum effect you want to detect and σ² is the expected variance (Evan Miller). For a conversion rate, variance is p(1-p). Say your baseline conversion is 10%. You want to detect a lift to 12% (δ = 0.02). Variance is 0.1 × 0.9 = 0.09. Then n = 16 × (0.09 / 0.0004) = 3,600 per variation. That's your target. Don't stop before you hit it.

You'll also need to set power. Before a test, the desired power is usually 80%, along with the minimum detectable effect (MDE) (Optimizely). Power is the probability of correctly rejecting a false null (Stat Trek). Power increases with larger sample size and larger effect size, and decreases with more variability or a lower significance level (Stat Trek). So if you're chasing a tiny effect in a noisy metric, you'll need a massive sample.

3. Don't peek — or use sequential methods if you must

Repeated significance testing always increases false positives (Evan Miller). If you peek ten times, what you think is 1% significance is actually 5% (Evan Miller). In the worst case, where you stop as soon as 5% significance appears and test after every observation, the actual false-positive rate can reach 26.1% (Evan Miller). That's not a test; that's a random number generator with a dashboard.

You have two options. First, pre-commit to a fixed end date and don't look at results until then. Second, use sequential testing methods that adjust thresholds for multiple looks (Optimizely). Sequential sampling lets you stop early if the treatment looks like a winner, addressing the peeking problem (Evan Miller). Evan Miller's simple sequential procedure: choose a sample size N, assign subjects 50/50, and stop when the treatment-minus-control success count reaches 2 × √N (declaring a winner) or the total success count reaches N (declaring no winner) (Evan Miller). In some circumstances, this can reduce the number of observations needed by 50% or more (Evan Miller).

4. Interpret p-values and confidence intervals correctly

A p-value is the probability of observing data as extreme as what you saw, assuming the treatment has no effect. A p-value of 0.03 means such a result would occur by chance 3% of the time (Optimizely). It does not mean the treatment has a 97% chance of working (Optimizely). That's a subtle but critical distinction.

Results are typically declared statistically significant when p

Use confidence intervals. A 95% confidence interval is a range that, if the experiment were repeated many times, would contain the true effect 95% of the time (Optimizely). If your interval for relative lift includes zero, you don't have a clear winner.

5. Watch for hidden traps: SRM, novelty, and multiple comparisons

Even with perfect peeking discipline, things go wrong. A Sample Ratio Mismatch (SRM) is when the observed sample ratio differs from the expected ratio; it's a symptom of data quality issues (KDD 2019). Ignoring an SRM without knowing its root cause may result in a bad product modification appearing to be good and being shipped (KDD 2019). Run A/A tests quarterly to validate your setup (Optimizely). If an A/A test reports a statistically significant difference between two identical versions, check your implementation (Optimizely).

Novelty and primacy effects can mislead you. Novelty is the desire to use new technology that diminishes over time; primacy is growing engagement as adoption increases (Sadeghi et al.). A test that shows a big lift in week one might fade. Run longer if you suspect either effect.

Multiple comparisons are another trap. When paired with classical statistics, peeking and testing many goals and variations at once can increase the chance of incorrectly declaring a winner by over 5x (Optimizely Stats Engine). Optimizely re-analyzed 48,000 historical experiments and found roughly 20% fewer variations with a false discovery rate below 0.1 than with a false positive rate at the same level (Optimizely Stats Engine). So limit your metrics. Decide on one primary metric and a few secondary ones.

6. What can go wrong if you ignore all this

You'll ship losers. At Bing, less than a third of ideas tested move the metrics they were designed to improve (Bing). And it's not uncommon for Bing experiments to unintentionally impact revenue positively or negatively by 1% (Bing). At Microsoft, a 1% revenue improvement equals about $10 million annually in the US (Kohavi et al.). So a false positive isn't just embarrassing; it can cost real money.

You might also miss real wins. A Type II error is a false negative — missing a real improvement (Spotify Confidence). If your test is underpowered, you'll fail to detect a genuine effect. That's why you set power at 80% and use variance reduction techniques like CUPED. CUPED adjusts metrics using pre-experiment data to reduce variance while keeping the estimator unbiased (Deng et al.). At Bing, three experiments using CUPED showed variance reductions of 45%, 52%, and 49% (Deng et al.). Reducing variance by about 50% effectively achieves the same power with half the users or half the duration (Deng et al.).

Sources

  • Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Evan Miller - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Stat Trek - https://stattrek.com/hypothesis-test/power-of-test
  • Nielsen Norman Group - https://www.nngroup.com/articles/practical-significance/
  • KDD 2019 - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
  • Deng et al. - https://exp-platform.com/Documents/2013-02-CUPED-ImprovingSensitivityOfControlledExperiments.pdf

Share this article:

Comments (0)

No comments yet. Be the first to comment!