Skip to main content
Statistical Analysis

Peeking Is Killing Your A/B Tests: Stop It Now

Checking results early and stopping at significance is a trap. It inflates false positives dramatically. Pre-commit or use sequential testing. Here's the fix.

You're probably peeking at your A/B test results. And that's a problem. Most people check daily and stop as soon as they see significance. That's wrong. It can inflate your false-positive rate from 5% to as high as 30% (Optimizely).

This is the single biggest statistical mistake in A/B testing. It's not sample size, not p-value interpretation, not even multiple comparisons. It's peeking. And it's rampant.

The Myth of the 5% Significance Level

That 5% significance threshold only works if you look at the data once, at a pre-defined time. The moment you peek and decide whether to stop, you're running a different test. The math changes. Repeated significance testing always increases false positives (Evan Miller).

Let's be concrete. Say you peek at your test ten times. What you think is 1% significance is actually just 5% significance (Evan Miller). You're fooling yourself. And if you stop at the first moment you see 5% significance, your real false-positive rate could hit 26.1% (Evan Miller). That's not a margin of error—that's a coin flip.

Why Peeking Is So Tempting (and So Dangerous)

Peeking feels productive. You're eager to ship a winner. But you're trading a false sense of certainty for real risk. The more you look, the more likely you'll see a spurious result. It's not just about significance—it's about the decisions you make from it. You might roll out a change that does nothing, or worse, hurts your metrics.

At Bing, where over 250 experiments run on a typical day, the stakes are enormous (Bing). A single experiment can move revenue by 1%—that's $10 million annually at Microsoft's scale (Kohavi). You can't afford to make decisions on a flawed statistical foundation.

The Solution: Pre-Commit or Go Sequential

There are two ways out. First, pre-commit to a fixed sample size and end date. Decide before you start. Then don't look until it's over. This is the classic approach, and it works if you have the discipline.

Second, use sequential testing. Sequential methods let you peek at any time with always-valid results (Optimizely Stats Engine). They adjust for multiple looks. You don't need to guess a minimum detectable effect in advance. Optimizely's Stats Engine, built with Stanford statisticians, can call a winner up to 2.5 times faster for large experiments (Optimizely).

Evan Miller's simple sequential procedure is even more direct: pick a sample size N, assign 50/50, and stop when the treatment-minus-control count reaches 2 times the square root of N for a winner, or the total success count hits N for no winner. It can cut observations by 50% or more, especially with low conversion rates (Evan Miller).

What About the Counter-Argument: "But We Need Quick Answers"

You'll hear this: "We can't wait weeks for a result. We need to move fast." I get it. But peeking to stop early doesn't give you a quick answer—it gives you a wrong answer. You'll just have to re-run the test later. That's slower, not faster.

Sequential testing is the answer to speed. It's designed to let you stop early when a winner is clear, without inflating false positives (Evan Miller). Netflix uses sequential canary tests to catch regressions quickly while controlling false detections (arXiv - Netflix). So you can be fast and correct.

What I'd Actually Do

Here's my recommendation: Stop peeking. If you must peek, use a sequential method. In practice, I'd do this:

  • Run an A/A test quarterly to check your setup (Optimizely).
  • Choose a primary metric and pre-register your hypotheses.
  • Use CUPED to reduce variance—Bing saw 45-52% variance reduction, meaning you need half the users or half the time (Deng).
  • If you need flexibility, use a sequential testing tool like Optimizely's Stats Engine.

And remember, statistical significance isn't everything. A 0.2 percentage point difference can be significant (p = 0.03) but practically meaningless (Nielsen Norman). Focus on effect size and business impact.

Quick tip: If you peek and see significance, don't stop. Let the test run to its planned end. You'll sleep better.

Sources

  • Optimizely – https://www.optimizely.com/optimization-glossary/ab-testing/
  • Evan Miller – https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Optimizely Stats Engine – https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
  • Kohavi et al. – https://exp-platform.com/large-scale/
  • Deng et al. – https://exp-platform.com/Documents/2013-02-CUPED-ImprovingSensitivityOfControlledExperiments.pdf

Share this article:

Comments (0)

No comments yet. Be the first to comment!