Skip to main content
Statistical Analysis

Stop Peeking: Why Your A/B Test Results Are a Lie

Peeking at A/B test results inflates false positives to 30%. Pre-commit to sample size or use sequential testing. Here's the fix.

Thirty percent. That's the chance that your "statistically significant" A/B test result is actually a false positive if you've been peeking at the data daily and stopping the moment significance appears (Optimizely). Most teams don't realize they're baking this lie into their roadmap. The fix isn't more discipline—it's a different statistical framework.

The Peeking Problem

Here's the dirty secret of classic A/B testing: every time you look at your results, you roll the dice. Run a test, check it on day three, see a win, and stop? You've just inflated your false-positive rate from the nominal 5% to as high as 30% (Optimizely). That's not a rounding error—that's a systemic failure.

The math is brutal. If you peek ten times, your 1% significance level is actually just 5% (Evan Miller). In the worst case, stopping as soon as 5% significance appears—when you check after every observation—can push the real false-positive rate to 26.1% (Evan Miller). This isn't hypothetical. It's what happens on every dashboard that updates in real time.

Why does this happen? Because the p-value—that sacred 0.05 threshold—assumes you decided in advance how many observations you'd collect. Peeking breaks that assumption. The more you look, the more likely you are to stumble on a random fluctuation that looks like a winner.

Pre-Commit or Go Sequential

The classic fix is to pre-commit to a fixed sample size. You decide before the test starts how many visitors you need, calculate it based on your baseline conversion rate and minimum detectable effect (Optimizely sample size guide), and then you don't peek until that number arrives. That works—if you have the willpower of a monk. Most teams don't.

The better fix is sequential testing. Sequential methods let you peek as often as you like without inflating false positives. They're designed for continuous monitoring. Optimizely's Stats Engine, built with Stanford statisticians, uses sequential testing and can call winners up to 2.5 times as fast for experiments over 50,000 visitors (Optimizely Stats Engine). Netflix uses sequential canary tests to detect streaming regressions like PlayDelay, flagging severe issues while strictly controlling false detection (arXiv - Netflix canary testing).

Evan Miller's simple sequential procedure is even easier: pick a sample size N, assign visitors 50/50, and stop when the treatment-minus-control success count hits 2√N (winner) or total successes hit N (no winner). It can cut required observations by 50% or more, especially with low conversion rates (Evan Miller - Simple Sequential A/B Testing). No p-hacking, no peeking penalty.

The Counter-Argument: Just Use a Bigger Sample

"But we have millions of visitors," you say. "We can afford a huge sample size and just wait." Sure, you can wait—but that doesn't fix peeking. Even with massive samples, if you check results weekly and stop early, you're still inflating false positives. And there's a subtler trap: with very large samples, statistical significance becomes trivial. A microscopic 0.03% improvement in form-field errors can yield p

So bigger samples don't solve the peeking problem; they just make it more expensive. Sequential testing solves it directly, and it also handles the multiple-comparisons issue—testing many metrics at once. In a re-analysis of 48,000 historical experiments, Optimizely found that fixed-horizon statistics declared a winner or loser in 36% of tests, while sequential Stats Engine declared in only 22%—uncovering 39% fewer conclusive results (Optimizely Stats Engine). That's not a bug; it's the honesty you need.

What I'd Actually Do

Stop using fixed-horizon p-values for your everyday A/B tests. Switch to a sequential method—either Optimizely's Stats Engine, Evan Miller's simple procedure, or a multi-armed bandit if you want to reallocate traffic dynamically (VWO). Pre-commit to a minimum detectable effect and a significance level, but let the sequential method tell you when to stop, not your gut.

Here's my concrete recommendation: if you're running a test on a conversion rate around 5%, use Evan Miller's sample size formula: n = 16 × (σ² / δ²), where σ² = p(1-p) and δ is your minimum effect. For a 5% baseline and a 1% absolute effect (δ = 0.01), that's roughly 16 × (0.05 × 0.95 / 0.0001) = 76,000 per variation. That's a lot. But with sequential testing, you might stop earlier. And if you're using Optimizely, just enable Stats Engine—it's designed for this.

One more warning: run A/A tests quarterly to check your setup (Optimizely A/A testing glossary). If your A/A test shows significance above 95%, your implementation is broken. Fix that before you trust any result.

Stop peeking. Start using sequential methods. Your roadmap will thank you.

Sources

  • Optimizely (A/B testing) - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Evan Miller - Simple Sequential A/B Testing - https://www.evanmiller.org/sequential-ab-testing.html
  • Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
  • Optimizely sample size guide - https://www.optimizely.com/insights/blog/how-to-calculate-sample-size-of-ab-tests/
  • Nielsen Norman Group - Practical Significance - https://www.nngroup.com/articles/practical-significance/

Share this article:

Comments (0)

No comments yet. Be the first to comment!