If you peek at your A/B test daily and stop as soon as you see significance, you're not running an experiment—you're gambling. In fact, that habit can inflate your false-positive rate from the nominal 5% to as high as 30% (Optimizely). That means nearly one in three of your “winners” is actually a loser. The fix isn't more willpower; it's a better tool.
The Peeking Problem Is Real, and It's Worse Than You Think
Peeking—checking results before the test is complete—is the most common way experimenters fool themselves. The math is unforgiving: if you peek ten times, what you think is 1% significance is actually just 5% (Evan Miller). And in a worst-case scenario where you stop as soon as 5% significance appears, the true false-positive rate can reach 26.1% (Evan Miller). That's not a rounding error; that's a broken process.
Why does this happen? Because every time you look at the data and consider stopping, you're implicitly running a new hypothesis test. The more you look, the more chances you give chance to produce a misleading result. It's like rolling dice and deciding to stop when you hit double sixes—eventually you will, but that doesn't mean the dice are loaded.
Pre-Commit or Use Sequential Testing
The classic defense is to pre-commit to a fixed sample size and end date before you start. That works, but it's rigid. If the effect is larger than expected, you waste time; if it's smaller, you might miss it. A better approach is sequential testing, which lets you peek as often as you like without inflating false positives (Optimizely Stats Engine). Sequential methods like the Sequential Probability Ratio Test (SPRT) can also reduce required sample sizes (Optimizely sample size guide).
Evan Miller's simple sequential procedure is a great starting point: choose a sample size N, assign users 50/50, and stop when the difference in successes reaches 2√N (declare a winner) or total successes hit N (declare no winner). In some cases, this can cut the number of observations needed by 50% or more, especially with low conversion rates (Evan Miller). That's not just a statistical nicety—it means you can ship winning changes faster, or run more tests with the same traffic.
Bayesian Testing: Another Way to Skip the Peeking Penalty
If you're not ready for sequential methods, consider Bayesian testing. Unlike frequentist methods, Bayesian testing outputs direct probabilities like “there's a 92% chance Variant B is better” and allows continuous monitoring without inflating false positives (Optimizely). VWO reports that Bayesian approaches can deliver actionable results almost 50% faster while still focusing on statistical significance (VWO). That's a practical advantage when you're waiting days for a winner.
But be careful: Bayesian testing isn't a magic bullet. It still requires you to define priors and interpret results sensibly. And no statistical method can save you from a bad metric or a poorly designed test. The key is to pick a method that matches your workflow and stick to it.
The Counter-Argument: “But I Have Deadlines”
You might say, “I can't wait weeks for a result—my boss wants an answer by Friday.” I get it. But here's the thing: a rushed, peeked-at result is worse than no result, because you'll act on a false positive and potentially harm your business. A 1% revenue impact at Microsoft is about $10 million annually in the US (Kohavi et al. - Online Controlled Experiments at Large Scale). Do you really want to risk that on a test you stopped early because it looked good on Wednesday?
If you truly can't wait, then use sequential testing or a multi-armed bandit. Bandits dynamically allocate traffic to better-performing variations, which is a form of continuous optimization that doesn't suffer from peeking issues (VWO). They're not a replacement for A/B testing in all cases, but they're a legitimate tool when you need to make decisions quickly.
Bottom Line
Stop peeking. Pre-commit to a sample size, or better yet, switch to sequential testing. Your false positives will drop, your winners will be real, and you'll actually learn something from your experiments. The best move is to adopt a sequential testing method today—your future self (and your boss) will thank you.
Sources
- Optimizely - A/B testing glossary - https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
- VWO - Multi-Armed Bandit Testing - https://vwo.com/glossary/multi-armed-bandit-testing/
- Kohavi et al. - Online Controlled Experiments at Large Scale - https://exp-platform.com/large-scale/
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!