We've all done it: you launch an A/B test, wait a few hours, and then refresh the dashboard. The conversion rate for the variation is up 10% and the p-value is 0.03. You're ready to declare victory and ship it. But here's the hard truth: if you stopped your test the moment you saw significance, your result is likely a lie. The common misconception is that you can check your results anytime and trust what you see. That's wrong. Peeking at your data and stopping early is the fastest way to turn a 5% false-positive rate into a 30% one (Optimizely).
The Fixed-Horizon Trap
Classical A/B testing is built on a fixed-horizon design: you decide your sample size in advance, run the experiment, and analyze the results once. This design assumes you'll only look at the data one time, at the end. But we're human. We're curious. We peek. And every peek you take at an ongoing experiment increases your chance of a false positive (Evan Miller). In fact, if you peek ten times, what you think is 1% significance is actually just 5% significance (Evan Miller). In the worst case, if you stop as soon as 5% significance appears and check after every observation, your actual false-positive rate can reach 26.1% (Evan Miller). That's not a rare edge case; that's a systemic problem.
The root of the issue is that a p-value is a conditional probability: it tells you the probability of seeing data as extreme as what you saw, assuming the treatment has no effect (Optimizely). It does not tell you the probability that the treatment works. And when you peek, you're effectively running multiple tests on the same data, which inflates the chance that you'll find a significant result by pure luck. This is why Optimizely's own data shows that combining peeking with testing many goals and variations can increase the chance of incorrectly declaring a winner by over 5x (Optimizely Stats Engine).
The Fix: Pre-Commit or Go Sequential
The solution is simple: stop peeking. Pre-commit to a fixed end date or a fixed sample size before you start. This is the classic remedy (Evan Miller). But if you can't resist glancing at the data, then design your test for peeking. Use sequential testing methods that adjust your thresholds for multiple looks. These methods are designed to let you monitor your experiment continuously without inflating your false-positive rate. Optimizely's Stats Engine, for example, is built on sequential testing and allows users to see results that are always valid any time they peek (Optimizely Stats Engine). The team at Stanford that helped develop it showed that sequential tests can call a winner or loser up to 2.5 times as fast as fixed-horizon statistics for experiments over 50,000 visitors (Optimizely Stats Engine).
Even better, sequential testing can free you from having to guess your minimum detectable effect (MDE) in advance. With fixed-horizon statistics, you need to specify an MDE to calculate sample size. But with sequential methods, you can just let the data accumulate and stop when you have enough evidence (Optimizely Stats Engine). This is a huge advantage in practice, because we're often wrong about effect sizes. At Bing, for example, less than a third of ideas move the metrics they were designed to improve (Bing - Large Scale Experimentation at Bing). So why pretend we know the exact effect size ahead of time?
The Counter-Argument: Sample Size Discipline
Now, some statisticians will argue that the real problem isn't peeking; it's that people don't calculate sample sizes at all. They'll say, 'If you just set your sample size and stick to it, you'll be fine.' And there's truth to that. A rule-of-thumb sample size formula is n = 16 times (sigma-squared / delta-squared), where delta is the minimum effect you wish to detect and sigma-squared is the expected sample variance (Evan Miller). For a binomial conversion rate, variance is p(1-p). If you follow that, you'll have 80% power to detect your effect at a 5% significance level, which is the standard (Optimizely).
But here's the problem: sample size calculations assume you're going to run the test for exactly that many visitors and then stop. If you're human, you'll peek. And even if you don't peek, fixed-horizon tests are inefficient. They can't stop early even if the effect is huge. They have to run to the full sample size. That's a waste of time and money, especially when you're running hundreds of experiments at scale. At Bing, over 250 experiments run on a typical day, and almost every user is exposed to about 15 different experiments simultaneously (Bing - Large Scale Experimentation at Bing). That's a lot of experiments that could be shortened by sequential testing.
So while sample size discipline is important, it's not enough. The real defense against peeking is to change the statistical framework. Sequential testing, as championed by Evan Miller's simple sequential procedure, can reduce the number of observations required for a successful experiment by 50% or more, and it works extremely well with low conversion rates (Evan Miller). That's a practical benefit you can't ignore.
The Bigger Picture: Practical Significance
Even with perfect statistical design, you still need to ask: is the effect worth caring about? Statistical significance (p
So my recommendation is this: use sequential testing for your A/B tests, and always look at the confidence interval and the effect size, not just the p-value. Pre-commit to a stopping rule—either a fixed sample size or a sequential threshold—and then let the data speak. And when you see a significant result, ask yourself: is this effect big enough to matter? If not, move on.
Bottom line
The single best move you can make in A/B test design is to stop peeking and switch to sequential testing. It will save you time, reduce false positives, and give you honest results you can act on. Pre-commit to a stopping rule, or use a tool that does it for you, and you'll sleep better knowing your experiment is telling the truth.
Sources
- Optimizely - A/B testing: https://www.optimizely.com/optimization-glossary/ab-testing/
- Optimizely Stats Engine: https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
- Evan Miller - How Not to Run an A/B Test: https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Evan Miller - Simple Sequential A/B Testing: https://www.evanmiller.org/sequential-ab-testing.html
- Nielsen Norman Group - Practical Significance: https://www.nngroup.com/articles/practical-significance/
- Bing - Large Scale Experimentation at Bing: https://blogs.bing.com/search-quality-insights/August-2013/Large-Scale-Experimentation-at-Bing/
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!