Which A/B Testing Method Should You Trust When the Stakes Are High?
You've just launched a new checkout flow, and after two days, the variation is winning with 95% significance. Do you pull the trigger? If you're like most practitioners, you've felt that itch. But peeking and stopping early can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely). That's the classic pitfall of fixed-horizon testing. In this head-to-head, we compare fixed-horizon (the traditional approach) against sequential testing (the modern fix) using real-world case studies from Bing, Netflix, and Optimizely. By the end, you'll know which method to adopt depending on your team's risk tolerance and decision speed.
Method A: Fixed-Horizon Testing – The Old Guard
Fixed-horizon testing is what most of us learned: pick a sample size, run the test until you hit that number, then check significance. The theory is sound—if you pre-commit to a sample size, your p-values are valid. But in practice, teams rarely resist peeking. Evan Miller's analysis shows that if you peek ten times, what you think is 1% significance is actually just 5% (Evan Miller). And if you stop as soon as significance appears, the actual false-positive rate can reach 26.1% (Evan Miller).
Consider Bing's experience. At Bing, less than a third of experiments move the metrics they were designed to improve (Bing). With that low a hit rate, false positives are costly—you might ship a feature that actually hurts revenue. Bing runs over 250 experiments daily, with every user exposed to about 15 simultaneous experiments (Bing). That scale amplifies the risk of false positives, so they have to be careful. Their solution? They alert experimenters to statistical interactions between experiments (Kohavi et al., KDD 2013). But even with those alerts, the fixed-horizon method's weakness remains: it tempts you to peek.
Fixed-horizon testing isn't all bad. If you can commit to a sample size and never peek, it's straightforward. But how many teams have that discipline? The data suggests few do.
Method B: Sequential Testing – The Always-Valid Alternative
Sequential testing flips the script: you can peek anytime, and your results are always valid. The math behind it, such as always-valid p-values, allows continuous monitoring without inflating false positives (arXiv - Johari et al.). Optimizely's Stats Engine, built with Stanford statisticians, uses sequential testing to let you call winners up to 2.5 times faster for experiments over 50,000 visitors (Optimizely).
Netflix's case study is instructive. They use sequential canary tests to detect regressions across the entire distribution, not just the mean. For a metric like PlayDelay—the time for a title to start once play is pressed—a heavier tail can signal a severe regression. Their sequential method flags such issues while strictly controlling the false-detection probability (arXiv - Lindon et al.). That's a real-world, high-stakes example where sequential testing caught a problem that fixed-horizon might have missed.
Optimizely re-ran 48,000 historical experiments and found that fixed-horizon statistics declared a winner or loser in 36% of tests, while sequential testing did so in only 22%—39% fewer conclusive results (Optimizely). That sounds conservative, but it's a feature: fewer false positives. The trade-off is that you'll need larger effects to call a winner, but you'll trust the ones you do call.
Head-to-Head Comparison
| Criterion | Fixed-Horizon | Sequential |
|---|---|---|
| False-positive control | Good only if you never peek; otherwise up to 26.1% actual false-positive rate (Evan Miller) | Always valid, even with peeking; false-positive rate stays at nominal level |
| Decision speed | Must wait until full sample size; no early stopping | Can stop as soon as result is conclusive; up to 2.5x faster for large experiments (Optimizely) |
| Sample size requirements | Fixed upfront; often larger than needed | Can reduce required sample size by 50% or more in some cases (Evan Miller) |
| Implementation complexity | Simple; standard statistical tools | Requires specialized methods (e.g., SPRT, always-valid p-values), but available in platforms like Optimizely |
Who Should Use Which?
Fixed-horizon suits teams with:
- Strict discipline to pre-commit and never peek
- Simple metrics and low test volume
- Resources to run tests to full length without pressure
Sequential testing suits teams that:
- Need fast decisions (e.g., high-traffic sites)
- Can't resist peeking (be honest)
- Run many concurrent tests where false positives multiply
In practice, most teams—especially those at scale like Bing or Netflix—benefit from sequential testing. The always-valid property removes the peeking problem entirely. Even if you think you have discipline, the organizational pressure to check results is real. Sequential testing lets you look anytime without penalty.
But there's a catch: sequential testing can be conservative in declaring winners. If your goal is to identify small but genuine improvements quickly, you might need huge sample sizes. Fixed-horizon with a well-powered sample size can detect smaller effects, provided you stick to the plan. Evan Miller's simple sequential procedure is a middle ground: it picks a sample size N and stops early only if the advantage is large, cutting sample sizes by 50% in some cases (Evan Miller).
My recommendation? If you're running experiments on a high-traffic platform where false positives are expensive, switch to sequential testing. If you're a small team with low traffic and can commit to a fixed sample size, fixed-horizon is fine. But for most of us, the cost of peeking is too high. Use sequential testing and let the data speak whenever you're ready to listen.
Sources
- Optimizely - A/B testing glossary - https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
- arXiv - Always Valid Inference - https://arxiv.org/abs/1512.04922
- arXiv - Netflix canary testing - https://arxiv.org/abs/2205.14762
- Kohavi et al. - Online Controlled Experiments at Large Scale - https://exp-platform.com/large-scale/
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!