I once watched a product manager stop an A/B test after three days because the p-value hit 0.03. The variation was a button color. Two weeks later, conversion was flat. That one decision cost two sprints and a lot of credibility. The culprit? Peeking. Checking results daily and stopping as soon as significance appears can inflate the false-positive rate from a nominal 5% to as high as 30% (Optimizely). That's not a rounding error—it's a coin flip dressed up as science. In this head-to-head, I'll compare the three most common statistical approaches to A/B testing: fixed-horizon frequentist, sequential testing (including always-valid inference), and Bayesian. I'll judge them on false-positive control, speed to decision, ease of use, and flexibility. My bias is clear: I'll take sequential methods over fixed-horizon almost every time, but there are conditions where each wins.
The problem with fixed-horizon testing
Fixed-horizon testing is the classic approach: you pick a sample size in advance, run the test until you hit that number, then compute a p-value. The p-value is the probability of seeing data as extreme as yours, assuming no real difference (the null hypothesis). If p
Evan Miller showed that if you peek at an ongoing experiment ten times, what you think is 1% significance is actually just 5% significance. In the worst case—stopping as soon as 5% significance appears, with a test after every observation—the actual false-positive rate can reach 26.1% (Evan Miller). That's a disaster. You'll ship losers and kill winners. The defense is to pre-commit to a fixed end date or use sequential methods that adjust thresholds for multiple looks (Optimizely). I've seen teams try to enforce a no-peeking policy. It fails. The temptation is too strong, and stakeholders demand updates.
Sequential testing: peeking without the penalty
Sequential testing flips the script. Instead of a fixed sample size, you define a stopping rule that's valid at any time. Evan Miller's simple sequential procedure is beautifully concrete: choose a sample size N, assign subjects randomly 50/50, and stop when the treatment-minus-control success count reaches 2 times the square root of N (declare a winner) or when the total success count reaches N (declare no winner). In some circumstances, this can reduce the number of observations required for a successful experiment by 50% or more, and it works extremely well with low conversion rates (Evan Miller).
Optimizely's Stats Engine, developed with Stanford statisticians and launched in 2015, takes a different sequential approach. It uses always-valid p-values and confidence intervals, so you can peek anytime without inflating false positives. For experiments over 50,000 visitors, it can call a winner or loser up to 2.5 times as fast as fixed-horizon statistics (Optimizely Stats Engine). That speed matters. But there's a cost: when Optimizely re-ran 48,000 historical experiments, fixed-horizon declared a winner or loser in 36% of tests, while sequential declared winners or losers in only 22%—39% fewer conclusive results. So you trade some conclusiveness for validity under peeking.
I'll take that trade. Here's why: a false positive is worse than an inconclusive test. An inconclusive test tells you to run another experiment or move on. A false positive ships a bad feature and erodes trust in the process. Sequential methods also let you stop early when the effect is huge, saving time and money. Netflix uses sequential canary tests that compare entire distributions, not just means, to catch regressions like a heavier tail in PlayDelay—the time for a title to start playing (Netflix canary testing). That's the kind of rigor I want.
Bayesian: the continuous monitoring alternative
Bayesian testing outputs probabilities like 'there is a 92% probability that Variant B is better' and allows continuous monitoring without inflating false positives (Optimizely). That's a huge selling point. VWO reports that the Bayesian approach can provide actionable results almost 50% faster while focusing on statistical significance, and it can directly state the probability that one variation has a lower conversion rate than another (VWO). For teams that struggle to interpret p-values, Bayesian is a breath of fresh air. It incorporates prior knowledge from previous tests, which can be powerful if you have strong priors.
But I'm wary. Bayesian methods require choosing a prior, and that choice can be subjective. If your prior is wrong, your results are wrong. Frequentist sequential methods make fewer assumptions. Also, Bayesian continuous monitoring is not a free lunch—it controls false positives under certain conditions, but it's not a magic bullet for multiple comparisons. When Optimizely compared false discovery rate to false positive rate on 48,000 experiments, they found roughly 20% fewer variations with a false discovery rate below 0.1 than with a false positive rate at the same level (Optimizely Stats Engine). That's a reminder that testing many metrics and variations at once is costly, regardless of framework.
Head-to-head comparison
| Criteria | Fixed-Horizon Frequentist | Sequential Testing | Bayesian |
|---|---|---|---|
| False-positive control under peeking | Poor (up to 30% with daily peeking) | Excellent (always-valid p-values) | Good (continuous monitoring allowed) |
| Speed to decision | Slow (must wait for sample size) | Fast (can stop early; up to 2.5x faster for large tests) | Fast (almost 50% faster per VWO) |
| Ease of interpretation | Moderate (p-values misunderstood) | Moderate (confidence sequences) | High (direct probability statements) |
| Flexibility (multiple metrics, early stopping) | Low (requires fixed sample, penalizes multiple looks) | High (designed for continuous monitoring) | High (but prior choice matters) |
Who each option is for:
- Fixed-Horizon: Teams with very low traffic, simple one-off tests, and iron discipline to never peek. Or those using tools that don't support sequential methods.
- Sequential: Most product teams, especially those with continuous experimentation, multiple metrics, and stakeholders who demand updates. It's the safest default.
- Bayesian: Teams that value intuitive probability statements and have reliable prior information. Also good for communicating with non-technical stakeholders.
Which wins? For me, sequential testing wins by a nose. It directly addresses the peeking problem that plagues real-world A/B testing, and it lets you stop early when the effect is large. The trade-off—fewer conclusive results—is acceptable because inconclusive tests are not failures; they're learning opportunities. Bayesian is a close second, and I'd recommend it for teams that struggle with p-values or have strong priors. Fixed-horizon is only acceptable if you have strict no-peeking policies and low traffic.
My recommendation and a concrete example
Imagine you run an e-commerce site with a baseline conversion rate of 3%. You want to detect a 10% relative lift (to 3.3%). Using the rule-of-thumb sample size formula n = 16 * (sigma-squared / delta-squared), where sigma-squared is p(1-p) = 0.03*0.97 = 0.0291 and delta = 0.003, you get n = 16 * (0.0291 / 0.000009) = 16 * 3233 = 51,733 per variation (Evan Miller). That's over 100,000 visitors total. If you peek daily and stop at the first significant result, you'll likely stop too early and get a false positive. With sequential testing, you can monitor continuously and stop when the evidence is strong—or when you hit the maximum sample size. If the effect is real and large, you might stop 50% earlier, saving weeks. If it's not, you'll correctly conclude no difference.
I've also seen the pain of ignoring data quality. A Sample Ratio Mismatch (SRM)—where the observed traffic split differs from the expected—is a symptom of data quality issues, and ignoring it can make a bad product modification appear good (KDD 2019). Always run A/A tests quarterly to validate your setup (Optimizely). If an A/A test shows a significant difference, your implementation is broken.
So here's my final take: adopt sequential testing as your default. It's the only method that lets you peek without lying to yourself. If you can't, use Bayesian and be honest about your priors. Only use fixed-horizon if you have the discipline of a monk and the traffic of a small blog.
Sources
- Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
- VWO - https://vwo.com/ab-testing/
- Netflix canary testing - https://arxiv.org/abs/2205.14762
- KDD 2019 - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!