You think your last A/B test win was real? Let me hit you with a number: if you check results daily and stop as soon as significance appears, your false-positive rate can inflate from 5% to 30% (Optimizely). That means nearly one in three of your “wins” could be pure noise. And you're probably doing exactly that.
I'm going to be blunt: most A/B testing case studies are garbage. They're cherry-picked, peeked-at, and underpowered. The dirty secret of our industry is that less than a third of ideas tested at Bing moved the metrics they were designed to improve—and in search, that fraction is even lower (Bing). So when you read a case study claiming a 20% lift from changing a button color, your default assumption should be that it's a fluke. Here's why, and what to do instead.
The peeking problem is worse than you think
You've done it. You launch a test, check the dashboard every morning, and the moment that p-value dips below 0.05, you call it. Congratulations, you've just committed the cardinal sin of experimentation. Repeated significance testing always increases false positives, and if you peek ten times, what you think is 1% significance is actually just 5% (Evan Miller). In the absolute worst case—stopping the moment 5% significance appears—the actual false-positive rate can reach 26.1% (Evan Miller).
So your “95% confident” winner? It's more like 74% confident. That's not a win; that's a coin flip with extra steps. The defense is simple: pre-commit to a fixed end date, or use sequential testing that adjusts thresholds for multiple looks. If you're not doing either, you're not testing—you're gambling.
And before you tell me you're “just monitoring,” stop. Monitoring is peeking with a nicer name. Decide your sample size in advance and don't stop early unless you're using a method designed for it.
Your sample size is probably a joke
Most teams run tests for a week because that's what feels right. That's not a sample size calculation; that's a vibe. To detect a realistic effect—say, a 2% relative lift on a 5% baseline conversion rate—you need thousands of visitors per variation. The rule of thumb is n = 16 × (σ² / δ²), where δ is the minimum effect you want to detect (Evan Miller). For a conversion rate, σ² is p(1-p). Plug in p = 0.05 and δ = 0.001 (a 2% relative lift), and you get n ≈ 16 × (0.0475 / 0.000001) = 760,000 per variation. That's 1.5 million visitors total. If your site gets 10,000 visitors a day, that's 152 days. You're not running that test. You're running a 7-day test and pretending it's enough.
The result? Your test is underpowered, and you're likely to miss real effects (Type II errors) while chasing phantom ones. Power should be 80% (Optimizely), but with tiny samples and noisy metrics, you're lucky to hit 50%. And when you do find a “significant” result, it's often a Type I error—a false positive.
Statistical significance ≠ practical significance
Here's a case study from my own experience: a client ran a test on their checkout flow. They had 2 million visitors, and the variation improved completion rate from 85.0% to 85.2%. The p-value was 0.03—statistically significant! They wanted to ship it. I asked them: “Would a 0.2 percentage point improvement be noticeable to your users? Would it move your revenue enough to justify the engineering cost?” They couldn't answer. That's the problem. With a very large sample, even a microscopic improvement can produce a p-value below 0.001 (Nielsen Norman Group). Statistical significance only tells you a result is unlikely due to chance; it says nothing about how large or valuable the effect is (Nielsen Norman Group).
You need to define your minimum detectable effect (MDE) before the test, and you need to ask: if I see exactly that effect, will I actually care? If the answer is no, don't run the test. Or run it for learning, but don't call it a win.
What about Bayesian and sequential methods?
You might be thinking: “Fine, I'll just use Bayesian testing. It allows continuous monitoring without inflating false positives.” That's true—Bayesian methods output probabilities like “there is a 92% probability that Variant B is better” and let you peek without penalty (Optimizely). VWO even claims Bayesian can provide actionable results almost 50% faster (VWO). But here's the catch: Bayesian results depend on your prior, and if you're not careful, you're just swapping one set of assumptions for another. Sequential testing, like Optimizely's Stats Engine, is another option. It gives you always-valid p-values and can call a winner up to 2.5 times as fast for experiments over 50,000 visitors (Optimizely). But it also declared 39% fewer conclusive results in a re-analysis of 48,000 experiments (Optimizely). That's the trade-off: you get validity, but you lose some power to detect small effects.
My take: for most teams, fixed-horizon testing with a pre-computed sample size is still the gold standard. If you must peek, use a sequential method, but understand that you're trading sensitivity for flexibility. Don't mix and match.
What I'd actually do
Stop treating A/B tests as magic. Start treating them as a discipline. Here's my concrete recommendation:
First, run an A/A test quarterly. If you see a “significant” difference between two identical versions, your setup is broken. Fix it before you trust any result. Second, calculate your sample size before every test using the n = 16 × (σ² / δ²) rule, and don't launch until you have the traffic to hit it. If you can't, don't run the test—or accept that you're doing qualitative research, not experimentation. Third, pre-commit to a fixed duration and don't peek. If you absolutely must monitor, use a sequential method like Optimizely's Stats Engine or a Bayesian approach, but pick one and stick with it. Fourth, always check for sample ratio mismatch (SRM). An SRM is a symptom of data quality issues, and ignoring it can lead to shipping a bad feature that looks good (KDD 2019). Finally, define practical significance thresholds in advance. A 0.2 percentage point lift on a 85% baseline is not worth your time.
You want better case studies? Then run better tests. The numbers don't lie—but your dashboards do.
Sources
- Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Nielsen Norman Group - https://www.nngroup.com/articles/practical-significance/
- Bing - https://blogs.bing.com/search-quality-insights/August-2013/Large-Scale-Experimentation-at-Bing/
- KDD 2019 - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!