Are you checking your A/B test results every day and feeling the urge to stop it as soon as it hits 95% significance? If so, you're not alone, but you're also fooling yourself. I’ve been there, and it took a few painful lessons to learn the right way. This article is for anyone who runs A/B tests—marketers, product managers, or developers—who wants to stop making decisions on shaky ground and start trusting their experiments.
1. The Temptation of Peeking and Its Hidden Cost
When I first ran A/B tests, I’d peek at the numbers daily, ready to declare a winner the moment the p-value dipped below 0.05. It felt productive. But here's the kicker: peeking at your data and stopping as soon as you see significance can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely). That means nearly one in three of my 'wins' were probably false. The reason is that if you test repeatedly, you increase the chance that random fluctuations will cross the threshold. Evan Miller, who's written extensively on this, points out that if you peek ten times, what you think is 1% significance is actually just 5% (Evan Miller). In a worst-case scenario where you stop as soon as 5% appears, your actual false-positive rate can hit 26.1% (Evan Miller). Yikes.
So, what's the fix? The first step is to pre-commit to a sample size before you start. Decide how many visitors you need based on the baseline conversion rate and the minimum effect you care about. A rule of thumb I use comes from Evan Miller: n = 16 × (sigma² / delta²), where delta is the smallest effect you want to detect and sigma² is the variance (for a proportion, that's p(1-p)). For example, if your current conversion is 10% and you want to detect a 1% absolute increase, you'd need a lot of users—about 16 × (0.1 × 0.9) / (0.01²) = 14,400 per variation. That might sound like a lot, but it's better than chasing noise.
2. How to Choose Your Sample Size and Duration
Once you know your sample size, you can estimate how long to run. If you need 14,400 visitors per variation and you get 2,000 visitors a day, you'll need about 14.4 days per variation—multiply by two variations for a total of about 29 days. But let me be clear: longer isn't always better. You have to balance the need for power with the reality of business cycles and seasonal changes (Optimizely). Running a test for a month might be fine, but if you're in retail, you don't want to span Black Friday unless that's your goal.
Another thing to watch out for is the minimum detectable effect (MDE). If you set your MDE too small, you'll need a massive sample. If you set it too large, you might miss a meaningful change. I've learned to be realistic—if a 0.2% difference in completion rate is statistically significant (p=0.03), it might not be worth changing your design (Nielsen Norman Group). That's because with a huge sample, even tiny differences become significant. So, focus on what matters for your business, not just what's statistically true.
3. What Can Go Wrong: A/A Tests and Sample Ratio Mismatches
Now, you might think you've got everything figured out, but there's a sneaky issue: your test setup might be broken. That's why I run A/A tests—where you compare two identical versions—every quarter (Optimizely). If an A/A test reports a significant difference between identical pages, your implementation is flawed, and you can't trust any results from that setup (Optimizely).
Another red flag is a Sample Ratio Mismatch (SRM), which is when the actual number of visitors in each group doesn't match what you expected. Think of it as a fever—it's a symptom of a bigger problem (KDD 2019). Ignoring an SRM can lead you to ship a bad change or miss a good one (KDD 2019). For instance, if you're running a 50/50 split but you see 60% of traffic going to the variation, something is off. Maybe your redirect code has a bug, or your targeting is wrong. Always check for SRM before you trust your results.
4. Better Alternatives: Sequential Testing and CUPED
If you're tired of the fixed-sample constraint, there's a better way: sequential testing. Unlike fixed-horizon tests, sequential methods let you peek anytime without inflating your false positives. Optimizely's Stats Engine uses this approach, and it can call a winner or loser up to 2.5 times as fast as fixed-horizon statistics for large experiments (Optimizely). When they re-analyzed 48,000 historical experiments, sequential testing found 39% fewer conclusive results than fixed-horizon—meaning many of those 'conclusive' results were probably false (Optimizely).
If you're not using a platform that does this, you can use Evan Miller's simple sequential procedure: pick a sample size N, assign subjects 50/50, and stop when the difference in successes reaches 2 × sqrt(N) (declaring a winner) or when total successes hit N (declaring no winner) (Evan Miller). In some cases, this can cut your required sample size by half or more (Evan Miller).
Another technique to speed things up is CUPED (Controlled-experiment Using Pre-Experiment Data), which uses data from before the test to reduce variance. At Bing, CUPED reduced variance by 45%, 52%, and 49% in three experiments, effectively halving the number of users you need to achieve the same power (Deng et al.). So, instead of waiting an extra week, you might be done in half the time.
Bottom Line
If you take away one thing from this, it's this: pre-commit to a sample size and stop peeking. If you can't resist, switch to a sequential testing approach that lets you monitor continuously without lying to yourself. Your future self—and your conversion rate—will thank you.
Sources
- Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
- Optimizely A/A testing glossary - https://www.optimizely.com/optimization-glossary/aa-testing/
- KDD 2019 - Diagnosing Sample Ratio Mismatch - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
- Deng et al. - CUPED - https://exp-platform.com/Documents/2013-02-CUPED-ImprovingSensitivityOfControlledExperiments.pdf
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!