Who This Is For
If you're a product manager, marketer, or developer who runs A/B tests and has ever been tempted to check the results early—or has done it and felt that nagging doubt—this guide is for you. I've been there, and I've learned the hard way that our instincts about when to stop a test are often dead wrong. This is a practical walkthrough of how to design a test that you can trust, even when you're watching it like a hawk.
1. Start with the Right Hypothesis
Before you write a line of code or set up a variation, you need a clear, falsifiable hypothesis. The null hypothesis—that there's no real difference between your control and variation—is your starting assumption (Spotify Confidence). Your alternative is what you believe might be true. I always write down both. For example: "Null: The new checkout button color has no effect on conversion rate. Alternative: It increases conversion rate by at least 1%." This forces me to define what I'm testing and what kind of difference I care about. Too many tests are just "let's see if this works"—that's not a hypothesis, that's a gamble.
2. Choose Your Metrics and Set Your Significance Level
Decide on your primary metric before you start. In A/B testing, that's often conversion rate, but it could be revenue per user, click-through rate, or something else. You also need to set your significance level (alpha), conventionally 0.05, meaning you're willing to accept a 5% chance of a false positive (NIST e-Handbook). That's the "Type I error"—concluding a difference when there is none (Spotify Confidence). But here's the catch: if you're testing multiple metrics or variations, your error rate balloons. Optimizely found that when they re-analyzed 48,000 historical experiments, using a false discovery rate (which controls for multiple comparisons) led to roughly 20% fewer variations being declared winners than using a false positive rate at the same level (Optimizely Stats Engine). So, I limit my primary metric to one or two, and I'm wary of celebrating a p-value on a secondary metric unless it's dramatic.
3. Calculate Your Sample Size—and Don't Wing It
I can't stress this enough: you need a sample size calculation before you start, not after. The formula from Evan Miller is a rule of thumb I trust: n = 16 * (sigma^2 / delta^2), where delta is the minimum effect you want to detect and sigma^2 is the variance (for a binomial conversion rate, that's p(1-p)) (Evan Miller - How Not to Run an A/B Test). For a conversion rate, you need to estimate your baseline conversion rate and decide on your minimum detectable effect (MDE). Power is usually set at 80%, meaning you have an 80% chance of detecting the MDE if it truly exists (Optimizely sample size guide). If you don't do this, you're flying blind. I once ran a test without a proper sample size calculation, and after two weeks I saw a 3% lift that looked promising—but the confidence interval was so wide I couldn't tell if it was real. I had to kill it and start over with a proper plan.
4. Pick a Fixed Duration (or Go Sequential)
Once you have your sample size, you can estimate how long the test will run based on your daily traffic (Optimizely sample size guide). But here's where I've changed my mind: I used to be a stickler for fixed-duration tests, because peeking at your data and stopping as soon as you see significance can inflate your false-positive rate from 5% up to 30% (Optimizely A/B testing). That's a huge deal. Evan Miller showed that if you peek ten times, what you think is 1% significance is actually only 5% (Evan Miller - How Not to Run an A/B Test). So, the traditional advice is: decide on a sample size, run the test for that many visitors, and don't stop early. But that's hard in practice—executives want results, and you're curious. That's why I now prefer sequential testing methods, which allow you to peek at any time without inflating your error rates (arXiv - Always Valid Inference). Tools like Optimizely's Stats Engine use sequential testing, and they claim they can call a winner up to 2.5 times faster for large experiments (Optimizely Stats Engine). The trade-off is that sequential tests are more conservative—they require larger sample sizes to achieve the same power if there is no true effect, but they save you from your own impatience.
5. Watch for the Silent Killers: SRM and Novelty Effects
Even with a good design, things can go wrong. The first thing I check after launch is Sample Ratio Mismatch (SRM)—that's when the observed ratio of users in control vs. variation differs from your intended split (KDD 2019 - Diagnosing Sample Ratio Mismatch). An SRM is like a fever; it signals something is off—maybe a bug in your code, or some users are being excluded. If you ignore it, you might ship a bad change because your data is garbage. I had an SRM once because of a caching bug, and it almost led me to declare a winning variation that was actually a loser. The second killer is novelty and primacy effects: when users react to something just because it's new, or they need time to learn it (arXiv - Novelty and Primacy). These effects can make a good change look bad—or vice versa—if your test is too short. I once tested a new onboarding flow that showed a big drop in first-week retention, but after three weeks it turned positive as users got used to it. If I'd stopped early, I'd have killed a great feature.
6. Interpret with Practical Significance, Not Just p-Values
Let's say your test is done, and you have a p-value of 0.03. That means there's a 3% chance of seeing this data if the null hypothesis is true (Optimizely A/B testing). But that doesn't mean the effect is worth your time. Statistical significance is not the same as practical significance (Nielsen Norman Group). For example, a completion-rate difference of 0.2 percentage points—from 85.0% to 85.2%—can be statistically significant with a large enough sample, but it's probably not worth redesigning your checkout flow for that. I always look at the confidence interval, which tells me the range of possible true effects (Optimizely A/B testing). If the interval includes effects that are too small to matter, I'm skeptical. In one test, we had a p-value of 0.01, but the confidence interval was [0.1%, 1.5%], and the business case required at least 2% lift. So we didn't ship it. Practical significance is the final filter.
What I'd Actually Do
Here's my concrete recommendation: For most tests, use a fixed-horizon test with a pre-calculated sample size and a pre-committed end date, but use a sequential testing method if your platform supports it (like Optimizely's Stats Engine or Evan Miller's simple sequential procedure). My rule of thumb is to run a quick A/A test first to validate your setup, and run it quarterly as a sanity check (Optimizely A/A testing glossary). Then, for each experiment, I write a plan that includes my hypothesis, primary metric, significance level (0.05), power (0.8), MDE, and expected duration. I also set up alerts for SRM. I do not peek at the results until the test is complete, unless I'm using a sequential method. If I have to peek, I use the sequential approach so I don't fool myself. And when the test ends, I look at the effect size and confidence interval, not just the p-value. I've seen too many teams chase statistical significance and ship changes that had no real impact—or worse, hurt the business. Don't be that team.
Sources
- Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
- Spotify Confidence - https://confidence.spotify.com/glossary/frequentist-ab-testing
- Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
- Nielsen Norman Group - https://www.nngroup.com/articles/practical-significance/
- KDD 2019 - Diagnosing Sample Ratio Mismatch - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!