Skip to main content
Optimization Tools

Stop Peeking at Your A/B Test: Why Fixed Sample Sizes Are Your Friend

You're ruining your A/B tests by peeking. Here's how to stop, why sequential testing helps, and why you need to pre-commit to a sample size.

Imagine you're a product manager at a mid-sized SaaS company. You've launched an A/B test on your pricing page, and after just two days, you see a 10% lift in conversion for the variant. Your heart races. You want to call it a win and ship it. But if you stop the test now, you might be fooling yourself.

This is the classic peeking problem. In A/B testing, peeking—checking results daily and stopping as soon as significance appears—can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely). That means nearly one in three times you call a winner, you're actually chasing noise. The more you peek, the more your reported significance levels are off. If you peek ten times, what you think is 1% significance is actually just 5% significance (Evan Miller).

The solution isn't to ban peeking; it's to use methods that make peeking safe. You have two solid options: pre-commit to a fixed sample size, or use sequential testing. My blunt advice: stop running fixed-horizon tests like they're a buffet; you need to decide how much data you'll collect before you start.

The Seduction of the Peek

We've all been there. The test is running, and the numbers look promising. You refresh the dashboard, and the variant is winning by a comfortable margin. It's tempting to end the experiment early and move on. But here's the thing: when you repeatedly test significance as data trickles in, you're not just getting a single p-value; you're getting a range of p-values, and the chance that at least one of them crosses the 0.05 threshold is much higher than 5%. In a worst-case scenario, if you stop as soon as 5% significance appears after every observation, the actual false-positive rate can reach 26.1% (Evan Miller).

The fix is simple: decide on a sample size in advance, based on the effect you care about, and don't stop until you reach it. This is the core of classical hypothesis testing. You set your significance level (usually 0.05), your power (often 0.8), and your minimum detectable effect (MDE). Then you calculate the sample size you need. The formula is straightforward: for a given variance, you need roughly 16 times the variance divided by the square of the effect you want to detect (Evan Miller). For example, if your baseline conversion rate is 5% and you want to detect a 10% relative lift (so an absolute increase of 0.5 percentage points), the variance is p(1-p), which is about 0.05*0.95 = 0.0475. The squared effect is (0.005)^2 = 0.000025. So n = 16 * 0.0475 / 0.000025, which is about 30,400 per variation. That's a lot of visitors, but it's what you need to be confident.

If you don't pre-commit, you're just guessing. And guessing is how you end up shipping a change that actually hurts your metrics.

Why Sequential Testing Is a Game-Changer (But Not a Silver Bullet)

Now, some of you might be thinking, "But I want to check results early! I can't wait weeks for a test to finish." That's where sequential testing comes in. Sequential testing allows you to peek at any time and still have valid results, without inflating your false-positive rate. It's like having a test that adapts to the data, letting you stop early if the treatment is a clear winner or loser.

Evan Miller's simple sequential procedure is one elegant method: you choose a sample size N, assign subjects randomly 50/50, and stop when the treatment-minus-control success count reaches 2 times the square root of N (declaring a winner) or the total success count reaches N (declaring no winner). In some cases, this can reduce the number of observations needed for a successful experiment by 50% or more, especially with low conversion rates (Evan Miller).

But here's the catch: sequential testing requires you to plan for it. You can't just switch mid-test. And it's not a free lunch—you might need a larger maximum sample size to get the same power, but you'll often stop earlier if the effect is real. Optimizely's Stats Engine, which uses sequential methods, has shown that it can call a winner or loser up to 2.5 times as fast as fixed-horizon statistics for experiments over 50,000 visitors (Optimizely Stats Engine).

The Counterargument: "But I Have a Business Deadline"

You might argue, "I can't wait for a fixed sample size because I have a product launch deadline. I need results now." I hear you. But here's the thing: if you can't wait for a properly powered test, you shouldn't be running an A/B test at all. You should be making a decision based on your best judgment, not on a statistically meaningless result. Running an underpowered test and acting on it is worse than not testing, because you'll be confident in a result that is likely false.

Consider this: at Bing, less than a third of ideas move the metrics they were designed to improve (Bing). So the base rate of a winning idea is low. If you're running an underpowered test, you're likely to miss even the few good ideas, and you'll be fooled by noise. A test that isn't powered to detect a meaningful effect is a waste of time.

If you're truly time-constrained, use a Bayesian approach or a multi-armed bandit, which can adapt faster. For instance, VWO reports that Bayesian testing can provide actionable results almost 50% faster while directly stating the probability that one variation beats another (VWO). But even then, you need to set a stopping rule before you start, or you'll fall into the same trap.

Don't Forget the A/A Test and SRM Checks

Before you even run a test, you should validate your setup. Run an A/A test periodically—Optimizely recommends doing this quarterly (Optimizely). If your A/A test reports a statistically significant difference between two identical versions, something is broken with your implementation. Fix it before you trust any results.

Also, check for Sample Ratio Mismatch (SRM). An SRM is when the observed sample ratio differs from the expected ratio, and it's a symptom of data quality issues (KDD 2019). Ignoring an SRM can lead you to ship a bad change because a bug in your assignment caused the control and treatment groups to differ in ways unrelated to your variation. Always check for SRM before analyzing your results.

The Practical Significance Trap

Even if you do everything right, you can still be misled by statistical significance. With a very large sample size, even a microscopic improvement—like a 0.03% drop in form-field errors—can produce a p-value below 0.001 (Nielsen Norman Group). But is that difference meaningful to your business? Probably not. Statistical significance only tells you that a result is unlikely due to chance; it says nothing about the size or value of the effect.

Conversely, an effect that is not statistically significant may still be actionable. For example, if you see an 80% drop in task completion in a small-sample usability test, that's a huge red flag even if the p-value is above 0.05 (Nielsen Norman Group). Use your judgment. Don't let the p-value make your decisions for you.

What to Do Instead: A Practical Plan

So here's my blunt advice: before you launch your next test, use a sample size calculator to determine how many visitors you need. Pick a minimum detectable effect that is practically meaningful—not the smallest effect you can detect, but the smallest effect that would matter for your business. Set your significance level at 0.05 and power at 0.8. Then, don't peek until you've reached that sample size. If you absolutely must peek, use a sequential method like Optimizely's Stats Engine or Evan Miller's simple procedure. And always run an A/A test and check for SRM to ensure your data is trustworthy.

One more thing: be wary of running too many variations at once. The more you test, the higher your chance of a false positive. In fact, when paired with classical statistics, peeking and testing many goals and variations can increase your false-positive rate by over 5x (Optimizely Stats Engine). So keep your tests simple.

Quick tip: If you're using a fixed-horizon test, set a calendar reminder for the date you'll have enough data, and don't look at the results until then. This is the single most effective way to avoid peeking.

In the end, the most important thing to remember is this: an A/B test is only as good as its planning. If you don't pre-commit to a sample size and a stopping rule, you're not doing science—you're just guessing with data. And that's how you end up with a 30% false-positive rate.

Sources

  • Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
  • Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Evan Miller - Sample Size Calculator - https://www.evanmiller.org/ab-testing/sample-size.html
  • Optimizely A/A testing glossary - https://www.optimizely.com/optimization-glossary/aa-testing/
  • KDD 2019 - Diagnosing Sample Ratio Mismatch - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom

Share this article:

Comments (0)

No comments yet. Be the first to comment!