Skip to main content
Case Studies

Why Your A/B Test Results Are Lying to You: A Case Study Guide

Peeking, p-hacking, and tiny effects: A/B testing is full of traps. Learn from real case studies how to get trustworthy results and avoid common mistakes.

Why can't I trust my A/B test results?

You've run an A/B test, the p-value dips below 0.05, and you're ready to ship the winner. But hold on. If you peeked at the data daily, stopped early, or tested a dozen metrics, your 'significant' result might be pure noise. At Bing, where they run over 250 experiments a day, less than a third of ideas actually move the metrics they were designed to improve (Bing - Large Scale Experimentation at Bing). So, the first lesson: your A/B test results are only as reliable as your methodology. You need to stop lying to yourself.

Is peeking really that bad?

Yes. Peeking—checking results daily and stopping as soon as significance appears—can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely (A/B testing)). If you peek ten times, what you think is 1% significance is actually just 5% (Evan Miller - How Not to Run an A/B Test). In the worst case, stopping as soon as 5% significance appears, the real false-positive rate can hit 26.1% (Evan Miller - How Not to Run an A/B Test). So, the classic advice: decide your sample size in advance, run the test to completion, and resist the urge to peek. Or, better yet, use sequential testing.

Can I stop my test early if it's winning?

You can, but only if you're using sequential methods. Sequential testing lets you peek any time and get valid results, without inflating false positives. Evan Miller's simple sequential procedure: pick a sample size N, assign subjects 50/50, and stop when the treatment-minus-control success count reaches 2 times the square root of N (declaring a winner) or the total success count reaches N (declaring no winner). This can reduce required observations by 50% or more in some cases (Evan Miller - Simple Sequential A/B Testing). Optimizely's Stats Engine, which uses sequential testing, can call a winner or loser up to 2.5 times faster than fixed-horizon statistics for experiments over 50,000 visitors (Optimizely Stats Engine). So, yes, you can stop early—just do it the right way.

Why is my 'significant' result not practically significant?

Statistical significance (p

Should I use Bayesian testing instead?

Bayesian testing offers a different approach: it outputs probabilities like 'there is a 92% probability that Variant B is better' and allows continuous monitoring without inflating false positives (Optimizely (A/B testing)). It can also incorporate prior knowledge from previous tests (VWO A/B testing guide). VWO reports that the Bayesian approach can provide actionable results almost 50% faster (VWO A/B testing guide). But it's not a magic bullet. The key is to pick a methodology and stick to it. If you prefer frequentist, use sequential testing. If you prefer Bayesian, use it properly. Just don't mix them or switch mid-test.

How do I avoid these mistakes in my own tests?

Here's a blunt, practical checklist:

First, run A/A tests quarterly to validate your setup. If an A/A test reports a statistically significant difference between two identical versions, your implementation is broken (Optimizely A/A testing glossary). Second, check for Sample Ratio Mismatch (SRM)—when the observed sample ratio differs from the expected ratio. SRM is a symptom of data quality issues, and ignoring it can lead you to ship a bad change (KDD 2019 - Diagnosing Sample Ratio Mismatch). Third, use CUPED to reduce variance. At Bing, CUPED reduced metric variance by about 50%, which effectively gives you the same statistical power with half the users or half the duration (Deng et al. - CUPED (WSDM 2013)). Fourth, pre-commit to a fixed sample size or use sequential testing. Fifth, focus on practical significance, not just p-values.

Quick tip: If you're testing a new design, measure the effect on the entire distribution, not just the mean. Netflix uses sequential canary tests to flag regressions in metrics like PlayDelay—the time for a title to start once play is pressed—because a heavier treatment tail can spell trouble even if the mean looks fine (arXiv - Netflix canary testing (Lindon et al.)).

What's the single most important thing to remember?

Your A/B test is only as good as your methodology. Peeking, multiple looks, and ignoring practical significance will lead you astray. Use sequential testing, pre-commit to your sample size, and always ask: is this effect big enough to care about?

Sources

  • Optimizely - https://www.optimizely.com/optimization-glossary/ab-testing/
  • Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
  • Optimizely Stats Engine - https://www.optimizely.com/insights/blog/statistics-for-the-internet-age-the-story-behind-optimizelys-new-stats-engine/
  • Nielsen Norman Group - Practical Significance - https://www.nngroup.com/articles/practical-significance/
  • Optimizely A/A testing glossary - https://www.optimizely.com/optimization-glossary/aa-testing/
  • KDD 2019 - Diagnosing Sample Ratio Mismatch - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
  • Deng et al. - CUPED - https://exp-platform.com/Documents/2013-02-CUPED-ImprovingSensitivityOfControlledExperiments.pdf

Share this article:

Comments (0)

No comments yet. Be the first to comment!