At Bing, fewer than one in three ideas that get tested in a controlled experiment actually move the metrics they were designed to improve. In the optimized search domain, it's even lower (Bing - Large Scale Experimentation at Bing). Read that again. The world's most sophisticated experimentation platform, running over 250 experiments on a typical day, with almost every user exposed to about 15 simultaneous tests, still bats under .333 on good ideas.
So when you read a glowing case study about a button color change that lifted conversions 34%, your first instinct should be suspicion, not inspiration. That number almost certainly came from a test that peeked, or a sample size that was never pre-computed, or a metric chosen after the fact because it looked good. We have stopped treating published case studies as playbooks. They are anecdotes wearing lab coats.
The case study you should actually copy is not a case study at all
Here's our thesis: the most valuable A/B testing "case study" is not any single experiment. It's the process discipline that lets you run thousands of them without drowning in false positives. The experiments themselves are disposable. The guardrails are the asset.
Consider what happens when you skip the guardrails. Checking results daily and stopping the moment significance appears can inflate your false-positive rate from a nominal 5% to as high as 30% (Optimizely - A/B testing). Evan Miller's classic demonstration is blunter still: if you peek at a test ten times, what you think is 1% significance is actually just 5% significance. And in the worst case, where you stop as soon as 5% significance appears after every single observation, the real false-positive rate reaches 26.1% (Evan Miller - How Not to Run an A/B Test). That's not a rounding error. That's a coin flip with extra steps.
We've seen this movie. A team runs a test, peeks on day two, sees p = 0.04, ships the change, and declares victory. Three months later the metric has reverted. Nobody connects the dots because the case study was already written.
What we do instead: pre-commit, calibrate, and watch for fevers
Before a single user is exposed, we compute sample size. The formula isn't exotic: you need a baseline conversion rate, a minimum detectable effect, a significance level (usually 0.05), and power (often 0.8). The minimum detectable effect is the smallest lift your test can reliably catch; anything below it is indistinguishable from noise (Evan Miller - Sample Size Calculator). If you can't articulate your MDE, you're not ready to test. You're ready to guess.
We also run A/A calibration tests quarterly. That means testing two identical versions against each other. If your software reports a statistically significant difference between them, your setup is broken, not your hypothesis (Optimizely - A/A testing glossary). It's the cheapest insurance policy in experimentation.
And we treat Sample Ratio Mismatch like a fever. An SRM — when the observed traffic split diverges from the expected split — is a symptom, not a disease. It can be caused by bot filtering, redirects, slow pages, or a dozen other plumbing failures. Ignoring an SRM without knowing its root cause can make a bad product change look good and get shipped to users (KDD 2019 - Diagnosing Sample Ratio Mismatch). We've killed more tests for SRM than for bad ideas.
One more thing: we log every experiment in a central registry. Not just the ones that win. The failures, the inconclusive ones, the ones that never launched because of an SRM. That log has become our most valuable asset. It's how we spot patterns — like the fact that tests run on Fridays are 20% more likely to have SRM issues, probably because of traffic anomalies. You don't get that from a case study.
The counter-argument: sequential testing and Bayesian methods let you peek safely
The strongest objection to our discipline-first position is that it's outdated. Sequential methods and Bayesian inference genuinely do allow continuous monitoring without inflating false positives. Always-valid p-values let you look at data whenever you want and still make valid inferences (Johari et al. - Always Valid Inference). Optimizely's Stats Engine, built with Stanford statisticians and launched in January 2015, can call a winner up to 2.5 times as fast as fixed-horizon statistics for experiments over 50,000 visitors (Optimizely - Stats Engine). Netflix uses sequential canary tests that catch distributional regressions — a heavier tail on PlayDelay, for instance — while strictly controlling false-detection probability (Lindon et al. - Netflix canary testing). Bayesian tools can give you a straight answer like "there is a 92% probability that Variant B is better" and let you monitor continuously (Optimizely - A/B testing).
All true. And still not a rebuttal. Sequential and Bayesian methods don't eliminate the need for pre-commitment; they change what you pre-commit to. You still pick a stopping rule. You still decide what metric matters. You still need an SRM check and an A/A calibration. The guardrails move; they don't disappear. A team that peeks with a sequential test and no pre-registered metric is just as lost as a team that peeks with a fixed-horizon test — arguably more lost, because the tooling gives them false confidence.
Why we care about practical significance more than p-values
A p-value of 0.03 does not mean your treatment has a 97% chance of working. It means that if the treatment did nothing, you would see data this extreme about 3% of the time (Optimizely - A/B testing). That's it. It's a statement about surprise, not value.
The Nielsen Norman Group's example is the one we tape to the wall: a completion-rate difference of 0.2 percentage points — 85.0% versus 85.2% — can be statistically significant at p = 0.03 and still be too small to justify a design change (Nielsen Norman Group - Practical Significance). At Bing's scale, a 1% revenue improvement is worth about $10 million annually in the US (Kohavi et al. - Online Controlled Experiments at Large Scale). But a 0.2-point completion-rate lift in a checkout flow might be worth nothing at all if it doesn't change customer behavior in a way that compounds.
We have a rule: if the lift is real but the business case isn't, we don't ship. We write it down and move on.
The case studies worth reading are not the ones with the biggest lifts. They're the ones that tell you what failed, why it failed, and what guardrail caught it. Bing's experimentation system is credited with increasing annual revenues by hundreds of millions of dollars — and with identifying many negative features that were never deployed (Kohavi et al. - Online Controlled Experiments at Large Scale). The second half of that sentence is the part nobody puts in a headline.
So here's our recommendation, and we'll keep it short: stop collecting case studies. Start collecting failure post-mortems and calibration logs. Run your A/A tests quarterly. Pre-compute your sample size. Treat an SRM as a stop sign, not a speed bump. And when someone shows you a 34% lift from a button color, ask them how many times they peeked. The answer will tell you more than the number ever could.
Sources
- Bing - Large Scale Experimentation at Bing - https://blogs.bing.com/search-quality-insights/August-2013/Large-Scale-Experimentation-at-Bing/
- Optimizely - A/B testing - https://www.optimizely.com/optimization-glossary/ab-testing/
- Evan Miller - How Not to Run an A/B Test - https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- KDD 2019 - Diagnosing Sample Ratio Mismatch - https://www.kdd.org/kdd2019/accepted-papers/view/diagnosing-sample-ratio-mismatch-in-online-controlled-experiments-a-taxonom
- Nielsen Norman Group - Practical Significance - https://www.nngroup.com/articles/practical-significance/
- Kohavi et al. - Online Controlled Experiments at Large Scale (KDD 2013) - https://exp-platform.com/large-scale/
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!