On this article

The ROI of Experimentation: Why 2% of Your Tests Pay for the Rest

The counterintuitive economics of A/B testing: most tests fail, a rare few win huge, and near-zero test costs make the math pay off.
This is some text inside of a div block.

The ROI of Experimentation: Why 2% of Your Tests Pay for the Rest

Every experimentation program eventually meets the same challenge from a skeptical executive: does all this testing actually pay for itself? Most experiments don't win, and running them takes time and engineering. It's a fair question — and the research on the economics of experimentation gives a clear, if counterintuitive, answer. The return doesn't come from being right most of the time. It comes from being cheaply, systematically exposed to the rare occasions when you're spectacularly right.

The short version: Experimentation pays off because of a counterintuitive fact about innovation — the value of new ideas is extremely "fat-tailed." On Microsoft Bing's platform, the top 2% of ideas produced roughly three-quarters of all the gains, while most experiments did little or nothing. Because the marginal cost of an online test is near zero, the return-maximizing strategy isn't a few big bets — it's running many cheap experiments to catch the rare outliers. It's why testing added double-digit annual revenue growth at companies like Microsoft, and why a single Bing experiment was worth over $100 million a year. The ROI comes not from a high win rate, but from cheaply catching the few oversized winners.

The counterintuitive economics: fat tails

The finding that reframes the whole ROI question comes from a 2020 Journal of Political Economy study, "A/B Testing with Fat Tails" (Azevedo, Deng, Montiel Olea, Rao & Weyl). The authors asked how a company should best spend limited experimentation resources, and showed the answer depends on the shape of idea quality. If good ideas cluster near average, a few big, high-powered tests are optimal. But if the distribution is fat-tailed — dominated by rare, unpredictable, enormous wins — the optimal move flips: run many smaller experiments to fish for the outliers.

Then came the empirical punchline. Measuring the real distribution on Bing's experimentation platform, they found extremely fat tails — the top 2% of ideas accounted for about 74.8% of all the gains, an extreme version of the 80/20 rule. Their models suggested that even modest changes to how a company experiments could sharply raise its innovation productivity — testing just 20% more ideas on the same traffic increased productivity by around 17%. In plain terms: in the real world, the winners are so disproportionately valuable that trying more ideas beats trying to perfectly pick a few.

Most experiments fail — and that's the model working

The flip side of fat tails is a humbling hit rate. In well-optimized products like Bing and Google, only about 10–20% of changes move the target metric in the right direction (Kohavi & Thomke, Harvard Business Review, 2017). Across Microsoft more broadly, Kohavi's rule of thumb is that roughly a third of ideas win, a third do nothing, and a third are actually negative. Airbnb fits the pattern: Kohavi has recounted that of 250 ideas tested, only about 20 moved the key metric — but those winners drove a ~6% lift in booking conversion, worth hundreds of millions.

A 80–90% "failure" rate isn't a broken program. It's a portfolio doing exactly what the math predicts — most bets are duds, and a few are franchise-makers.

Why the math works: a test costs almost nothing

Fat tails alone wouldn't guarantee ROI if each experiment were expensive. The second half of the equation is that online experimentation is now nearly free at the margin. As Kohavi, Tang, and Xu document, mature platforms drive the marginal cost of an experiment close to zero — which is exactly why Google, LinkedIn, and Microsoft each run more than 20,000 controlled experiments a year, and Booking.com around 25,000.

Put the two together and the logic is inescapable: near-zero-cost shots at a fat-tailed jackpot. You don't need a good hit rate when each attempt is cheap and the occasional hit is worth nine figures.

The documented returns

The returns aren't theoretical. Companies have shared the numbers:

  • A single Bing experiment worth $100M+. A small, cheap change to how Bing displayed ad headlines — an idea shelved for six months because it looked unremarkable — lifted revenue about 12% when finally tested: more than $100 million a year in the US alone (HBR, 2017). One outlier, caught cheaply.
  • Double-digit annual growth. Experimentation collectively raised Bing's revenue per search on the order of 10–25% per year — the compounding of many small wins.
  • Speed has a price tag. Bing found every 100 milliseconds of latency improvement was worth roughly $18 million a year — experimentation quantifying the value of otherwise-invisible engineering work.
  • Smaller changes, real money. A "larger but fewer ads" test was worth about $50M/yr; a barely-perceptible title-color change, validated by replication, about $10M/yr; relocating Amazon's credit-card offer to the cart added tens of millions in annual profit.

The cost of not experimenting

The mirror image is just as real. Without experimentation, you ship on conviction — and conviction is expensive when it's wrong. Bing spent more than $25 million building a heavily anticipated social-integration feature that experiments revealed produced negligible value; without a controlled test, it would have been booked as a success. The ideas everyone is certain will win frequently don't, and shipping them blind is the most expensive form of "measurement."

This is the deeper cost of a "feature factory" that measures output instead of outcomes: budget pours into features that never move a metric, and no one finds out. When teams ship faster but don't learn faster, the waste never appears on a dashboard — but it's ROI leaking out the side.

Reframe ROI as expected value, not batting average

The mindset that makes this work is Jeff Bezos's. In Amazon's very first shareholder letter he framed failure and invention as "inseparable twins" — arguing that if you know in advance something will work, it isn't an experiment. Years later, in the 2015 letter, he sharpened the economics: given a 10% chance of a 100x payoff, you should take that bet every time, even though you'll be wrong nine times out of ten. That's fat-tails economics stated as a leadership principle — judge a program by its expected value, not its win rate.

But ROI isn't automatic — it scales with program maturity

Here's the catch the research is equally clear about: the favorable economics only materialize if the program is run well. Most companies never get there — as Stefan Thomke found, rather than running thousands of impactful tests, many firms "run no more than a few dozen a year that have little impact." Three things separate the high-ROI programs:

  • Trustworthiness. If your wins don't replicate, you're not banking returns — you're shipping noise and celebrating it. Guardrails, a clear success metric, and A/A tests are what make a reported ROI real, which is why avoiding the analytics mistakes that kill experimentation is a financial concern, not just a statistical one.
  • Honest measurement. The ROI you can actually bank is the incremental one — the lift your change truly caused, not what last-click attribution credits it. Reported returns almost always overstate reality, which is why measuring true incrementality matters before you claim a number.
  • Velocity and culture. Because the returns live in the tails, your odds of finding one scale with how many trustworthy tests you run — "triple your experiment rate and you triple your success rate," as Kohavi puts it. Booking.com's rule that anyone can test anything, without sign-off, is how that velocity is achieved. Researchers even map programs along a maturity curve — Crawl, Walk, Run, Fly — where ROI rises at each stage as experimentation goes from occasional to the default for every change.

As Expedia Group's former CEO put it, in an increasingly digital world, companies that don't experiment at scale don't survive.

The takeaway

The ROI of experimentation isn't a story about a high success rate — it's a story about cheap options on a fat-tailed upside. Most of your tests should, and will, fail. The rare winners you couldn't have predicted more than pay for the rest. The only ways to break the math are to make your tests expensive, or to run them so poorly you can't trust the wins. Get those right, and experimentation is one of the highest-return activities a product or growth team has.

See what your experimentation program is actually returning

Most teams have no honest read on whether their program is banking real returns or just running tests. Start with our free Experimentation Gap Assessment for a quick diagnosis of where yours stands — and if you want a deeper look at whether your program is set up to capture real ROI, our Experimentation Readiness Audit maps exactly what to fix first.

Book a call with our team →

Frequently asked questions

Does A/B testing actually have a positive ROI?
Yes, when run well — but not because most tests win. Research on the economics of experimentation shows idea quality is "fat-tailed": on Bing's platform, the top 2% of ideas produced roughly three-quarters of the gains. Because the marginal cost of an online test is near zero, running many cheap experiments to catch those rare winners is a high-return strategy.

Why do most experiments fail, and is that a problem?
It's normal and expected. In well-optimized products, only 10–20% of changes improve the target metric; across a broad portfolio, roughly a third win, a third are flat, and a third are negative. That's not a broken program — it's the portfolio working. The return comes from expected value across all tests, not from a high individual win rate.

What's the biggest hidden cost of not experimenting?
Shipping the wrong things with confidence. Without tests, teams invest heavily in ideas that feel certain but move nothing — like the $25M+ feature one company built and then abandoned. That wasted spend rarely shows up on a dashboard, which is what makes it so dangerous.

How do you actually measure the ROI of an experiment?
Focus on incremental impact — the lift your change genuinely caused versus what would have happened anyway — rather than the numbers attribution credits to it. Reported returns almost always overstate real ones, so trustworthy, incremental measurement is what turns a claimed ROI into a bankable one.

Does more experimentation always mean more ROI?
Only if the tests are trustworthy. More volume raises your odds of catching a fat-tail winner, but volume built on false positives and un-replicable "wins" destroys value instead of creating it. Velocity pays off only when it's paired with trustworthiness and program maturity.

Related articles

Quick Tip
5min

Why Experimentation Programs Fail to Scale (and What Actually Works)

More shipping isn't more learning. The feature-factory trap, why AI makes it worse, and what actually scales an experimentation program.
Guide
5min

How to Measure Impact When You Can't Run a Clean A/B Test

No clean A/B test? Triangulate. How to combine lightweight tests, before/after data, analytics, and qual into a credible read.
Deep Dive Article
5min

The Four Metric Buckets Every Experiment Needs (Primary, Secondary, Guardrail, Learning)

A framework from our experimentation webinar: the four metric buckets that turn messy test results into clear ship/no-ship decisions.

Get in touch!

Adasight is your go-to partner for growth, specializing in analytics for product, and marketing strategy. We provide companies with top-class frameworks to thrive.

Gregor Spielmann adasight marketing analytics