41 Shades of Blue and a $100 Million Headline: What the World's Best Companies Know About Testing
The short version: The biggest lesson from the world's most data-driven companies is that experimentation beats intuition. Even at elite firms, only about a third of ideas improve the metric they were built to move — and the winners are nearly impossible to predict in advance. So the advantage goes to whoever can run the most trustworthy tests, cheaply and fast: Booking.com runs tens of thousands a year, and a single overlooked experiment at Bing turned out to be worth more than $100 million. Test more ideas, overrule the highest-paid opinion, and treat failed experiments as the price of finding the rare, outsized wins.
A $100 million idea nobody wanted to build
In 2012, a Microsoft employee suggested a small change to how Bing displayed its ad headlines. It was cheap — a few days of engineering — but it was one of hundreds of ideas in the queue, so it sat untouched for more than six months. When an engineer finally ran it as an experiment, the results tripped a "too good to be true" alarm. The tweak had lifted revenue by about 12% — more than $100 million a year in the US alone — with no harm to user experience. It became the single most valuable idea in Bing's history (Harvard Business Review).
The uncomfortable question: if Microsoft's own experts couldn't spot a nine-figure idea sitting in their own backlog, what makes any of us think we can rank our ideas by gut?
Most of your ideas are wrong — and that's the point
The data is humbling. When Microsoft evaluated well-designed experiments meant to improve a key metric, only about a third actually did. Another third changed nothing, and a third made things worse. At a heavily optimized product like Bing, the win rate falls to 10–20%. Ronny Kohavi — who built experimentation programs at Amazon, Microsoft, and Airbnb — has recalled that of 250 ideas Airbnb tested, only around 20 moved the key metric. But those 20 were worth hundreds of millions.
A low success rate isn't a sign of a weak team, then. It's the normal physics of innovation. The teams that win aren't the ones with better instincts — they're the ones who reframe failure as learning and optimize for how fast they can learn.
Small changes, massive impact
If you can't predict winners, the logical move is to test more of them — and to drop the assumption that big bets beat small ones. The famous example is Google's "41 shades of blue": unable to agree on the right blue for its links, the team tested a spectrum of them to see which drove the most clicks. Google's lead visual designer, Douglas Bowman, left partly over that culture, worn down by being asked to justify a three-versus-five-pixel border. A Google UK executive later claimed the winning shade was worth around $200 million a year — a number best treated as an executive's estimate rather than an audited figure, but the direction is unmistakable: a color choice moved real money.
Velocity is the moat
Once you accept that winners are unpredictable and small changes matter, experiment volume becomes a strategy. Booking.com is the textbook case: it runs more than a thousand experiments simultaneously — on the order of 25,000 a year — and lives by one rule: anyone can test anything, without a manager's permission (Harvard Business Review). When a director once proposed a radical homepage redesign the CEO doubted, leadership didn't veto it — because vetoing would have broken the company's core tenet. They tested it. LinkedIn runs the same playbook at staggering scale, with tens of thousands of experiments live at once.
The lesson: tools don't create an experimentation culture. Permission does.
But volume without trust is just confident noise
Here's the catch. Running more tests only helps if you can believe the results — and getting a number is far easier than getting a number you can trust. Three traps sink most programs:
- Peeking. Checking a test over and over and stopping the moment it hits significance manufactures false winners. Decide your stopping point before you start — the statistics behind this are worth understanding.
- Novelty effects. A shiny new change can spike, then fade as the novelty wears off. One Microsoft test showed a 28% click lift that steadily decayed — a sign of confusion, not delight. Run tests a full week or more and watch the trend, so you don't ship a false winner.
- Twyman's Law. Any result that looks amazing is probably wrong. That "too good to be true" alarm on the Bing headline? A guardrail metric doing its job — validating a real win instead of shipping a bug.
The discipline that separates a serious program from a lucky one is trustworthiness: agree on one success metric, set guardrails, and stay suspicious of your own good news. (If you want a repeatable structure, we've written up an 8-step framework for reliable experiments.)
Kill the HiPPO
The most expensive bias in most companies has a name: the HiPPO — the Highest Paid Person's Opinion. The classic story is Amazon's. In its early days, engineer Greg Linden built shopping-cart recommendations; a senior VP was "dead set against it" and forbade further work. Linden tested it anyway. It won so decisively that not shipping it was costing Amazon real money — so it launched. The people closest to the problem are often right; the highest-paid person is often confidently wrong. Experimentation quietly replaces hierarchy with evidence.
Budget for failure
The mindset that ties it all together comes from Jeff Bezos, who has written that failure and invention are "inseparable twins" — that to invent, you have to experiment, and if you already know it'll work, it isn't an experiment (Amazon's 2015 shareholder letter). Amazon treats a string of failed bets as the cost of the occasional hundredfold win. Organizations that punish failed experiments are, whether they intend to or not, punishing invention itself.
The real takeaway
Come back to that Bing headline. The idea wasn't valuable because someone was brilliant enough to foresee it — nobody was, for six months. It won because someone finally had the tools and the permission to test it cheaply. That's the whole game: not better guesses, but more shots on goal, measured honestly. Triple your rate of trustworthy experiments and you'll triple both your failures and your wins — and it's the wins that compound.
Build the engine, not just the occasional test
Most teams don't have an idea problem — they have a testing-velocity and trustworthiness problem. If you want to build an experimentation practice that ships tests you can actually believe, consistently, our Experimentation Programs are built for exactly that.
Frequently asked questions
What percentage of A/B tests actually succeed?
At Microsoft, only about a third of well-designed experiments improved their target metric — a third did nothing, and a third made things worse. At highly optimized products like Bing, the success rate drops to 10–20%. A modest win rate is normal, even at elite companies.
What was Google's "41 shades of blue" experiment?
Unable to agree on the best blue for its links, Google tested a range of shades to see which drove the most clicks — a now-famous example of data settling a design debate. A Google UK executive later claimed the winning shade was worth roughly $200 million a year in ad revenue, though that figure is an executive estimate rather than an audited number.
What does HiPPO mean in experimentation?
HiPPO stands for the "Highest Paid Person's Opinion" — the tendency for decisions to be made by seniority rather than evidence. A core benefit of experimentation is replacing that opinion-based veto with data, letting the people closest to the problem prove their case.
Why do so many A/B test "wins" fail to hold up?
Usually because of peeking (stopping a test the moment it looks significant), novelty effects (a temporary bump that fades), or samples too small to trust. Deciding your stopping rule up front, running tests for a full cycle, and using guardrail metrics are what make results reliable.
How many experiments do top companies run?
A lot. Booking.com runs on the order of 25,000 experiments a year with more than a thousand live at once, and companies like LinkedIn, Google, and Microsoft each run well over 20,000 a year. High volume — run trustworthily — is the defining trait of mature experimentation cultures.
Sources: Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments (2020); Kohavi & Thomke, Harvard Business Review (2017); Thomke, Harvard Business Review (2020); Amazon shareholder letters (2015, 2018); Douglas Bowman, "Goodbye, Google" (2009); The Guardian (2014); Greg Linden, "Early Amazon" (2006).





