On this article

The Four Metric Buckets Every Experiment Needs (Primary, Secondary, Guardrail, Learning)

A framework from our experimentation webinar: the four metric buckets that turn messy test results into clear ship/no-ship decisions.
This is some text inside of a div block.

The Four Metric Buckets Every Experiment Needs (Primary, Secondary, Guardrail, Learning)

At our recent webinar, Adasight x Cherto: 5 Lessons Learned from 1,000s of Experiments, experimentation expert and Adasight advisor Dr. Simon Jackson took a sharp question from the audience: when you run tests across several product areas, you end up watching a lot of metrics — and the more metrics you watch, the more likely some drift by chance, and the greater the risk of shipping a false positive. So how do you keep a clean read on impact and catch one team's win quietly hurting another's numbers? His answer was one of the most practical frameworks of the session.

Prefer to watch it?

The short version: To keep experiments trustworthy — especially when you run many at once — sort each test's metrics into four buckets. A primary metric (the single decision-maker), secondary metrics (behavioral signals close to the change), guardrail metrics (things that must not get worse — including other teams' primary metrics, which is how you catch cannibalization), and learning metrics (exploratory signals). Then set one rule: ship only if the primary is positive and no guardrail goes negative. Every test gets an unambiguous decision, and no team's win silently costs another.

The problem: more metrics, more ways to be wrong

The instinct when you run experiments across multiple product verticals is to track everything. But every metric you add is another chance for a random up-or-down movement to look meaningful — a false winner waiting to happen. Worse, across verticals you face cannibalization: a change that lifts your numbers may be stealing from another team's. Watch too little and you miss it; watch too much and you drown in noise.

Option 1: the statistical fix (and why it isn't your first move)

There's a rigorous statistical answer — Family-Wise Error Control, which corrects for the inflated false-positive rate you get when testing many metrics at once. Good experimentation platforms bake it in. But as Simon noted, it comes with a cost: it reduces your statistical power on the metrics you actually care about, and it's genuinely statistician territory. It's worth knowing the statistics underneath, but for most teams it isn't the everyday tool.

Option 2: metric buckets (the framework to actually use)

The more useful discipline for most teams is to mindfully sort every experiment's metrics into categories, each with its own job and its own decision rule. Simon's minimum set is four buckets.

Primary — the one metric that decides

Typically a single metric: the one that determines whether you ship. Ideally it's a shared business metric — if you have the traffic and power to move it. If you don't, it's the metric for your specific product vertical. The discipline is having one decision-maker, not five. (If you're not sure how to choose it, our 8-step framework for reliable experiments walks through it.)

Secondary — did customers react the way you expected?

Behavioral metrics that sit close to the change you made. Their job is to confirm the mechanism — that users are actually responding the way your hypothesis predicted, not that the primary moved for some unrelated reason.

Guardrail — what must not get worse (and how you catch cannibalization)

Metrics that should not go negative. This is the bucket that does the heavy lifting on cannibalization: as Simon put it, if you have other product verticals, put their primary metrics into your guardrails. Now, if your win comes at their expense, the guardrail flags it before you ship.

Learning — the exploratory signals

Metrics you're curious about and want to learn from, but won't make a ship decision on. Keeping them out of the primary bucket stops them from muddying the call.

The payoff: one clear decision rule

The reason to bucket metrics isn't tidiness — it's that each bucket gets its own decision criteria. That collapses a messy dashboard into a single rule: ship if the primary is positive and none of the guardrails are negative. No debating which of eight metrics "really" mattered after the fact. The decision was defined before the test ran.

Stop lumping everything into "primary"

The most common mistake — and the realization the audience member had live on the call — is that teams have primary, secondary, and guardrail buckets, but then dump every high-level vertical metric into "primary." That's what makes results ambiguous. Moving other verticals' metrics into the guardrail bucket is exactly what turns a confusing set of ups and downs into a clean, cannibalization-aware decision.

The final piece is alignment: agree with the other teams in your business on what the shared guardrails should be, so everyone's experiments protect the same things. (There's a deeper world here — different statistical tests per bucket, non-inferiority tests — but the hygiene of clear buckets and shared guardrails already puts you ahead of most teams.)

Watch the full session

This was one moment from a session packed with hard-won lessons from thousands of experiments.

Watch the full webinar recording

Where does your experimentation program actually stand?

As Simon noted at the close, most teams have some structure but haven't pressure-tested whether it holds up at scale. Start with our free Experimentation Gap Assessment for a quick read on where your biggest gaps are — and if you want a deeper, hands-on diagnosis, our Experimentation Readiness Audit maps exactly what to fix first.

👉Book a call with our team →

Frequently asked questions

What are guardrail metrics in experimentation?
Guardrail metrics are the things that should not get worse when you ship a change — like page-load time, or another team's key metric. They protect against unintended harm, and they're the main tool for catching cannibalization: if you include other product verticals' primary metrics as your guardrails, you'll see when your win comes at their expense.

How do you stop A/B tests from cannibalizing each other?
Put the other product verticals' primary metrics into your experiment's guardrail bucket, and set the rule that you only ship if your primary is positive and no guardrail goes negative. That way a change that lifts your numbers by stealing from another area gets flagged before it ships.

What's the difference between primary, secondary, and guardrail metrics?
The primary metric decides whether you ship (ideally just one). Secondary metrics are behavioral signals close to the change that confirm customers reacted as expected. Guardrail metrics are things that must not get worse. A fourth bucket — learning metrics — captures exploratory signals you won't decide on.

What is Family-Wise Error Control?
It's a statistical method that corrects for the higher false-positive rate you get when testing many metrics at once. Good experimentation platforms include it, but it reduces statistical power on the metrics you care about, so it's a more advanced tool than the everyday metric-bucketing approach.

How many primary metrics should an experiment have?
Ideally one. A single decision metric keeps the ship/no-ship call unambiguous. Additional metrics you care about belong in the secondary, guardrail, or learning buckets — not stacked into "primary," which is what makes results hard to read.

Related articles

Deep Dive Article
10min

41 Shades of Blue and a $100 Million Headline: What the World's Best Companies Know About Testing

Why velocity beats intuition in experimentation—real lessons from Bing's $100M test, Google's blues, and Booking.com's culture.
Quick Tip
5min

How to Fix Fragmented User Identity in Amplitude (When Users Don't Log In by Email)

Phone-number or non-email login can fragment users in Amplitude. Why identity breaks—and how to stitch sessions back together.
Guide
5min

How to Set Up Amplitude Guides & Surveys (Step-by-Step)

Set up Amplitude Guides & Surveys step by step: the three components, building a guide and survey, and targeting with cohorts.

Get in touch!

Adasight is your go-to partner for growth, specializing in analytics for product, and marketing strategy. We provide companies with top-class frameworks to thrive.

Gregor Spielmann adasight marketing analytics