EXPERIMENTATION GUIDE
The Pre-Launch Experiment Checklist
Most invalid or wasted experiments were doomed before they ever launched: no agreed success metric, an underpowered sample, broken tracking, no guardrails. This guide walks through the five things that make a test trustworthy by design: hypothesis quality, metric structure, statistical setup, tracking integrity, and stakeholder alignment. It ends in a full, printable checklist you can run through 24-48 hours before any test goes live.

The Cost of Skipping This
Picture a test that ran for two weeks, hit "significance" on day 4, and got shipped. Three months later, someone notices the metric it was supposed to lift hasn't actually moved. Digging in, the traffic split was never checked, the variant had been quietly getting 46% of traffic instead of 50%, a bucketing bug nobody caught. The "win" wasn't real. It was noise dressed up as a result, and it took a quarter for anyone to notice.
This is the default failure mode in most experimentation programs — not dramatic mistakes, but small, invisible gaps that turn a test into a coin flip wearing a lab coat. None of them show up in the results. They show up months later, as a metric that mysteriously never moved the way the "winning" test promised.
This guide, and the checklist at the end of it, exists to catch every one of these before launch, not after.
Why Download?
Intro: Get a technical breakdown of what actually makes an experiment trustworthy, not just a list of best practices, but the reasoning behind each one.
✅ Why a hypothesis without a stated mechanism can't teach you anything, win or lose
✅ How MDE and sample size actually relate, and what happens when nobody sets one
✅ What Sample Ratio Mismatch (SRM) is, why it happens, and why it invalidates an entire test
✅ The "peeking problem", why checking significance early inflates your false-positive rate
✅ A full, printable 18-item checklist covering hypothesis, metrics, statistics, tracking, and alignment

Who is this for?
- Growth & CRO Leads — You've shipped a "winning" test before, only to watch the metric it promised never actually move three months later. This gives you the exact five checks that would have caught it — before the launch, not after the postmortem.
- Experimentation Program Managers — Your team is scaling test velocity, and you're starting to worry that speed is coming at the cost of trust. This is the guardrail that lets you scale without quietly shipping garbage results.
- Data Analysts & Statisticians-in-Training — You know sample size and MDE matter, but explaining why to a stakeholder who wants to launch tomorrow is harder. This gives you the plain-language reasoning to back up the math.
- Product Managers Running Their First Tests — You're new to experimentation and don't yet have the instinct for what can silently go wrong. This is the checklist that replaces instinct with a repeatable process.
- Engineering & QA Leads — You've watched a "statistically significant" result get shipped, only to find later the tracking was broken the whole time. This gives you language and a checklist to insist on tracking validation before any test launches.