Statistical significance
Why it matters
Without the check, businesses roll out changes based on results that were never real. A small sample can look decisive purely by luck: three replies against one from a tiny list tells you almost nothing. The opposite risk is moving too slowly, so the bar should match the decision. A cheap, reversible change can go ahead on weaker evidence than an expensive one that will be lived with for years.
How to apply it
- Decide the sample size and the decision rule before the test starts, and run to that size instead of judging by eye halfway.
- Do not stop a test the first time it crosses the threshold. Repeated peeking inflates false positives.
- Read significance together with effect size. A result can be statistically solid and still too small to act on.
- Treat patterns in data that was never set up as a test with the same caution. A trend across a handful of sessions is a hunch.
- Make sure the test has enough statistical power, or a non-significant result tells you little.
What it is
Whenever two options are compared, such as two subject lines or two landing pages, the numbers almost never come out identical, even when both options are equally good. Chance alone creates a gap. Statistical significance is a way of asking how likely it is that a gap this big would appear by luck if the two options were really the same.
The usual convention is a threshold of 5 percent, written as a p-value below 0.05. Passing it means the gap is unlikely to be chance. It does not prove the change caused the gap, and it says nothing about whether the gap is big enough to matter.