P-value
Why it matters
Random variation looks exactly like a win. Run enough small experiments and one will look positive purely by luck. A p-value stops a business rolling out noise as if it were a discovery.
It answers a narrower question than most people think. It does not give the probability that the change works, and it does not say how large or how valuable the effect is. A tiny lift can have a very small p-value on a large audience and still be worth nothing.
How to apply it
- Decide the sample size and the significance threshold before starting. Statistical power explains how big the sample must be.
- Run the test to its planned end. Stopping the moment it looks good makes a fluke far more likely.
- Read the p-value together with the size of the effect and a confidence interval.
- Treat a result just under the threshold as weak evidence.
What it is
Every test starts with a default assumption, called the null hypothesis, that the change makes no difference. The p-value asks how surprising the data would be if that were true. A p-value of 0.02 means that, if the change did nothing, a gap this large or larger would still appear by chance about two times in a hundred.
By convention, many teams treat a p-value below 0.05 as statistically significant. That threshold is a habit, and it should be agreed before the test runs.
Common mistakes
- Reading a p-value of 0.03 as "a 97 per cent chance this works". It does not mean that.
- Testing ten things and celebrating the one that crossed the line.
- Treating a large p-value as proof of no effect. It may only mean the sample was too small to tell.