P-value

Definition
The probability of seeing a result at least this extreme if the change being tested actually did nothing at all.

Why it matters

Random variation looks exactly like a win. Run enough small experiments and one will look positive purely by luck. A p-value stops a business rolling out noise as if it were a discovery.

It answers a narrower question than most people think. It does not give the probability that the change works, and it does not say how large or how valuable the effect is. A tiny lift can have a very small p-value on a large audience and still be worth nothing.

How to apply it

  • Decide the sample size and the significance threshold before starting. Statistical power explains how big the sample must be.
  • Run the test to its planned end. Stopping the moment it looks good makes a fluke far more likely.
  • Read the p-value together with the size of the effect and a confidence interval.
  • Treat a result just under the threshold as weak evidence.

What it is

Every test starts with a default assumption, called the null hypothesis, that the change makes no difference. The p-value asks how surprising the data would be if that were true. A p-value of 0.02 means that, if the change did nothing, a gap this large or larger would still appear by chance about two times in a hundred.

By convention, many teams treat a p-value below 0.05 as statistically significant. That threshold is a habit, and it should be agreed before the test runs.

Common mistakes

  • Reading a p-value of 0.03 as "a 97 per cent chance this works". It does not mean that.
  • Testing ten things and celebrating the one that crossed the line.
  • Treating a large p-value as proof of no effect. It may only mean the sample was too small to tell.
Worked example

Suppose an online accountancy tests a new colour for its book-a-call button. Before the test starts, the team agrees on a p-value threshold of 0.05 and to run until each version has 10,000 visitors. The test is set up in VWO, which splits traffic between the two versions.

At the end, the original converts at 3.0 per cent and the new button at 3.6 per cent. The p-value is 0.02, so a gap this large would appear by chance about two times in a hundred if the colour made no difference. The team also checks the size of the gain and its confidence interval, which is wide. It rolls out the change, but it does not promise a 20 per cent lift every month, because one test measures one period.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    Statistical significance

    The judgement a p-value feeds into.

  2. Article

    A/B testing

    The practice where p-values are used most.

  3. Article

    Sample size

    How many people the test needs.

Where it shows up