Confidence interval

Definition
A confidence interval is the range a true result probably falls in, given that only a sample was measured rather than everyone.

Why it matters

A single number looks exact and is not. An interval shows how much weight that number deserves. If the range crosses zero, the data cannot rule out that the change did nothing, or made things worse. A team that only reads the headline lift treats a lucky sample as a proven fact and then rolls it out to everyone.

Width depends mainly on how much data there is. Roughly four times the data halves it.

How to apply it

  1. Quote every lift with its interval. Write "+1.0 points (95 per cent interval: -0.8 to +2.8)", not just "+1.0 points".
  2. Treat any range that includes zero as unproven. The data cannot yet rule out no effect, or a harmful one.
  3. Decide the sample size before the test starts. Use a sample size calculator with the smallest lift worth acting on, and run until you reach it.
  4. Wait for more data when the range is wide. Roughly four times the sample halves the width, so a range of plus or minus 2 points needs about 16 times the data to reach plus or minus 0.5.
  5. Judge by the worst plausible end. If the lowest value in the range would still justify the change, ship it. If not, either collect more data or accept the risk knowingly.
  6. Do not peek and stop. Check at the planned end date, not every day, or false wins creep in.

What it is

Every number measured from a sample is an estimate. A confidence interval puts a range around that estimate. Say a new email subject line gets a 5.0 per cent click rate against 4.0 per cent for the old one. The measured lift is one point, but the interval might read from 0.8 points worse to 2.8 points better. The data cannot yet tell those outcomes apart.

The usual 95 per cent interval has a precise meaning. If the same test were repeated many times, about 95 in 100 of the intervals built this way would contain the true value. In daily use, read it as the range of results the data can reasonably support.

Common mistakes

  • Reading a 95 per cent interval as a 95 per cent chance that the truth is inside it. It describes the method, not this one result.
  • Stopping a test the first day the range clears zero. Checking repeatedly and stopping on a good moment inflates false wins.
Worked example

Suppose a consultancy tests two versions of its landing page. Over a fortnight, version A converts 4.0 per cent of 1,000 visitors and version B converts 5.0 per cent of the same number. The one-point lift looks like a win, but the interval runs from about 0.8 points worse to 2.8 points better, so the data cannot rule out no change. The team keeps the test running in VWO, which runs A/B tests on websites and tracks conversion, until each version has about 10,000 visitors. At that size the same rates give an interval of roughly 0.4 to 1.6 points better, and the change is almost certainly helping. These figures are illustrative. Every reported lift now carries its interval, so nobody rolls out a lucky sample to the whole site.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    Statistical significance

    The yes or no question of whether an effect is real.

  2. Article

    Sample size

    The main lever on how narrow the range gets.

  3. Article

    Statistical power

    Whether a test can detect an effect at all.

  4. Article

    A/B testing

    Where most business intervals come from.

Where it shows up