Confidence interval
Why it matters
A single number looks exact and is not. An interval shows how much weight that number deserves. If the range crosses zero, the data cannot rule out that the change did nothing, or made things worse. A team that only reads the headline lift treats a lucky sample as a proven fact and then rolls it out to everyone.
Width depends mainly on how much data there is. Roughly four times the data halves it.
How to apply it
- Quote every lift with its interval. Write "+1.0 points (95 per cent interval: -0.8 to +2.8)", not just "+1.0 points".
- Treat any range that includes zero as unproven. The data cannot yet rule out no effect, or a harmful one.
- Decide the sample size before the test starts. Use a sample size calculator with the smallest lift worth acting on, and run until you reach it.
- Wait for more data when the range is wide. Roughly four times the sample halves the width, so a range of plus or minus 2 points needs about 16 times the data to reach plus or minus 0.5.
- Judge by the worst plausible end. If the lowest value in the range would still justify the change, ship it. If not, either collect more data or accept the risk knowingly.
- Do not peek and stop. Check at the planned end date, not every day, or false wins creep in.
What it is
Every number measured from a sample is an estimate. A confidence interval puts a range around that estimate. Say a new email subject line gets a 5.0 per cent click rate against 4.0 per cent for the old one. The measured lift is one point, but the interval might read from 0.8 points worse to 2.8 points better. The data cannot yet tell those outcomes apart.
The usual 95 per cent interval has a precise meaning. If the same test were repeated many times, about 95 in 100 of the intervals built this way would contain the true value. In daily use, read it as the range of results the data can reasonably support.
Common mistakes
- Reading a 95 per cent interval as a 95 per cent chance that the truth is inside it. It describes the method, not this one result.
- Stopping a test the first day the range clears zero. Checking repeatedly and stopping on a good moment inflates false wins.