Sample size

Definition
Sample size is how many data points a decision is based on. Too few and the result you see may just be chance rather than a real pattern.

Why it matters

A small sample lies with confidence. Ten emails or five lost deals can show a striking pattern that disappears at fifty, because randomness has more room to look meaningful when there is little data to average it out. A rate of 10% from fifty emails could, by chance alone, really sit anywhere between roughly 4% and 21%.

Acting on a small sample means changing a template, a script or a price based on noise, then wondering why the improvement never repeats.

How to apply it

  • Before declaring a winner in any test, ask how many data points each side had.
  • Decide the sample before the test starts, not after peeking at early results.
  • For email and ad tests, count responses per variant, not only sends or opens.
  • Where volume is naturally low, run the test for longer instead of in one burst.
  • Look for the same pattern across several batches rather than trusting one.
  • Remember the maths: to halve the margin of error, the sample must be about four times larger.

What it is

Sample size is how much data sits behind a number. If five of fifty emails got a reply, the sample is fifty and the reply rate is 10%. If five hundred of five thousand got a reply, the rate is the same but the evidence is far stronger.

Think of flipping a coin. Seven heads in ten flips is unremarkable and says nothing about the coin. Seven hundred heads in a thousand flips would be hard to put down to luck. The proportion is identical. Only the sample differs.

Common mistakes

  • Stopping a test the moment one side leads. Early gaps are mostly noise. Set the sample first and wait for it.
  • Counting sends instead of responses. A thousand emails with five replies is a sample of five where it matters.
  • Pooling different audiences. A result from 200 CEOs and 200 junior staff mixed together describes neither.
  • Testing many variants on a small list. Each variant gets only a thin slice of the data.
  • Treating a big sample as a cure for a biased one. Large volume from the wrong audience still gives the wrong answer.
  • Confusing a small effect with a real one. A tiny gap needs a much larger sample to show up.
Worked example

Suppose a B2B consultancy tests two subject lines on its outbound emails. Version A goes to 50 prospects and gets 5 replies, a rate of 10%. Version B gets 7 replies, which looks like a 40% lift on paper. With so few replies, the gap could easily be chance, so the team sends the test again to 500 prospects per version. The rates settle at 10% and 11%, a difference too small to justify changing the template. The same discipline applies on a website, where a tool such as VWO can split traffic between two landing pages. The team fixes the number of visitors in advance and reads the result only once that number has arrived, not the first time one version pulls ahead.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    Statistical significance

    The test of whether a difference is likely real.

  2. Article

    Statistical power

    The chance that a test of a given size will detect a real effect.

  3. Article

    Confidence interval

    The range that shows how uncertain a rate is.

  4. Article

    Booking rate

    Easy to misread from a small batch of leads.

Where it shows up