Control group

Definition
The slice of people deliberately left unchanged during a test, so a result can be compared against what would have happened anyway.

Why it matters

Without a control, a lift after a change could be a strong week, a seasonal swing, a price promotion or an unrelated launch. Crediting the change for something else wastes the lesson the test should teach. A control group is what moves a result from correlation towards cause. It matters most on long or costly tests, where a wrong conclusion costs weeks, not an afternoon.

How to apply it

  • Assign people to groups at random. Picking the groups by hand lets bias in.
  • Make the control truly comparable: same source, same segment, same time window, not a different list from a different month.
  • Change one thing at a time, or the result cannot be traced.
  • Run the test through ordinary week-to-week swings. Two to four weeks is a sensible floor for most B2B tests.
  • Wait for statistical significance before declaring a winner.

What it is

A control group is the baseline in an experiment. One group sees the new thing: a different subject line, a new landing page, an extra follow-up email. The control group sees the usual version or nothing at all, in the same period and from the same source. Any gap that appears between the two groups can then be put down to the change, because everything else was shared.

Common mistakes

  • Comparing this month's results with last month's and calling last month the control.
  • Ending the test the moment the test group pulls ahead. Early leads often reverse.

Types of control

  • No change: the group gets nothing new, which shows the effect of doing something at all.
  • Current version: the group keeps the existing page or email, which is the normal A/B testing set-up.
  • Holdout: a small share of an audience is left out of a whole programme, such as retargeting ads, to measure what it adds.
Worked example

Suppose a B2B software company tests a new pricing page against the current one. Its paid traffic is split at random by VWO: half the visitors see the existing page and half see a version with a shorter plan comparison. Both groups come from the same ad sets over the same three weeks, so the two results can be compared directly. Say the control converts at 2.4 per cent and the new page at 2.9 per cent. The gain can be credited to the page because nothing else differed between the groups. Had the team compared the new page with last month's figures instead, a price rise or a quiet week could have produced the same lift, and the company would have rolled out a change that did nothing.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    A/B testing

    The wider practice a control group belongs to.

  2. Article

    Statistical significance

    Whether the gap is real.

  3. Article

    P-value

    The number a significance test produces.

  4. Article

    Sample size

    How many people each group needs.

Where it shows up