A/B testing

Definition
Showing two versions of one thing to separate, random slices of an audience, and judging which wins on a number that matters.

Why it matters

Without a test, a redesign or a new subject line is judged on opinion, and a page that looks better can easily convert worse. The random split is what makes the result trustworthy. Because both groups come from the same audience at the same time, a difference between them can be put down to the one change, not to who happened to see it or to what week it was.

How to apply it

  • Write the hypothesis and the deciding metric before starting: what changes, and which number says it worked.
  • Change one element at a time. If the headline, button and form all change together, a win cannot be traced to any of them.
  • Split the audience randomly and run both versions at the same time.
  • Decide the sample size or the run length up front, then wait for it. Check statistical significance before calling a winner.
  • Record every result, including the losses, so the next test builds on what was learned.

What it is

Version A is what you run today, usually called the control. Version B changes exactly one thing. The audience is split at random, so half see A and half see B, and after enough people have seen each, the two are compared on one number chosen in advance. A newsletter sender might test two subject lines and compare open rates. A shop might test a short product page against a long one and compare purchases.

Common mistakes

  • Stopping the moment one version pulls ahead. Early gaps often close with more data.
  • Testing a tiny audience. A list of a few hundred people cannot reveal a small difference, so test a bolder change or skip the test.
Worked example

Suppose a twelve-person B2B services firm wants more demo requests from its homepage. The current page converts about 2 in every 100 visitors, and the marketing lead believes a shorter headline would help. The team sets up a test in VWO: version A is today's page, and version B changes only the headline. Visitors are split at random, and the deciding number, demo requests per visitor, is fixed before the test starts. After four weeks each version has had 5,000 visitors. Version B converts at 2.4 per cent against 2.0 per cent for A. Before calling a winner, the team checks that the gap is larger than chance would produce, then records the result in a shared log, win or loss, for the next test.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    Control group

    The unchanged version a test is measured against.

  2. Article

    Statistical significance

    The check that a result is real, not chance.

  3. Article

    P-value

    The number a significance test produces.

Where it shows up