Hypothesis testing

Definition
Hypothesis testing is the practice of stating a specific, falsifiable prediction and running a controlled experiment to see whether the evidence supports it.

Why it matters

Opinions about what will work are cheap and often wrong. A test turns the question into evidence. It also forces a team to agree on what counts as success before the data arrives, which stops a weak result from being reframed as a win afterwards.

It matters most when being wrong is expensive. A small test before a large build can save months spent on the wrong thing.

How to apply it

  • Write the hypothesis in one sentence: if X is changed, then Y will happen, measured by Z.
  • Pick one main metric. Several metrics make it easy to find something that looks like a win.
  • Decide the sample size and the test length before starting, and do not stop early because the numbers look good.
  • Keep a control group that sees the original version, so the result is a comparison and not a before-and-after guess.
  • Change one thing at a time where possible, so the result has one explanation.
  • Record every test, including failures, so the same idea is not tested twice by someone who forgot.

What it is

A hypothesis is a prediction that can turn out wrong. "A shorter sign-up form will increase trial starts" is a hypothesis. "Let's improve the form" is not, because nothing could prove it false.

In a formal test, the team also states the default assumption that nothing changed, called the null hypothesis. The experiment then asks whether the data is strong enough to reject it. A result is judged with statistical significance, commonly shown as a p-value.

Common mistakes

  • Peeking at the results daily and stopping at the first good-looking number.
  • Reading "no significant difference" as proof that there is no effect. It may only mean the test was too small to see one.
Worked example

Suppose a B2B software firm believes its sign-up form asks for too much. The hypothesis is written as one sentence: if the form drops from seven fields to three, trial starts will rise, measured by trial starts per hundred visitors. The team picks that single metric and fixes the test length at two weeks before launch. It runs the test in VWO, which shows half the visitors the original form, the control group, and half the short version. Say the original converts at 4.0 per cent and the short form at 4.6 per cent. The team checks whether that gap is larger than normal variation before trusting it. It does not stop early because the first three days look good. The gap holds to the end, so the short form ships and the result is recorded for the next test.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    A/B testing

    The most common way to run a hypothesis test on a website or email.

  2. Article

    Minimum viable test

    A small, quick version of the same discipline.

  3. Article

    Heatmap

    A source of ideas worth testing.

Where it shows up