Pick one email and one question
Start with the email that gets the most sends, usually the first or second email of your welcome flow or your booking nurture. More volume means a faster answer.
Write the question you want answered in one line. For example: "Does a specific subject line beat a curious one for the first nurture email?" If you cannot write the question in one line, the test is too big.
Write your guess before you send
Note which version you think will win and why. This takes thirty seconds and stops you from rewriting history afterwards. Over a few months it also shows you where your instincts are wrong, which is where the learning is.
Change only the subject line
Keep the sender name, preview text, body and send time identical. Change one thing. Good pairs to try:
- Specific against curious: "Three steps before our call" against "Quick thought on your onboarding".
- Short against long: under five words against ten to fourteen words.
- Question against statement: "Still comparing options?" against "How buyers compare options".
- With the first name against without it.
- Benefit against plain label: "Cut proposal time in half" against "Our proposal guide".
Run the same style of test on the same email more than once before you call it a rule. One win is a hint, not a law.
Judge on clicks, not only opens
Some mail apps load images automatically, which counts as an open even when nobody read the email. Open rates are therefore noisy and often too high. Use them as a first signal, but pick the winner on clicks, replies or bookings.
A subject line that wins on opens and loses on clicks has promised something the email did not deliver. Do not copy it.
Test send times one slot at a time
Pick two slots and keep everything else the same. A sensible first pair for B2B is Tuesday at 08:30 against Thursday at 14:00, in the recipient's time zone if your tool supports it. Split the list at random, never by region or by lead source.
In a nurture flow the send time is often relative to sign-up, for example two days after the previous email. Test the delay as well as the clock time. Two days against four days tells you more about momentum than 08:30 against 09:30.
Be honest about sample size
Small differences need large lists. As a rough rule, aim for at least 200 recipients per variant before you trust a gap of ten percentage points, and far more for a gap of two. If your list is small, test big, bold differences and let the test run across several weeks of new sign-ups.
Do not stop a test the moment one version leads. Decide the end date or the number of sends in advance, then read the result.
Keep a test log
Use a simple sheet in Google Sheets or Notion. One row per test, with these columns: date, email, question, version A, version B, sends per version, metric, result, decision. When a test ends, write what you will do next time in one sentence.
HubSpot, ActiveCampaign, Mailchimp, Brevo and Customer.io all support some form of A/B testing. Check what yours does before you plan a test. Some tools only test subject lines on a one-off send, not inside an automated sequence, and in that case you split the sequence by hand.
Roll the winner out and set a rhythm
When you have a repeatable result, apply it to the other emails in your sequences. If specific subject lines beat curious ones twice, rewrite the weakest subject lines in your other flows the same way.
Run one test a fortnight. Alternate between subject lines and timing so you do not change two things at once. The monthly numbers from Review open and click rates monthly tell you which email to test next.
Common mistakes
- Testing two things at once, such as a new subject line and a new send time.
- Picking a winner after a day because it looks ahead.
- Using opens as the only metric.
- Running tests on emails with 40 recipients and trusting the gap.
- Declaring a global rule, such as "Tuesdays are best", from one test on one audience.
- Never writing the result down, so the same test gets run again next year.
How you know it works
After a quarter you have six or more logged tests and at least two rules you have confirmed twice. Your weakest sequence email has a better click rate than it did, and you can say why. New emails start from a tested subject line style instead of a blank page.