A/B Testing

A/B testing is an experiment that compares two versions of a message, page or feature by showing each to a random share of the audience and measuring which performs better on a chosen metric. In email, it is used to test subject lines, copy and calls to action.

Cold Email & OutreachUpdated September 30, 2026

In short

A/B testing sends two versions to similar groups and keeps the one that measurably performs better.

Key points

  1. A/B testing is a randomized controlled experiment with two variants, A and B [1].
  2. Test one variable at a time, pick the winning metric in advance and split recipients randomly [2][3].
  3. Small samples give noisy results; use a sample size calculator to see how many sends each variant needs [4].
  4. For Cold Email, judge tests by Reply Rate or Positive Reply Rate, not Open Rate, which privacy features distort.
  5. Common tests include the Subject Line and Preview Text, the Icebreaker, the Value Proposition and the Call to Action (CTA).

How an A/B test works

An A/B test starts with a hypothesis, such as "a question in the subject line will get more replies". You create two versions that differ only in that element, randomly split your audience and send each group one version. After enough sends, you compare results on the metric you chose beforehand. Wikipedia describes A/B testing as a form of randomized controlled experiment, and notes that with more than two variants or several changing elements it becomes multivariate testing [1]. Optimizely's glossary stresses defining the goal metric before the test starts [2], and Mailchimp's A/B tool lets you choose whether opens, clicks or revenue decide the winner for campaigns [3].

Sample size and significance

The most common mistake is calling a winner too early. With small numbers, random variation looks like a real difference. If variant A gets 6 replies from 100 emails and B gets 9, that gap is well within chance. Evan Miller's sample size calculator shows how many sends per variant are needed to detect a given lift with reasonable confidence [4]; for low baseline rates, such as a 3% reply rate, detecting a modest improvement can take thousands of sends per variant. Cold outreach teams with smaller volumes can test bolder changes, which produce larger effects, or run tests over several weeks. Keep the audience comparable, since a test across different segments measures the segment, not the message.

What to test in cold email

Test the elements with the biggest expected effect first. Targeting often matters most, such as comparing two segments of your Ideal Customer Profile (ICP), but that is a list test, not a copy test, so run it separately. For copy, the opening line and the Call to Action (CTA) tend to matter more than small wording changes elsewhere. Subject line tests are popular but only affect whether the email gets read, so measure them by replies. Record every test, its result and the sample size in one place, so learning accumulates. Retest winners occasionally, because what works changes as markets and inboxes change. Avoid testing tricks that could raise complaints.

Sources
  1. A/B testing — Wikipedia
  2. A/B testing — Optimizely Optimization Glossary
  3. About A/B Tests — Mailchimp
  4. Sample Size Calculator — Evan Miller
External sources open in a new tab.

Related terms

Mentioned in

Outreach without the busywork.

PineLead finds new B2B prospects every day, qualifies them against your criteria and writes the first email in your voice. You approve — PineLead sends.

Start free with 100 credits →