GTM Engineering

Outbound A/B testing

Outbound A/B testing means changing exactly one variable between two cold outbound versions, calculating the minimum sample size before the test starts, and picking a single win/loss metric in advance, instead of eyeballing a dashboard after a few days.

Most teams that think they're testing cold outbound are actually just changing something and watching what happens. Someone swaps a subject line, maybe the opener too, in the same send. A few days pass. Someone glances at reply counts, calls a winner, and rolls it out. Nothing about that sequence proves the winning version is actually better, because three decisions never got made in writing before the first email went out.

The three decisions that make it a real test

First, exactly one variable changes between the two versions, everything else, copy, list, send time, sequence step, stays identical. Second, the minimum sample size gets calculated, not guessed, using the baseline reply rate and the size of the lift worth detecting; smaller lifts need bigger samples, not smaller ones, because a smaller true difference is harder to tell apart from noise. Third, one metric gets named as the primary win/loss signal before anyone looks at a single reply, because reply rate, positive reply rate, and meetings booked will each tell a different story if all three stay in play.

A fourth failure sits on top of all three: peeking. Checking results daily and calling a winner the moment it looks good, before the calculated sample size or run window is reached, breaks a properly designed test just as thoroughly as never designing one in the first place. And the sample sizes involved are usually bigger than teams assume: detecting a 20% relative lift off a 3.43% baseline reply rate needs roughly 12,110 sends per variant at standard confidence and power settings, not the few hundred sends most teams assume is enough.

In practice

Paste a short checklist above the campaign before it launches: the variable being changed, the hypothesis, the primary metric, the minimum sample size per variant (calculated in a sample-size tool, not estimated), the minimum run window, and who signs off on the result. If a line can't get filled in, the test isn't ready to launch.

What people get wrong

Open rate gets treated as a clean, obvious metric to judge a test on. It isn't anymore. Apple's Mail Privacy Protection pre-loads images, tracking pixels included, on Apple's own servers before a human opens anything, which can push measured open rates toward 100% for Apple Mail recipients regardless of whether anyone read the email. The metric hierarchy that still holds up runs reply rate first, positive reply rate second, meetings booked third, picked in advance, not after someone's seen which number looks best.

Related terms
Where we use this
Updated July 25, 2026

Ready to engineer your GTM motion?

Tell us how your motion runs today. We'll show you what we'd engineer.

Contact us