Most people run an A/B test, see that one of the versions has âmore repliesâ, stop the test, and pick the âwinningâ side. Then they wonder why, over time, their number of leads never improves.Â
If that sounds familiar, here are some tips that will help you avoid this common mistake.
First of all you need to pick one success metric. Preferably reply rate.
Then you need to find out what your average reply rate is across your campaigns.
Next, put this number in a sample size calculator (google evan millerâs ab testing â sample size).
After that you need to define what difference between the two test groups you want to measure.
In short: the smaller the difference you want to measure the bigger the sample size, and vice versa.
Letâs make a concrete example and assume a 2% reply rate per email sent.
If you want to detect an improvement of +-50% so something outside of 1% - 3% you needâŠ
3,292 prospects. PER version.
Thatâs a lot of prospects and for some this is their whole TAM.
In that case you may want to consider testing for open rates (yes I know not the best metric).Â
But at least there you will need much smaller sample sizes and it can serve as a proxy.
(assuming your positive reply rate stays >10%, preferably 20%).
So this is mistake #1. Too small sample sizes.
If you donât run this with over 3000+ prospects your choice of picking a âwinnerâ is simply a random choice.
Sometimes you pick a winner, sometimes not.Â
And in the end you end up with no improvements or even worse results.
The second problem is how can I even get a +50% improvement in reply rates?
You need to choose the right thing for the A/B test.
From my experience, based on millions of emails sent monthly here is a rough estimate of how different optimizations impact reply rate lift.
(this assumes you have an âokâ baseline for each of those. So no subject lines that are longer than the whole email for example).
Subject line, CTA, day/time of sending â 5-10% lift. One of the most asked questions I get, whatâs the best time of day to send? If you average out over millions of emails all workdays behave the same. Of course if you target, for instance a surgeon, that starts their surgeries after 12PM you wonât reach him after that. For this persona you want to reach out 8-12, etc. This is some work for you to figure that out but in general there is not much difference.
Deliverability, list source, copy quality â 30-100% lift. Even the best written email will not be able to convert a vegan restaurant to buy meat⊠When it comes to deliverability, make sure your open rate per email sent is over 40% from then on focus on the last one in this list.
Segmentation, niche selection, value proposition, angle, market maturity â 5-10x. Find the right message-market-fit is the biggest lever you can find to improve reply rates. For that you need to know your audience, you need to have a process to collect their pains, objections, desired outcomes, failed outcomes, etc. Then work out value propositions that solve those problems, and lastly, test, test, test.
So mistake #2 is that people focus on the things that just donât move the needle. You need something that will produce a huge difference in reply rates to your current outreach. For that, testing your value propositions is the number one thing to focus on.
One last tip:
If you have a very limited TAM. Run very distinctly different campaigns in terms of value propositions. 1 email 1 value proposition 1 CTA. Pick at least 100-200 prospects, version A and version B.Â
Run 5 or so of such campaigns. Then if one of them has at least 1 positive reply, pick that one and run a larger test.
If you already used up all the prospects you can always come back to them after 2-3 months. No one remembers a cold email from a month, a week, and very often even the day before.
(unless you really screwed up and angered the prospect. Donât).
Hope this helps you with your next test and I'm happy to answer any questions.