Six versions of the same email, and no winner
Open rates across six versions ran from 8.0% to 41.4% on 130 delivered emails. The 41.4% arm has three opens we can date. Zero clicks, zero replies, six arms still tied.
A five-fold difference that means nothing
For the last fourteen days our outreach has been sending six different versions of the same email to independent hostels in Mexico and Colombia. Same list, same send window, same offer. Only the wording changes: one leads with money, one with OTA commission, one with how little of your time the work takes, one asks a question, one offers a trial, one offers a rate check.
This is what the scoreboard said this morning, across 130 delivered emails:
- trial — 29 delivered, 12 opens, 41.4%
- rate check — 12 delivered, 3 opens, 25.0%
- OTA commission — 17 delivered, 4 opens, 23.5%
- time — 23 delivered, 5 opens, 21.7%
- question — 24 delivered, 3 opens, 12.5%
- money — 25 delivered, 2 opens, 8.0%
Top to bottom that is 41.4 against 8.0, a bit over five times. If somebody put that table in front of you, the decision would look like it had already been made: keep the trial version, delete the other five, stop wasting sends on losers. We did not do that, and this post is why.
Count what an open actually is
The winning arm has twelve opens on record. Split by when they happened relative to the delivery: five landed within two minutes of it, three carry a timestamp well after it, and four cannot be matched to a delivery record at all.
On this list, an open logged two minutes after delivery has not been a person. It is the receiving mail host fetching the images in the message before the inbox has been touched, and a tracking pixel is an image. We saw the same thing on the click side a fortnight ago: fourteen clicks out of three mailboxes that never opened the message.
The four unmatched opens are an open event with no delivery to join to, so we cannot place them in time. They may be real people; we cannot show they read the mail rather than that a scanner touched it, so they decide nothing.
That leaves three opens on the winning arm that we would stand behind. Three. Re-rank all six on dated, later opens only:
- rate check — 2 of 12, 16.7%
- time — 3 of 23, 13.0%
- OTA commission — 2 of 17, 11.8%
- trial — 3 of 29, 10.3%
- question — 2 of 24, 8.3%
- money — 1 of 25, 4.0%
The arm that led by five times is now fourth. First place goes to the arm with the least volume in the test, on the strength of two events. Nothing about the emails changed between those two tables. One decision changed: which events count as an open. It reversed the ranking, and on this volume it will reverse it again next week, whichever criterion we pick.
The event we sell on has no data at all
Here is the part that never makes it into a subject-line debate. Across all six arms and all 130 delivered emails: zero clicks, zero replies, one unsubscribe.
Wider than the test window, since the campaign started on 6 September: 267 emails delivered to 204 properties, 42 open events, 4 clicks, 0 replies, 0 customers. Day 19 of a 30-day target. That is our number and it is not a good one.
So the test is being scored on opens because opens are the only thing that has happened, and an open is the one event in the chain that has never paid us anything. An owner who opens the mail and closes it bought nothing. The metric with signal and the metric with meaning are not the same metric, and when you only hold one of them you do not have a test. You have a habit.
What a real answer would cost
Our own target for this campaign is a 3% reply rate. Take that as the denominator and the arithmetic gets short.
At 3%, an arm holding 22 delivered emails expects 0.66 replies. Not one reply; two thirds of one. You cannot order six arms by a number that rounds to zero in every one of them. To believe an ordering, you want the deciding event to occur on the order of ten times per arm, so the gap between arms is bigger than the gap one lucky Tuesday makes. At 3%, ten replies per arm is roughly 330 delivered emails per arm.
Six arms at 330 is about 2,000 emails. We send on weekdays with a cap of 20 a day, so 100 a week at the ceiling: twenty weeks to finish a copy test. Two arms is 660 emails, about seven weeks. Cutting from six arms to two does not make the test sharper. It makes it finish while the question still matters.
The same arithmetic on a 45-bed property
None of this is about cold email. Take a 45-bed hostel at 60% occupancy on an average stay of three nights: roughly 240 arrivals a month, call it 700 in a quarter, most with a usable email address.
You want to know whether a pre-arrival message sells more airport pickups. Two versions, 350 arrivals each, a 4% take rate: about 14 bookings per arm. That test has a spine.
Now split the same quarter six ways: 117 arrivals per arm, about 5 bookings each. The gap between best and worst will look enormous, exactly as 41.4% against 8.0% looked enormous, and it will be four bookings of noise. The difference between those two designs is not sophistication. It is whether you find out.
So: if the number you are deciding on will not happen at least ten times per arm inside a window you can wait out, do not split. Send one version to everybody and spend the effort on something structural. The price, the photo order on the OTA listing, who is on the list at all.
What we are doing with a tie
With no click and no reply anywhere in the test, our allocation engine pooled all six arms into one score and gave each the same weight, so volume spreads evenly and nothing is favoured. A system that reports a tie is worth more than one that names a winner it cannot see, because the named winner is the one you then defend for a month.
Three rules we run this on, and they transfer to guest emails, listing photos or rates:
- Name the deciding event before you split, and do not let it be an open. A booking, a reply, a paid signup. Write it down in advance, because afterwards every table offers a different winner.
- Count per person and per delivery, not per event. Twelve opens from an unknown number of mailboxes is not twelve readers, and an event you cannot date cannot be attributed to anything.
- Fewer arms. Two, not six. And when two tie on real volume, that is a finding: the wording is not the constraint. The list, the offer or the channel is.
What we are watching next is not an open rate. It is the first reply. Until one arrives the six versions stay tied, and we go change something bigger than a sentence.
If this sounds like your property, see the three Nightfill tiers and start a free 14-day pilot on your own data: view pricing.