Data
6,043 cold emails: our real bounce and reply rates
Published benchmark posts are almost always someone else's aggregate. This is ours - eight campaigns from our own outbound, including the one that bounced at 25.6% and taught us more than the seven that worked.
Data through August 2026
Short answer: across 4,433 contacts and 6,043 sends, seven campaigns held a bounce rate between 0.34% and 0.99%. One hit 25.63%. And the campaign group with the fewest sends produced every single positive reply and every meeting we booked.
Every campaign, including the bad ones
Eight campaigns, 6,043 emails sent to 4,433 contacts. Lists were built by scraping and cross-matching marketplace sellers, then double-verified before sending. Campaign names are anonymised; the numbers are not.
| Campaign | Contacts | Sent | Bounce | Replies | Reply rate |
|---|---|---|---|---|---|
| Scrape batch 1 | 256 | 575 | 0.87% | 7 | 1.22% |
| Scrape batch 2 | 296 | 588 | 0.34% | 4 | 0.68% |
| Scrape batch 3 | 593 | 741 | 0.54% | 8 | 1.08% |
| Cross-match (two signals) | 415 | 417 | 0.72% | 8 | 1.92% |
| Scrape batch 4 | 361 | 364 | 0.55% | 6 | 1.65% |
| Broad industry list | 1,695 | 1,624 | 0.99% | 3 | 0.18% |
| A/B test campaign | 460 | 1,221 | 25.63% | 5 | 0.41% |
| Marketplace list | 357 | 513 | 0.78% | 3 | 0.58% |
Weighted across the seven healthy campaigns, the bounce rate was 0.75%. Include the eighth and it becomes 5.78%. One campaign moved the company-wide number by nearly eight times, which is the single most useful thing in this table.
The campaign that bounced at 25.6%
1,221 sends, roughly 313 hard bounces. It is the only campaign we have ever run above 1%, and it is the reason our verification process now looks the way it does.
The immediate cost was not the wasted sends. It was the sending reputation. Mailbox providers read a sustained pattern of hard bounces as evidence that the sender is working from a purchased or scraped list, and that judgement attaches to the infrastructure rather than to the campaign. Every other campaign sending from those domains inherited the damage.
What the data shows plainly is a list that was verified at build time and then sent to later, at volume - 1,221 sends against 460 contacts, the highest send-to-contact ratio of any campaign we ran. B2B contact data does not sit still. People change jobs continuously, and a verification result is a statement about the day it was run, not a permanent property of the address.
The fix was a process change rather than a tooling change: verification now happens immediately before send rather than when the list is assembled, and any list older than about 30 days is re-verified from scratch before it is touched. We covered the sending-side setup separately in cold email deliverability: SPF, DKIM and DMARC.
A list 4x smaller returned 10x the reply rate
The clearest finding in the dataset. Our broad industry list held 1,695 contacts and produced 3 replies. A cross-matched list of 415 - one quarter the size - produced 8.
| Broad industry list | Cross-matched list | |
|---|---|---|
| Contacts | 1,695 | 415 |
| Emails sent | 1,624 | 417 |
| Replies | 3 | 8 |
| Reply rate | 0.18% | 1.92% |
| Positive replies | 0 | 2 |
| Meetings booked | 0 | 1 |
The difference between those two lists is one filter. The broad list was everyone in a category. The cross-matched list was sellers who appeared on two marketplaces at once - a group for whom our opening line was accurate rather than plausible.
That is what an enrichment waterfall is actually for. Not to append more columns to a spreadsheet, but to find the intersection where a specific claim about the prospect is verifiably true. One accurate signal outperformed four times the volume.
Where every meeting actually came from
The five signal-built campaigns sent 2,685 emails and produced all 9 positive replies and all 3 meetings. The three broader campaigns sent 3,358 emails - more volume - and produced zero of either.
| Broader lists (3 campaigns) | Signal-built lists (5 campaigns) | |
|---|---|---|
| Emails sent | 3,358 | 2,685 |
| Replies | 11 | 33 |
| Positive replies | 0 | 9 |
| Meetings booked | 0 | 3 |
Reply rate is the metric everyone reports and the one that flatters you most. The broader lists did generate replies - 11 of them - but not one was positive. Measuring at the reply line would have shown a campaign that was merely underperforming. Measuring at the positive line showed a campaign returning nothing at all.
This is the entire argument for treating list construction as engineering work rather than admin. The difference between those two columns is not effort or copy. It is whether someone built the filter that made the outreach true.
How to run this measurement on your own campaigns
Four numbers per campaign, tracked per list rather than per account: bounce rate, reply rate, positive-reply rate, and meetings. The fourth is the only one that pays you, and the first three only matter as leading indicators of it.
- Track by list source, not by campaign. Our batch 2 result - reply rate roughly half of batch 1 on identical copy - was only legible because the lists were logged separately. Attributed to the campaign, it would have looked like a copy problem and we would have rewritten a working email.
- Record bounce rate per sending domain as well as per campaign. Campaign-level bounce tells you the list was bad. Domain-level bounce tells you how much reputation you spent finding out.
- Report positive replies separately from replies. Two of our campaigns had respectable reply rates and produced nothing. The gap between the two numbers is where self-deception lives.
- Log the hypothesis before the send. Every campaign above was a stated test with a written expectation. Without that, a result is an anecdote you can interpret whichever way is most comfortable.
What this data is not
Eight campaigns and 6,043 sends is a real sample but a small one, drawn from a single market - e-commerce marketplace sellers - over a few months in 2026.
Reply rates are heavily category-dependent, and anyone selling into a different market should expect different absolute numbers. Three meetings is a small denominator, and the difference between 3 and 0 is not statistically robust on its own.
What travels is the shape, not the magnitude: verification timing drives bounce rate, list construction drives reply quality, and the positive-reply line tells you things the reply line hides. We publish the failures because a benchmark set with no bad campaigns in it is a marketing asset rather than a measurement.
Frequently asked questions
What is a good cold email bounce rate?
Under 1%. Seven of our eight campaigns ran between 0.34% and 0.99% on double-verified lists. Above about 2% means the list was not verified immediately before sending, and sustained rates above 5% put the domain itself at risk.
What reply rate should I expect?
Ours averaged 0.73% across all sends, but the spread ran from 0.18% to 1.92%. Treat the spread as the real finding - it is a property of the list far more than of the copy.
Does a bigger list get more replies?
Not in our data. A 415-contact cross-matched list produced 8 replies; a 1,695-contact broad list produced 3.
How do you keep bounce under 1%?
Verify twice with different providers, verify immediately before send rather than at build time, and re-verify any list older than about 30 days from scratch.
Who builds this in a team?
A GTM engineer - the person who builds enrichment, verification and scoring, as opposed to the rep who works the resulting list.
Want lists built this way?
Book a 30-minute discovery call. We'll map the role, show you matched GTM engineer profiles, and have your pod live within 7 days.