If your outbound reply rate fell off a cliff in the last eighteen months and nobody on the team can explain why, the answer is probably not your copy. Something structural has changed about how cold email arrives. Mailbox providers have moved from filtering unwanted mail into a spam folder to rejecting it outright at the point of delivery, and they have published the conditions under which they will do so. That shift has consequences for anyone doing outbound, and it lands hardest on small teams who set up their sending infrastructure once, years ago, and have not touched it since. This is a practical guide to what changed, how to check whether it affects you, and what to fix in what order. It is deliberately unglamorous. Deliverability work is plumbing, and plumbing does not benefit from creativity.
What actually changed
The short version is that authentication moved from best practice to entry requirement, and complaint rates became a hard ceiling rather than a soft signal. Google and Yahoo introduced sender requirements for bulk senders in early 2024, defining bulk as roughly 5,000 or more messages per day to their consumer domains. The requirements covered SPF and DKIM authentication, DMARC alignment, one-click unsubscribe for commercial mail, and a spam complaint rate kept below 0.3 per cent. Microsoft followed with equivalent requirements for its consumer Outlook domains, including outlook.com, hotmail.com and live.com, announced in April 2025 and taking effect from 5 May 2025. Microsoft's own guidance uses the same 5,000-messages-per-day threshold and requires SPF, DKIM and DMARC, with DMARC at a minimum policy of p=none and alignment to SPF or DKIM. Microsoft also stated that mail failing these requirements would ultimately be rejected rather than routed to junk. Three details in that matter more than the headline. The first is that rejection is different from filtering. A message routed to spam still counts as delivered and still shows up in your sending tool's delivery statistics. A rejected message is refused at the protocol level, which means your platform sees a bounce. This is genuinely better for you, because it is honest feedback, but it will make your dashboards look worse before it makes your results better. The second is the complaint rate threshold. Below 0.3 per cent sounds generous until you do the arithmetic. On a thousand-message campaign, three complaints puts you at the ceiling. Three people out of a thousand marking a message as spam is not an unusual outcome for badly targeted cold email. It is a routine one. The third is that these thresholds apply to consumer mailboxes, and a great deal of B2B outbound goes to business domains hosted on Google Workspace or Microsoft 365. The published requirements do not technically cover those, but the filtering models behind them are informed by the same reputation signals. Assuming business inboxes are exempt is a mistake that takes a quarter to become visible and much longer to recover from.
Reply rates in context, and why the benchmarks are useless
Before fixing anything, it is worth being honest about what a normal reply rate is, because a lot of teams are trying to fix a problem that does not exist while ignoring one that does. Published cold email benchmarks are close to worthless for comparison, and it is worth understanding why rather than just picking a number. Different providers count different denominators. Some measure replies against emails sent, others against emails delivered, others against contacts in a sequence regardless of how many messages each received. Some count out-of-office replies and unsubscribes as responses. Almost all of the data comes from vendors reporting on their own customers, which is a self-selecting sample of people who bought a cold email tool. The result is that published averages for cold email reply rates in 2026 span a range wide enough to make the average meaningless. Figures around 3 to 4 per cent appear frequently across large sample analyses, and figures above 15 per cent appear in write-ups of heavily personalised, low-volume campaigns. Both can be true and neither tells you whether your campaign is working. The only benchmark worth anything is your own, measured consistently. Pick a definition of reply rate, write it down, and stop comparing yourself to blog posts. If your reply rate was 4 per cent last year on the same definition and it is 1.2 per cent now, you have a real signal. If you have never measured it consistently, that is the first thing to fix, before touching infrastructure.
The technical checklist, in order
Deliverability advice tends to arrive as a long list with no priority, which is unhelpful when you have an afternoon. This is ordered by return on effort.
- Confirm SPF, DKIM and DMARC are published and passing for every domain and subdomain you send from. Not configured, passing. Send a message to a mailbox you control at a Gmail address and check the original headers for three explicit passes. Configured-but-failing is the most common state and looks identical to correct from inside your DNS panel.
- Set DMARC to at least p=none with a reporting address, and actually read the reports for a fortnight. This is the cheapest diagnostic available and it will tell you about sending sources you had forgotten existed, which for most small businesses includes an old invoicing tool and a form plugin on the website.
- Separate your cold outbound domain from your primary domain. If you send outbound from the same domain as your customer email, invoices and password resets, a reputation problem in outbound becomes a business continuity problem. Use a separate domain, warmed properly, and accept that it will perform slightly worse because it has no history.
- Add one-click unsubscribe headers to commercial messages and honour them immediately. There is a persistent belief in cold email circles that including an unsubscribe link marks a message as bulk and hurts delivery. The published requirements point the other way, and a one-click unsubscribe is strictly better for you than the alternative, which is the recipient using the spam button to achieve the same result.
- Verify your list before every send, not once when you bought it. Bounce rate thresholds are low and B2B contact data degrades continuously. Published decay estimates vary a great deal by source and method, with figures from roughly 20 per cent to well over 50 per cent annually appearing in vendor research, but the direction is not in dispute: average job tenure is short, companies get acquired, and email formats change. A list that was clean in March is not clean in September.
- Keep per-mailbox volume low and consistent. Sudden volume changes on a sending identity look like exactly what they usually are. Ramping is not a warm-up ritual you complete and forget, it is an ongoing constraint. That is the whole technical list. It is not long, and none of it is difficult. It is simply nobody's job in most small companies, which is why it does not get done.
The part that actually determines your complaint rate
Here is the uncomfortable conclusion of everything above. The 0.3 per cent complaint ceiling is not a technical constraint. It is a targeting constraint wearing technical clothing. People do not mark messages as spam because the DKIM signature is missing. They mark messages as spam because the message was irrelevant to them, arrived without any plausible reason, and looked like it had been sent to ten thousand other people. Every technical fix in the previous section improves your ability to reach an inbox. None of them affects what happens once you are there. Which means the deliverability work with the highest return for a small team is sending fewer, better-matched emails. Not as a philosophical position about respecting people's time, though that argument is available too, but because the complaint rate maths does not work any other way at volume. This is why the economics of cold outbound have shifted against high-volume approaches specifically. When the penalty for poor targeting was a lower response rate, volume was a rational way to compensate. Now the penalty compounds: poor targeting produces complaints, complaints damage sending reputation, damaged reputation reduces delivery for everything you send afterwards including well-targeted messages. The strategy eats itself. The practical implication is a smaller list with a real reason for each contact being on it. That reason should be specific enough to write into the email. If you cannot articulate why this person, at this company, this week, then the message will read as untargeted because it is untargeted, and some proportion of recipients will act accordingly.
The copy decisions that move complaint rate
Since complaint rate is now the binding constraint, it is worth being specific about which writing choices produce complaints, because the advice usually stops at "be relevant" and that is not actionable at 9am on a Monday. Subject lines that imitate an existing conversation are the single most reliable way to generate complaints. A "re:" prefix on a first contact, a fake forward, or a subject that suggests a prior exchange creates a moment of confusion followed by irritation, and irritation is what the spam button is for. Microsoft's guidance explicitly calls out subject lines that do not accurately describe the message. This tactic still circulates in outbound communities because it lifts open rates, which is precisely the wrong metric to optimise when the penalty lands on complaints. Fake familiarity in the opening line has the same effect. Claiming you were referred by someone unnamed, implying a previous conversation, or opening with a compliment so generic it could apply to any company all signal mass sending. Recipients have become good at recognising the pattern and the recognition itself produces the complaint. Personalisation that is obviously mechanical is worse than no personalisation. A first name in the subject line, a company name inserted mid-sentence in a way that reads awkwardly, or a reference to a LinkedIn post that clearly came from a scraper all announce that a machine wrote the message. Plain, direct copy with no personalisation at all performs better than visibly automated personalisation, because at least it is not pretending. The other consistent driver is a mismatch between the effort implied and the ask. A three-line email asking a single question is proportionate. A three-line email asking for thirty minutes on Thursday is not, and the gap between the two is where people decide you are not worth the reply. Length matters less than people think, but structure matters more. A message that makes the reason for contact clear in the first sentence gives the recipient an immediate decision, which is what they want. A message that spends two paragraphs establishing credibility before arriving at the point converts a quick no into an annoyed one.
Rebuilding a sending setup that has already been damaged
If your domain reputation is already poor, the honest answer is that recovery is slow and partial. There is no reset button, and services promising one are selling you a warm-up schedule with a markdown. The sequence that works starts with stopping. Pause the outbound entirely for a fortnight rather than continuing at reduced volume, because continued poor-quality sending prevents any recovery signal from forming. Use the fortnight to fix authentication, clean the list properly and rewrite the sequence. Then restart small on a new domain, keep volume genuinely low for several weeks, and prioritise the segment of your list most likely to reply positively. Early engagement on a new sending identity matters disproportionately. Sending your best-matched hundred contacts first is both better business and better infrastructure practice. Keep the old domain out of outbound permanently. Let it recover as a normal business domain over months. Trying to rehabilitate a burnt domain while still using it for cold email is the equivalent of resting an injury by walking on it slightly less. Finally, put a monthly check in place so this does not recur. Ten minutes reviewing bounce rate, complaint rate where your provider exposes it, and reply rate by segment. Deliverability degrades gradually and invisibly, and the teams who avoid a crisis are the ones who noticed the trend at month two. Empiraa Signal keeps sequencing, contact data and pipeline in one place, which removes one common source of this problem: lists that were verified in one tool, exported to another, and sent from a third without anyone checking what decayed in between. The infrastructure discipline still has to be yours. This gives a small team sequencing and contact data in one place without turning the workflow into another reporting burden. You can review Signal pricing before choosing a plan.
Knowing which change fixed it
One last operational point. Teams tend to fix deliverability by changing eight things in a week, watching the numbers improve, and having no idea which change mattered. That feels fine until it degrades again in six months and nobody knows what to undo. Sequence the changes and leave a week between the ones that are reversible. Authentication fixes can all go in at once because they are unambiguous and there is no scenario in which correct SPF, DKIM and DMARC makes things worse. After that, change one variable at a time: list quality, then volume, then copy, then send timing. Give each change at least a fortnight before judging it. Reputation signals do not update daily, and a week of good numbers after a change is as likely to be normal variance as improvement. Small teams are particularly prone to declaring victory early because the sample sizes are small enough that a single good day looks like a trend. Track three numbers and no more: bounce rate, reply rate on your fixed definition, and complaint rate if your provider exposes it. Adding more metrics at this stage produces the appearance of rigour without improving any decision. The three above are enough to tell you whether the problem is reaching the inbox, being read, or being resented. Write down what you changed and when, in the same place as the numbers. This sounds like unnecessary process for a two-person outbound function, and it is the single thing that separates teams who solve this once from teams who solve it repeatedly.
Frequently asked questions
What are the new bulk sender requirements for cold email?
Google and Yahoo introduced requirements in early 2024 and Microsoft introduced equivalent requirements for consumer Outlook domains effective 5 May 2025. All three apply to senders of roughly 5,000 or more messages per day to their consumer mailboxes, and require SPF and DKIM authentication, a DMARC record with a policy of at least p=none aligned to SPF or DKIM, an easy unsubscribe mechanism for commercial mail, and a spam complaint rate below 0.3 per cent. Microsoft has stated that non-compliant bulk mail will be rejected rather than filtered to junk.
Does the 5,000 messages per day threshold mean small senders are exempt?
Not in any way you should rely on. The published thresholds define when the requirements are enforced as a hard gate, but the underlying reputation and filtering systems evaluate all senders. A small sender with failing authentication and a high complaint rate will see poor delivery regardless of whether they cross a stated volume threshold. Treat the requirements as the minimum standard rather than a rule that applies to someone else.
What is a good cold email reply rate in 2026?
There is no reliable industry answer, because published benchmarks use inconsistent denominators and are almost entirely drawn from vendors reporting on their own customers. Figures around 3 to 4 per cent recur in large-sample analyses and much higher figures appear in reports on low-volume personalised campaigns. The useful benchmark is your own historical rate measured on a fixed definition. Choose whether you count replies against messages sent or contacts sequenced, whether out-of-office counts, and then keep that definition unchanged so the trend means something.
Should cold email include an unsubscribe link?
Yes, and preferably a one-click unsubscribe honoured immediately. The concern that an unsubscribe link marks a message as bulk mail is outweighed by the alternative: a recipient who wants your messages to stop and has no easy way to achieve it will use the spam button, which damages your sending reputation in a way an unsubscribe does not. Complaint rate thresholds are low enough that this trade is not close.
Should outbound be sent from the main company domain?
No. Use a separate domain for cold outbound so that a reputation problem there cannot affect delivery of customer email, invoices, password resets and other mail your business depends on. Expect the separate domain to perform slightly worse initially because it has no sending history, and warm it up gradually. Keeping the two functions on one domain works fine until the first time it does not, at which point the cost is disproportionate.
How often should a B2B contact list be verified?
Before each significant send, and at minimum monthly for any list in active use. Estimates of annual B2B data decay vary widely across published research, from around 20 per cent to well over 50 per cent depending on the field measured and the methodology, but the practical point is consistent: contacts leave, domains change and companies are acquired continuously. Given that bounce rate thresholds are strict, verification is cheaper than the reputation damage from sending to a stale list.


