There is a widely repeated claim that AI-written email lands in spam, and that you therefore need to humanise your outreach copy to reach the inbox. We went looking for the evidence behind it. There is none, for it or against it. What there is instead is a very well-documented set of things that genuinely do get you filtered, and none of them are about who wrote the words.
The claim, and what actually supports it
No major mailbox provider, not Google, not Microsoft, not Yahoo, has ever published a policy, a changelog entry or an engineer statement indicating that machine-written text is treated as a negative deliverability signal. There is no independent study demonstrating that AI-authored copy is filtered at a higher rate than human-authored copy from matched sending infrastructure.
Meanwhile almost the entire corpus of “AI cold email goes to spam” content is published by companies selling warmup services, verification, or humanising tools. That is not a reason to dismiss it. It is a reason to notice that a specific commercial interest is the only source.
Detecting AI-typical patterns is technically plausible. Providers penalising on that basis is nowhere evidenced. Vendors routinely present the first as proof of the second.
Here is the honest version of the connection, which is real but indirect. AI does not change how a message is filtered. It changes how many messages get sent and how similar they are to each other, and template similarity at volume is a well-established filtering signal. Near-identical messages differing only in merge fields get clustered and flagged regardless of who or what wrote them.
So the underlying advice, vary your messages, is correct. The stated reason is wrong. You are not defeating an AI detector. You are defeating similarity clustering, which existed long before generative models and would flag a human writing from a template just as fast.
What actually gets you filtered
The requirements are published, specific, and now enforced with rejections rather than warnings.
Google and Yahoo, from February 2024
A bulk sender is anyone sending more than 5,000 messages a day to Gmail addresses. The threshold is per sending domain, and once you cross it in a single day you are permanently classified as a bulk sender.
- SPF, DKIM and DMARC, with alignment. DMARC at minimum p=none.
- Spam complaint rate below 0.3% in Postmaster Tools. Google’s own guidance is to stay under 0.10% to build resilience, and that 0.1% figure comes from Google, not from vendor folklore.
- One-click unsubscribe per RFC 8058, honoured within two days.
- TLS in transmission, valid forward and reverse DNS, RFC 5322 compliant formatting.
Yahoo’s requirements mirror Google’s, with one detail worth knowing: Yahoo measures its 0.3% spam rate against mail delivered to the inbox, a stricter denominator than most senders assume.
Microsoft, from May 2025
This is the update most cold-email operations missed. Microsoft applied equivalent requirements to senders of more than 5,000 messages a day to Outlook.com, Hotmail.com and Live.com, with enforcement beginning 5 May 2025. Non-compliant mail was initially routed to Junk with a 5.7.515 rejection code.
Microsoft also weighs IP reputation considerably more heavily than Gmail, which is domain-dominant. For anyone rotating shared IPs across sending infrastructure, this is a meaningful difference.
The enforcement ramp, April 2024 to November 2025
This is where the timeline gets misreported, so it is worth being exact. Gmail started with temporary 421 deferrals in February 2024 and began issuing permanent 550-5.7.26 rejections to some non-compliant traffic as early as April 2024. November 2025 is when it escalated that to full-scale permanent rejection. Microsoft followed a similar path, junk-foldering first and moving to permanent rejection for persistent offenders in late 2025. Either way, the soft-enforcement window is closed.
Which inbox is actually hard to reach
Validity’s 2025 Email Deliverability Benchmark Report, based on seed-address monitoring across hundreds of mailbox providers through 2024, gives the most useful picture of where mail actually lands.
| Provider | Inbox | Spam folder | Missing entirely |
|---|---|---|---|
| Gmail | 87.2% | 6.8% | 6.0% |
| Yahoo and AOL | 86.0% | 4.8% | 9.2% |
| Apple | 76.3% | 14.3% | 9.4% |
| Microsoft | 75.6% | 14.6% | 9.8% |
| Global average | 83.5% | 6.7% | 9.8% |
Two things stand out. Microsoft is the hardest major inbox, with spam placement roughly double Gmail’s. For B2B outreach, where Outlook and Microsoft 365 share is far above the consumer figure, that is your binding constraint, and it is the provider whose requirements arrived a year later than everyone else’s.
And the “missing” column is larger than the spam column. Nearly one message in ten is accepted and then silently dropped, delivered nowhere. It does not bounce and it does not appear in a spam folder. If you are measuring deliverability by bounce rate alone, you are not measuring the largest failure mode.
List quality is the actual lever
Hard bounces are the clearest signal a mailbox provider has that a sender does not have a real relationship with the recipient. High bounce rates are, statistically, a proxy for purchased or scraped lists, which is exactly what cold outreach is, which is why the tolerance is so low.
The decay figures are the part people underestimate. ZeroBounce, processing over eleven billion addresses across 2025, reported that at least 23% of a typical list degrades per year, and that only around 62% of email submissions were valid at capture. Note that ZeroBounce headlines 28% in its own press release for the same report, so you will see both numbers quoted; 23% is the figure on the report page itself. Around 9% of addresses processed were catch-all domains. Spam traps were roughly 0.01% of volume, negligible as a percentage and disproportionate in damage.
Practical benchmarks converge across sources: keep cold-outreach bounce rates under 2%, treat anything over 3% as a problem, and expect throttling or suspension above 5%.
Spam traps deserve a specific note because they are the failure that ends a domain rather than degrading it. Pristine traps are addresses that never belonged to a human, planted specifically to catch scrapers, so hitting one is near-proof of list scraping. Recycled traps are abandoned real addresses repurposed after a period of dormancy, and hitting those signals poor hygiene rather than malice. Verification catches the second category reliably and the first only by inference.
Catch-all domains are where verification honestly runs out. A catch-all accepts mail to any address at the domain, so no verifier can confirm a mailbox exists. In one B2B test set, catch-alls were 28% of the list. Any vendor quoting a headline accuracy figure is usually computing it over the addresses it resolved, which means the effective coverage on a B2B list can be far below the number on the pricing page.
On verification accuracy claims
Nearly every verification vendor claims 97% to 99%-plus accuracy. Independent testing has put real-world accuracy for at least one major provider at 93% to 97% against a 99% claim. At ten thousand addresses, a two-to-six-point gap is two to six hundred misclassifications.
More importantly: no genuinely independent, methodologically transparent, reproducible email verification benchmark exists in the public record. Every comparison we could find is published by a party selling verification, adjacent tooling, or affiliate placements, including one detailed ten-thousand-address study in which the publishing vendor ranks itself first with a sixfold lead on the single differentiating metric.
The questions worth asking a verifier, since the accuracy number tells you nothing: what percentage of a B2B list do you return as unknown, how do you handle catch-alls specifically, and is your headline accuracy computed over the full list or only over resolved addresses.
We built BounceBlock because the answers to those three questions determine everything and almost nobody publishes them. We would rather tell you a quarter of your B2B list is unresolvable than report a confident number over the three quarters that were easy.
Stop measuring open rates
This is the most consequential measurement problem in email and it is still routinely ignored.
Apple Mail Privacy Protection, launched September 2021, pre-fetches tracking pixels through Apple’s proxy. An “open” is registered whether or not a human read anything. Apple Mail accounted for around 49% of all recorded opens as of early 2025. Gmail proxies images too.
The inflation is 15 to 20 percentage points. A campaign that genuinely achieved 28% engagement now displays as roughly 45 to 52% with identical behaviour. Security appliances that scan links generate both fake opens and fake clicks, with recorded bot-click volumes peaking in the millions per day and concentrated on business, government and education domains, precisely the domains B2B outreach targets.
Open rate is not a usable primary metric in 2026. Reply rate, positive reply rate, meetings booked and bounce rate are the defensible KPIs. Any vendor quoting open-rate lift as proof of anything should be discounted on that basis alone.
What good actually looks like
The largest current benchmark dataset comes from Instantly, covering campaigns across 2025. It is a vendor dataset and should be read as such, but it is the largest available.
Three findings from that dataset are worth acting on. 58% of all replies come from the first email, with the remaining 42% spread across every subsequent touch. Follow-ups matter, but the first message does most of the work, and sequences beyond about seven touches earn very little. Emails under 80 words perform best. And warmup should run at five to ten emails a day scaling over four to six weeks, not a volume ramp measured in days.
The checklist
- Authenticate properly. SPF, DKIM and aligned DMARC on every sending domain. Non-negotiable since 2024, and rejected rather than deferred since November 2025.
- Verify before every send, not once at import. A list decays roughly 23% a year, so a list verified six months ago is not verified.
- Treat catch-alls as a known unknown rather than assuming they are valid, and size your risk accordingly.
- Watch complaint rate above everything. Under 0.3% is the requirement; under 0.1% is Google’s own recommendation and the level that gives you room to recover from a bad week.
- Vary your messages substantively, with different angles, different evidence, different asks. Not spun synonyms of the same message, which is what similarity clustering is designed to catch.
- Test Microsoft separately. If you are sending B2B and only checking Gmail placement, you are testing the easy inbox.
- Measure replies, not opens.
- Send fewer, better messages. Everything in the enforcement regime (complaint thresholds, similarity clustering, bounce tolerance) penalises volume without differentiation. That is the same principle Google applies to web content, arrived at independently by a different industry.
Disclosure. Lacewing builds BounceBlock, an email verification and deliverability tool. We have said above that no independent benchmark exists in our own category and that vendor accuracy claims including ours should be interrogated rather than believed. That is not modesty; it is the actual state of the evidence.
The cross-industry pattern
It is worth noting how closely this mirrors the search side. Neither Google Search nor any mailbox provider penalises AI authorship. Both penalise the second-order consequence of cheap generation: undifferentiated output produced at volume to extract value from a channel. Google codified it as scaled content abuse. Mailbox providers enforce it as similarity clustering plus complaint and bounce thresholds.
AI is not the risk variable in either place. Volume without differentiation is, and AI is simply the cheapest way to produce it.
Sources and further reading
- Google, Email sender guidelines (support.google.com/mail/answer/81126)
- Yahoo, Sender best practices
- Microsoft, Outlook.com sender requirements, effective May 2025
- Reporting on stricter Google and Microsoft authentication enforcement, November 2025
- Validity, Email Deliverability Benchmark Report
- ZeroBounce, Email List Decay Report (vendor data, 11 billion addresses processed)
- Instantly, Cold Email Benchmark Report 2026 (vendor dataset)
- Apple Mail Privacy Protection documentation and subsequent open-rate inflation analyses
