ActiScan

DMARC

Why SPF and DKIM Break Mid-Year: The Client DNS Drift MSPs Miss

September 2, 2026

Rodney Hall, COO— AI-assisted and reviewed prior to publication.

Laptop showing DNS records on a workbench beside old and new cabling, symbolizing accumulated DNS drift

A client's SPF and DKIM pass every check the day they go live. Six months later, invoices are landing in spam and nobody on the client side has touched a DNS record. This pattern shows up often enough in MSP service desks that it deserves a name: DNS drift, the slow accumulation of small, individually reasonable changes that collectively break email authentication without any single change looking like the cause.

What Actually Causes SPF and DKIM to Break Without Anyone Changing Them?

SPF and DKIM rarely fail because someone edited the authentication record itself. They fail because the ecosystem around that record changed: a new SaaS tool added an SPF include, a marketing platform rotated its sending infrastructure, or a signing key aged past its useful life. The record that was correct in January silently stops being correct by August, and nothing in the client's change log points to email as the cause.

This is the operational reality behind a well-documented technical limit. Under RFC 7208, the standards document that defines SPF, a domain's SPF record can only trigger a fixed number of DNS lookups before receivers must reject it outright. Every include, a, mx, ptr, and exists mechanism, along with any redirect, counts toward that cap. When a client's marketing team adds a new email platform in March and a helpdesk tool in June, each one usually adds its own include: statement to the SPF record, and nobody connects those two unrelated purchases to a single shared risk.

Why Does the 10-Lookup Limit Catch MSPs Off Guard?

It catches MSPs off guard because the limit is invisible until it is exceeded, and exceeding it produces a permanent, silent failure rather than a warning. RFC 7208 specifies that SPF implementations must cap evaluation at 10 DNS lookups and return a PermError if that number is exceeded, and mailbox providers generally treat a PermError the same as an outright SPF fail.

The UK Government's own domain security guidance describes this plainly: SPF records generating more than 10 DNS lookups risk having their entire authentication result rejected, since the standard permits at most 10 lookups per evaluation and exceeding it means SPF processing may fail. The failure mode is what makes it dangerous for MSP operations: a record sitting at 8 or 9 lookups looks completely fine in a point-in-time check, and stays fine right up until the eleventh include gets added by a department that never files a change ticket with IT.

Nested includes make this worse. A single include: mechanism for a marketing platform or a security gateway often triggers its own chain of sub-lookups, so one vendor addition can consume two or three lookups instead of one, pushing a record from apparently safe to broken in a single afternoon with no visible warning sign at the DNS layer itself.

The Second Drift Vector: DKIM Keys Nobody Is Rotating

SPF lookup creep gets the attention, but DKIM has its own quiet failure mode: keys that are never rotated, selectors that get abandoned mid-migration, and signing configurations that silently point at a service the client stopped using a year ago. The Messaging, Malware and Mobile Anti-Abuse Working Group, the industry body most mailbox providers look to for authentication hygiene, states directly that DKIM keys should be rotated at least every six months to reduce the window in which a compromised or cracked key can be misused, and that frequent rotations also standardize the rotation process so institutional knowledge survives staff turnover. Most SMB clients rotate DKIM keys never, because nothing in their daily operations reminds them to.

The practical trap for MSPs is what happens during vendor migrations, not during quiet periods. When a client switches marketing platforms or replaces a helpdesk tool, the old DKIM selector and its matching SPF include frequently get left in DNS because removing them feels riskier than leaving them alone. That leftover record does not cause an immediate failure, but it does one of two things: it keeps consuming an SPF lookup slot for a service that no longer sends mail, or it leaves a stale public key in DNS that a receiver can still find and evaluate against messages that were never actually signed by that key. Either way, the client's authentication surface gets messier every time a vendor changes, and nobody circles back to clean it up because the mailbox that would show the failure symptom (invoices, password resets, vendor confirmations) is not the mailbox anyone is watching.

Why 2024 and 2025 Made This Problem More Urgent

Google and Yahoo began enforcing new bulk sender requirements in February 2024, and Google's own support documentation confirms that all senders, including Google Workspace users, must meet the requirements in Google's email sender guidelines when sending to personal Gmail accounts, with enforcement applied gradually and progressively through error codes on non-compliant traffic. Microsoft followed with its own enforcement timeline. According to authentication vendor dmarcian, Microsoft began rejecting non-compliant bulk mail outright rather than routing it to junk, with rejected messages producing error code 550 5.7.15 Access denied starting May 5, 2025.

The combined effect is that authentication failures which used to mean "slightly worse deliverability" now increasingly mean "message never arrives." A client sitting at p=none with a marginal SPF record could tolerate some sloppiness in 2022. That same client in 2025 is one HubSpot signup away from a PermError that Google or Microsoft treats as an outright fail, with no soft landing in the spam folder.

Drift triggerWhat breaksTypical time to symptom
New SaaS/marketing tool adds SPF include10-lookup ceiling exceeded, PermErrorWeeks to months
DKIM key never rotatedExtended exposure window, no visible symptom until compromise6+ months (per M3AAWG guidance)
Vendor migration leaves old selector/includeWasted lookup budget, stale signing surfaceImmediate but unnoticed
Nested includes from gateway/security vendorMultiple lookups consumed per single includeVaries, often invisible

Building Drift Detection Into MSP Operations

The fix is not a one-time SPF and DKIM audit at onboarding. It is a recurring check that catches drift between the reviews an MSP already schedules. A quarterly manual review misses a SaaS signup that happened in month two, and by the time the quarterly check runs, the client has already had weeks of silently failing invoices or vendor emails landing in spam.

Three habits close most of the gap:

  • Treat every new SaaS or marketing tool request from a client as a DNS change, not just an application change, since most of these platforms ask the client to add an SPF include or a DKIM selector during setup.
  • Track DKIM key age per client the same way patch levels get tracked, using the six-month M3AAWG window as the trigger for a rotation ticket rather than waiting for a client complaint.
  • Monitor SPF lookup count continuously rather than at review time, since the difference between 9 lookups and 11 is one vendor signup away and produces no warning from the DNS layer itself.

Continuous monitoring is where automated scanning earns its keep over spreadsheet-based tracking. A platform that checks every client domain's SPF lookup count, DKIM selector validity, and DMARC alignment on a recurring schedule catches the ninth include the week it lands, not the quarter after a client calls about bounced invoices. For an MSP standing up this workflow for the first time, ActiScan's getting-started guide walks through connecting client domains so lookup counts and key ages surface automatically instead of requiring a manual dig against every zone file.

Making Drift Monitoring Part of the Service, Not an Afterthought

The MSPs handling this well are not the ones with the most thorough onboarding checklist. They are the ones who treat SPF and DKIM as living infrastructure that needs the same drift detection applied to patch compliance or certificate expiry. A client's DNS zone changes every time they adopt a new tool, and every one of those changes is a candidate for breaking authentication that passed cleanly a year earlier.

Building that into a managed service does not require a large team. It requires visibility into every client domain at once, which is the gap that turns a five-minute fix into a Monday morning support ticket about vendor payments landing in spam. MSPs evaluating how this fits into existing per-client billing can review ActiScan's pricing page to see how domain monitoring scales across a client book, and the fastest way to see actual drift on a real domain set is to run it directly through a signup rather than estimating exposure from a single point-in-time check.

DNS drift is not a hypothetical. It is the predictable outcome of clients adopting new tools faster than anyone reviews the DNS footprint those tools leave behind, and the receivers evaluating that footprint have gotten considerably less forgiving since 2024.

← Back to all posts