Finally you can send cold emails for $9.99/mo with our new rebellion. Click here to learn more!
Cheap Inboxes Editorial9 min read

Cold email deliverability: find the broken layer before you change everything

Diagnose cold email deliverability across authentication, accounts, lists, copy, sending behavior, and measurement with one evidence packet.

DeliverabilityDomains & DNSInbox operations

When cold email performance drops, changing everything at once is the fastest way to lose the cause.

Do not replace the inbox provider, rewrite the copy, slow every mailbox, add a warmup tool, and upload a new list on the same day. Pick one affected message path. Capture the evidence. Find the broken layer. Change one thing and retest the same path.

There are six layers to check:

  1. DNS and authentication
  2. Provider and account state
  3. List quality
  4. Message and recipient fit
  5. Sending behavior
  6. Measurement

The point is not to prove that one layer never matters. It is to stop one vague symptom—“deliverability is down”—from becoming a full infrastructure rebuild.

Build the evidence packet first

Before touching the setup, choose one affected domain, one mailbox, one campaign, and a narrow time window.

Capture:

  • sending domain and mailbox;
  • mailbox provider and workspace or tenant;
  • sequencer and campaign;
  • recipient provider when known;
  • exact dates when the change started;
  • sent, delivered or accepted, bounced, replied, and complained counts available from the systems you use;
  • bounce codes and full error text;
  • message headers from representative mail;
  • current SPF, DKIM, and DMARC results;
  • list source, verification date, and recent changes;
  • subject, body, links, attachments, and tracking settings;
  • schedule, volume pattern, and concurrency changes;
  • provider notices, restrictions, or disabled-account events;
  • infrastructure, copy, data, or configuration changes near the start of the problem.

Do not begin with open rate. Image blocking, privacy features, scanners, and client behavior make it a weak diagnostic by itself. Replies, bounces, complaints, provider errors, final-recipient headers, and controlled placement tests answer narrower questions.

Now write the symptom in a way that can be tested.

Bad: “Outlook deliverability is terrible.”

Useful: “Replies from Microsoft-hosted recipients fell after July 30 for Campaign B on three mailboxes under example.net; Gmail-domain replies did not move, and no account restriction is visible.”

The second statement gives you a scope.

Layer 1: DNS and authentication

Start here when you see authentication errors, recent DNS work, new sending sources, or failures concentrated on one domain.

For a real message, confirm:

  • which domain appears in the visible From address;
  • which MAIL FROM or Return-Path domain SPF authenticated;
  • which DKIM d= domain signed the message;
  • whether SPF passed and aligned;
  • whether DKIM passed and aligned;
  • whether DMARC passed;
  • whether a forwarding or security hop changed the path.

A DNS checker proves that it can find a record. It does not prove that the failing message used the record you expected.

Use the final recipient's Authentication-Results header. If SPF and DKIM appear to pass but DMARC fails, the likely question is alignment. Follow the DMARC failure guide and put the three domains next to each other.

Also look for operational DNS mistakes:

  • more than one SPF policy at the same hostname;
  • an SPF processing error;
  • a DKIM selector that does not resolve;
  • signing disabled in the provider;
  • more than one DMARC record;
  • a stale record after changing tools;
  • nameservers moved without copying the required records;
  • forwarding or reply records pointing to the wrong system.

Fix the exact source and retest the same message path. Do not add unrelated vendors to SPF because they happen to be in the stack.

Evidence that clears this layer: representative messages pass the intended aligned authentication path, DNS resolves from authoritative servers, and the same symptom remains after a controlled retest.

Layer 2: provider and account state

If authentication is healthy, check whether the mailbox provider accepted the activity and whether the account is in a normal operating state.

Look for:

  • disabled or suspended users;
  • outbound restrictions;
  • unusual login or security events;
  • daily or rolling limit errors;
  • rejected connections;
  • provider warnings;
  • sudden differences between Google and Microsoft accounts;
  • one tenant, workspace, domain, or mailbox failing while peers continue normally.

Google and Microsoft control the underlying service behavior of standard Google Workspace and Microsoft 365 mailboxes. Buying those mailboxes from a different reseller does not create a different Google or Microsoft transport layer.

The inbox supplier still matters for the operating work around the account: workspace structure, administration, DNS, handoff, billing, and support. But a provider restriction should be diagnosed as a provider or account event, not described as a mysterious “bad inbox.”

Compare affected and unaffected accounts with the same campaign conditions.

CompareWhat it can tell you
Same domain, different mailboxPossible account-specific state or assignment issue
Same provider, different domainPossible domain or campaign issue
Same campaign, Google vs MicrosoftPossible provider-specific behavior or configuration
Same inbox, before vs after changePossible schedule, copy, list, or connection change

Do not treat published product limits as recommended cold email volume. They define service boundaries, not the appropriate behavior for your domain and recipients.

Evidence that clears this layer: no relevant restriction or error is present, account access and sending are normal, and a matched account test does not reproduce the failure.

Layer 3: list quality

A clean technical setup can still send to a bad audience.

List problems often show up as:

  • hard bounces clustered around one source;
  • catch-all or risky addresses dominating the batch;
  • stale data;
  • malformed or role-based addresses;
  • recipients who have no plausible reason to care;
  • duplicate contacts or broken suppression logic;
  • one enrichment or data vendor producing a different result from the rest.

Trace the list, not just the total bounce rate.

For each list segment, record:

  • source;
  • collection or enrichment date;
  • verification method and date;
  • exclusions applied;
  • country, industry, company size, or role filters;
  • number loaded, skipped, bounced, and replied;
  • whether the same contact appeared in another campaign.

Then compare like with like. If one data source has a materially different bounce or reply pattern under the same mailbox, copy, and schedule, isolate that source before changing the infrastructure.

Low replies are not automatically a deliverability failure. A technically delivered message to the wrong person can perform exactly as badly as a message in spam.

Evidence that clears this layer: the affected and control lists have comparable sourcing and verification, bounces do not cluster around a data source, and a small matched audience test reproduces the problem.

Layer 4: message and recipient fit

The message itself can create filtering or indifference.

Inspect the exact version sent, not the copy in a planning document.

Check:

  • subject and body;
  • personalization tokens and fallback values;
  • link count and destination domains;
  • redirect or tracking domains;
  • images and attachments;
  • HTML complexity;
  • unsubscribe mechanism;
  • From name, address, and reply path;
  • whether the offer makes sense for the segment;
  • whether the same template was used across several unrelated audiences.

Separate two questions:

  1. Did the message reach a mailbox?
  2. Did the recipient have a reason to respond?

Placement testing can help with the first question under controlled conditions. Real replies and negative responses tell you more about the second. Neither should be stretched into proof about the entire campaign.

Run a controlled message test:

  • same mailbox and schedule;
  • small, comparable recipient segments;
  • one version with the suspected link, tracking, or formatting choice;
  • one version without it;
  • no simultaneous list-source or infrastructure change.

If both the audience and message change, the result cannot tell you which mattered.

Evidence that clears this layer: a simpler or better-matched message does not change the scoped outcome, and the same template behaves differently only when another layer changes.

Layer 5: sending behavior

Sending behavior is the pattern around the messages:

  • new-contact volume;
  • follow-up volume;
  • schedule and time zone;
  • bursts versus steady activity;
  • number of active campaigns per mailbox;
  • mailbox rotation;
  • reply handling;
  • pauses and restarts;
  • sudden capacity changes;
  • whether the sequencer retried failed work.

Calculate total messages, not only new contacts. A sequence with several follow-ups creates a different mailbox load from a one-message campaign.

Look for step changes. Did the team double new leads, enable several campaigns, shorten delays, or reconnect mailboxes after an outage? Did a timezone mistake create a burst? Did the same mailbox get assigned twice?

There is no single “safe” daily number that applies to every account, domain, provider, list, and message. Use provider restrictions as hard boundaries, then make operating decisions from the evidence on the scoped setup.

The useful test is a controlled change:

  1. Pick the affected mailbox group.
  2. Hold list source and message constant.
  3. Change one schedule or volume variable.
  4. Retest across the same recipient-provider mix.
  5. Compare bounces, provider errors, replies, and placement evidence.

Evidence that clears this layer: the symptom persists under a stable schedule and controlled volume while account, list, message, and recipient conditions remain matched.

Layer 6: measurement

Sometimes the system changed. Sometimes the dashboard did.

Check how each metric is produced:

  • Does “sent” mean queued, handed to the provider, or accepted?
  • Does “delivered” exclude only hard bounces, or come from provider confirmation?
  • Are replies matched across forwarding and aliases?
  • Did a tracking-domain or webhook change interrupt events?
  • Did the sequencer change its calculation or reporting window?
  • Are automatic replies included?
  • Are privacy opens or scanner clicks inflating engagement?
  • Did the campaign, mailbox, or time-zone filter change?

Reconcile a small sample across systems. Take ten or twenty message IDs and trace them from the sequencer to provider events, bounces, replies, and the final inbox where possible.

A dashboard total without raw examples is not enough to diagnose a failure.

Evidence that clears this layer: message-level samples reconcile across systems, event collection is intact, and the trend survives a consistent definition and date window.

Use a decision log, not a pile of theories

For every test, record:

FieldEntry
SymptomNarrow, measurable statement
ScopeDomain, mailbox, campaign, provider, dates
Layer testedOne of the six layers
Evidence beforeHeaders, errors, counts, configuration
ChangeOne controlled change
ControlWhat stayed the same
Evidence afterSame measurements and path
DecisionCleared, implicated, or inconclusive
Next ownerPerson and exact follow-up

“We changed providers and it got better” is not a useful diagnosis if the domain, accounts, DNS, schedule, data, and copy changed too.

Stop conditions

Pause the test and escalate when:

  • an account is disabled or restricted;
  • provider errors indicate a policy or security event;
  • a domain is not under your control;
  • DNS changes are being made from more than one place;
  • hard bounces or complaints rise abruptly;
  • you cannot identify the sending source;
  • a client or customer asset is involved and the owner has not approved the change;
  • evidence contains credentials, private messages, or personal data that should not be shared in a worklog.

Redact sensitive material. Preserve the domain, result, timestamp, provider, and error family needed to understand the issue.

The order of operations

When the next “deliverability is down” message arrives, use this order:

  1. Narrow the symptom.
  2. Capture the evidence packet.
  3. Check authentication on a real message.
  4. Check provider and account state.
  5. Segment results by list source.
  6. Inspect the exact message and recipient fit.
  7. Check schedule and total sending behavior.
  8. Reconcile the measurement.
  9. Change one layer.
  10. Retest the same path and write down the result.

You may still decide to change providers, lists, copy, or sending software. The difference is that you will know which problem the change is meant to solve.

Sources