Switching Email Providers Without Duplicate Sends: Use a Routing Ledger

Switching Email Providers Without Duplicate Sends: Use a Routing Ledger

Prevent duplicate messages during provider migration or failover with persistent provider selection, operation identity and reconciliation of uncertain attempts.

SendDart Team

TL;DR

  • Use an application-owned routing ledger as the authoritative record to prevent duplicate sends during a provider switch, retaining intent and evidence while controlling whether work moves between services.
  • Decide and persist the provider before submission and record every attempt without overwriting history so retries consult stored choices and late callbacks correlate to the original operation.
  • Treat uncertain outcomes as held work and verify with reconciliation and targeted tests; validate failover by simulating lost responses and asserting the fallback does not auto-send.

The dangerous moment is an unknown first outcome

A provider switch can duplicate mail when the application cannot tell whether the original service accepted a request. Sending the same payload through a second service may seem like resilience, but both services can complete independently. Their idempotency systems do not generally share state.

This article focuses on that specific boundary rather than repeating a full migration checklist. The proposed solution is an application-owned routing ledger: a durable record of each intended notification, its selected provider and the evidence collected from attempts.

The ledger does not create a universal exactly-once guarantee. It gives your application a controlled basis for deciding whether work can move, must remain attached to the original provider or needs manual reconciliation.

Select the provider before submission

Assign the provider when the notification becomes eligible for execution and persist that choice atomically with the claim or routing decision. A retry should read the stored choice rather than recompute a percentage rollout rule.

For example, a fictional rollout routes a bounded cohort to the new service. If the cohort calculation changes tomorrow, a notification already attempted yesterday should not silently move to another provider. New notifications can follow the new rule; existing attempts retain their history.

Use business event and message purpose to prevent duplicate notification creation before routing. Two separate notification records with two provider choices cannot be reconciled reliably merely by noticing they have the same recipient and subject.

Store attempts without overwriting history

The ledger should retain notification ID, selected provider, operation key, content version, attempt stage and provider message ID when known. Keep timestamps and machine-readable response categories.

Do not overwrite the old provider ID when a later authorized operation uses another service. A late callback from the first provider still needs to correlate with the correct attempt. Store provider name with the external identifier to avoid ambiguity.

Separate the business notification's intent from individual network calls. Several recovery calls using one supported operation identity may represent one provider operation. A deliberate new resend is a different intent and should have its own reason and audit record.

Related reading: Scheduled Email Delivery: Store the Time, Content and Cancellation State.

Define which states may change providers

Known-unattempted work is the clearest candidate for rerouting. A documented pre-processing rejection may also allow a new decision under your policy, provided no progress evidence contradicts it. Accepted or uncertain operations require different handling.

Original evidenceRouting treatment
No submission attemptedMay select the current approved provider
Confirmed rejection before processingReview retry or reroute policy
Accepted with message identityContinue tracking the original provider
Partial batch progressReconcile each item before new work
Unknown outcomeStop automatic cross-provider resend

This table is a design framework. The adapter must translate each provider's actual contract into these evidence classes rather than inferring safety from status code alone.

Use idempotency within its documented scope

For supported SendDart send operations, preserve the same key and unchanged payload when recovering the original operation. Its documentation describes conflicts and partial-progress handling that the application must respect. SendDart SDK recovery contract.

A key sent to another provider is not proof of cross-service deduplication. Even if both services support a header with a similar name, their stores and retention rules are separate. Your ledger is what prevents the application from authorizing a second operation prematurely.

Keep content immutable during recovery. If a template or recipient changes, decide whether the business has authorized a new notification. Do not use a provider switch as an excuse to mutate an uncertain old operation without preserving its evidence.

Reconcile before releasing held work

When a request is uncertain, use any original message identifier and documented retrieval or event mechanism to establish what happened. Retain unresolved work visibly if the evidence remains insufficient.

A reconciliation worker should have a bounded schedule and an escalation path. It should not repeatedly create sends while pretending to investigate them. Operators need to know what evidence is missing and which action is safe.

For batches, keep item-level identities. Some items may be confirmed, others attempted but uncertain, and others never attempted. Only the known-unattempted portion should become eligible for a fresh operation after review of the contract.

Keep old callbacks alive during the transition

A provider that no longer receives new work can still report events for earlier messages. Retain its verified callback path and correlation records for the required operational window. Disabling them immediately at cutover loses evidence needed to resolve held work.

Normalize events into application concepts while preserving original provider data in restricted records. A delivery event from the old service should not update an unrelated new-service attempt merely because the recipient matches.

Carry suppression protections into the new routing path before moving traffic. The routing ledger solves duplicate authorization, not recipient eligibility. Both checks must pass before a worker submits a message.

Test the exact failover race

Simulate the first provider accepting a request while the response is lost. Trigger the fallback path and assert that it does not send through the second provider automatically. Then deliver a late event from the first provider and verify that the ledger resolves the original attempt.

Test a confirmed pre-processing rejection separately and confirm that the policy can make the intended retry or routing decision. Add concurrent workers, a deployment changing rollout percentages and a partially completed batch.

The desired result is not that every message instantly moves during an outage. It is that every intended notification has one explainable provider path, with uncertainty held rather than multiplied. That discipline makes migrations and failover safer for the people receiving your mail.

Related reading: Best Email API: Choose With a Production Acceptance Test.