Email Service Migration Checklist: A Runbook for a Controlled Cutover

Email Service Migration Checklist: A Runbook for a Controlled Cutover

Plan email migration with dependency mapping, suppression reconciliation, deterministic routing and explicit rollback rules for uncertain sends.

SendDart Team

TL;DR

  • Treat completion as a defined state: migration is complete only when new notifications use the intended service, old operations remain explainable, and recipient protections are preserved; write completion criteria first.
  • Follow the runbook phases: inventory all sending entry points and define an application-facing notification contract, then persist the provider selection with each notification to avoid recomputing on retries.
  • Use controlled verification and explicit stop conditions: validate imports with a small known set, avoid bulk live sends to discover failures, and set pilot stop conditions tied to meaningful operational thresholds.

Define what a completed migration means

An email migration is complete when new notifications use the intended service, old operations remain explainable and recipient protections are preserved. A successful test message does not establish those conditions. Write the completion criteria before changing credentials so the rollout has a clear destination.

Name an owner for application code, sender configuration, event ingestion and support readiness. One person may cover several roles in a small team, but the responsibilities should still be explicit. Record the reason for migrating and the requirement the new service must satisfy.

This runbook applies to application email migrations. It is not a vendor-specific claim that every configuration can be copied directly. Use the current source and destination documentation for their exact sender, template, event and recovery contracts.

Phase one: inventory and freeze the boundaries

List every sending entry point: web handlers, workers, scheduled jobs, administrative scripts and third-party integrations. Search for credentials, client construction and provider-specific message identifiers without printing secret values into reports. A forgotten scheduled script can keep using the old service after the main application moves.

Inventory templates, sender domains, receiving addresses, callbacks, suppressions and support dashboards. Record which components are active and which are obsolete. Do not migrate dead configuration merely because it exists, but retain the operational evidence needed for historical messages.

Define an application-facing notification contract before rewriting adapters. It should identify the business purpose, recipient, immutable content or template version, operation identity and selected provider. This gives the migration a stable boundary and makes rollback less dependent on scattered code changes.

Phase two: prepare the destination

Set up the new service through its documented process. Verify sender ownership and required DNS records without deleting unrelated records. Store credentials in trusted server configuration and test that the runtime, not the browser, receives them.

Render representative messages with fictional data. Compare subject, text alternative, HTML, links and attachments. Include edge cases such as long names, missing optional fields and non-ASCII text. A provider migration should not accidentally change the meaning of a receipt or invalidate a reset link.

For SendDart, inspect the official SDK's result and recovery contracts. Supported sending operations can use stable idempotency keys, but other resource operations do not inherit that guarantee merely because an option is accepted. SendDart SDK documentation.

Phase three: reconcile recipient protections

Export or query suppression information through supported access. Map reasons and scope into the destination and your application policy. A permanent delivery failure, a complaint and a newsletter preference are not interchangeable facts.

Keep original timestamps and source identifiers where available. Review ambiguous mappings instead of defaulting them to permission to send. If a customer opted out of a purpose, changing infrastructure should not silently reverse that decision.

Validate the import with a small known set before processing all records. Test a suppressed address in a controlled environment and confirm that the application's eligibility decision remains correct. Do not perform a bulk live send to discover whether the migration preserved protections.

Phase four: prove recovery before traffic

Test known rejection, duplicate job delivery, duplicate event delivery and an interrupted send. The application should preserve its original operation identity through recovery. If the provider returns partial progress or an original message identifier, retain that evidence and reconcile it before creating new work.

Simulate a crash after acceptance but before the database update. This is the point at which naive “retry failed jobs” logic often produces duplicates. A correct runbook distinguishes known-unsent work from uncertain work rather than treating both as a clean slate.

Verify support lookup for both providers. A customer may ask about an old message after the new integration is active. Store provider name with message ID and keep the historical event path available for the necessary investigation window.

Related reading: Best Email API: Choose With a Production Acceptance Test.

Related reading: Scheduled Email Delivery: Store the Time, Content and Cancellation State.

Phase five: route one bounded cohort

Choose a message class and a deterministic cohort rule. Persist the provider selection with each notification before submission. Recomputing the choice on every retry can move one logical event between services and undermine duplicate prevention.

Set explicit stop conditions for the pilot. Examples include uncorrelated events, unexpected sender rejection or a growing queue of uncertain operations. These are suggested categories; choose thresholds that match your workload and customer impact rather than copying arbitrary percentages.

Monitor queue age, rejection reasons and event correlation. Do not judge the migration solely from opens or a provider dashboard's total count. Reconcile application intent with actual provider attempts so missing and duplicate operations can be identified.

Phase six: roll back only what is safe to move

A rollback should redirect future or demonstrably unsubmitted notifications. It should not automatically resend every uncertain operation through the old service. The new service may already have accepted those messages.

Write a decision table for pending, accepted, rejected and uncertain states. Operators should know which states can be moved, which require the original operation key and which need investigation. Give the rollback tool the same authorization and audit controls as ordinary sending.

After the pilot passes, expand in deliberate steps. Continue consuming old events while earlier messages remain relevant. Revoke obsolete credentials only after confirming that no active job or operational process still needs them.

Keep the migration evidence

Save the mapping inventory, controlled tests, rollout decisions and unresolved exceptions. Mark each requirement as verified, intentionally changed or still open. This is more useful than declaring the migration done because deployment succeeded.

A controlled cutover leaves the application with clear message identity, preserved recipient protections and an auditable provider boundary. Those improvements are valuable even if the final decision is to remain with the original provider until a missing requirement is resolved.