Email Provider SLAs: Separate Credits From Recovery Capabilities

Email Provider SLAs: Separate Credits From Recovery Capabilities

Read email-provider SLAs alongside operational recovery requirements, separating covered availability, credit remedies and the work your application owns.

SendDart Team

TL;DR

  • Decide SLA value by separating bill credits from operational recovery: evaluate the commercial remedy and the application-level recovery as distinct purchasing requirements your team can use.
  • Assess providers with exercises and scorecards: run a tabletop recovery exercise to simulate outages, then compare a commercial and an operational scorecard to reveal practical gaps.
  • Verify recovery by demanding evidence, preserved records, and clear ownership so the team can determine which messages need reconciling and when to claim credits or take action.

An SLA is one part of reliability procurement

A service-level agreement can define a provider's commitment and a remedy when a covered target is missed. It does not by itself tell your application what to do with a delayed password reset or an uncertain send result. Evaluate the agreement and the recovery workflow as related but separate buying requirements.

The technical team should identify the service boundaries and evidence it needs. The appropriate contract reviewer should interpret the actual agreement. This guide provides an operational reading framework, not a substitute for that review or a claim that all providers use the same terms.

Identify the service the number describes

Start with the covered service. Does the commitment concern API availability, processing, a particular endpoint or another defined measure? Ask how it is measured and over which period. A percentage without those definitions is difficult to connect to your application.

Consider a hypothetical incident in which the API accepts requests but processing is delayed. Your users may experience a serious problem even if the availability measurement treats the endpoint as reachable. The agreement's definition determines whether that scenario is covered; your product's recovery requirements determine what you must do regardless.

Write several relevant scenarios and ask the provider to map them to the agreement. Do not infer coverage from a broad marketing phrase such as reliable delivery.

Related reading: Best Email API: Choose With a Production Acceptance Test.

Read exclusions and eligibility with the target

An agreement may define eligible plans, customer responsibilities, excluded events and claim requirements. These details affect the practical value of the commitment. Read them at the same time as the headline target rather than treating them as paperwork to inspect after an incident.

Resend's Enterprise Terms include service-level provisions and a process for requesting credits. Use the current document and your actual order terms when evaluating that offering. A commitment for one tier should not be assumed to apply to another account or provider.

Preserve the applicable version and identify who in your organization tracks eligibility and deadlines. Engineering may hold the incident evidence while finance owns the claim, so the handoff needs to be explicit.

Separate a credit from operational recovery

A service credit can reduce a future bill. It does not resend an expired invitation, explain a delayed notification to a customer or reconstruct evidence your application did not retain. Those are operational tasks that need their own design.

For each important message type, define what remains useful after a delay. An order update might still matter later; a short-lived action link may need a fresh generation path. Do not simply replay every queued message when service returns. The underlying business state may have changed.

This does not mean an SLA is unimportant. It means the commercial remedy and the user-facing recovery solve different problems. A buyer should evaluate both instead of treating a higher percentage as a complete resilience strategy.

Ask for evidence during an incident

Determine which provider records show acceptance, processing state and later outcomes. Ask how customers receive incident updates and whether the status page separates components relevant to your application. Confirm the identifiers and timestamps support needs for a focused investigation.

Your application should retain enough information to distinguish messages that were never submitted from messages with uncertain outcomes. A timeout does not establish that nothing happened. Recovery behavior should follow the provider's documented contract and your own durable business record.

For scheduled mail, preserve the intended time and cancellation state. SendDart's existing scheduled delivery guide discusses those distinctions. They remain relevant when assessing any provider's incident recovery behavior.

Run a tabletop recovery exercise

Use a fictional outage rather than disrupting production. Suppose sending is impaired for an hour while invitations, billing notices and a product campaign accumulate. Ask the team to decide which work pauses, which messages expire and who communicates with customers.

Then simulate recovery. Which messages can resume automatically? Which require a fresh eligibility check? Which have an uncertain prior outcome and need reconciliation? Record the evidence required for each decision and identify whether the provider exposes it.

This exercise can reveal a more important buying gap than a small difference in advertised availability. If your team cannot determine whether a message already left, a rapid retry path may create customer confusion even after the service is healthy.

Compare providers with two scorecards

Use a commercial scorecard for covered service, measurement period, exclusions, eligible plan and remedy process. Use an operational scorecard for evidence, status communication, queue behavior, investigation support and recovery decisions.

Mark unverified answers as unresolved. Avoid assigning a universal reliability rank based on an SLA alone, and do not treat a historical status page as a guarantee of future behavior. If you conduct a controlled test, describe exactly what it established and its limitations.

Apply the same questions to SendDart and alternatives. A provider without a requirement your organization considers mandatory should not receive an assumed pass because another feature is attractive.

Buy a commitment your team can use

The final decision should explain what the provider promises, how your organization would establish a covered failure and how the application protects users while service is impaired. Assign owners for both the incident response and any commercial follow-up.

A useful SLA is understandable and relevant to the service you depend on. A useful recovery plan is executable with the evidence your systems actually retain. Buying them together gives the team a clearer operating model than relying on a percentage to stand in for the whole customer experience.