Email Provider Rate Limits: Protect the Queue Instead of Flooding Retries
Handle email API rate limits with bounded queues, fair scheduling and clear separation between short-term request limits and account quotas.
TL;DR
- Treat a provider rate limit as a scheduling decision: stop flooding retries and let a queue absorb temporary pressure while preserving message priority, fairness and expiry rules.
- Use a queued, class-separated scheduler with bounded capacity per class and tenant to enforce product-level fairness and prevent a single workload from starving critical messages.
- Honor provider guidance and test capacity changes: persist next-attempt timestamps instead of sleeping workers, add deadlines and define operational triggers like sustained queue growth or increasing oldest age.
Rate limits are a scheduling input
An email provider rate limit tells your application that it cannot submit work at the current pace under the relevant constraint. The right response is usually a scheduling decision, not a larger burst of retries. A queue should absorb temporary pressure while preserving message priority, fairness and expiry rules.
First identify the type of limit. Requests per interval, recipients per request, account sending allowance and reputation-related restrictions describe different constraints. A smaller batch may fix a request-size issue but not an exhausted account allowance. More workers can worsen a request-rate limit rather than increase throughput.
Build your integration so the response category is visible to the scheduler. A generic “provider failed” error hides the distinction and encourages recovery behavior that cannot succeed.
Related reading: How to Reduce Email Bounce Rate with Better Diagnostics.
Read the documented scope
A limit may apply to an account, credential, source address, endpoint or another scope. Your application should coordinate at the appropriate boundary. If several workers share the same allowance, a separate limiter in each process may permit too much aggregate traffic.
SendDart's documentation distinguishes plan quotas from per-request limits and exposes structured quota errors. Consult the current account and API documentation for the actual limits relevant to your deployment rather than copying an old value into permanent code. SendDart quota documentation.
Keep the configuration reviewable. Record why a worker concurrency or pacing value was chosen and which provider constraint it addresses. If the provider changes a limit or the account moves to another plan, the team should know which setting to revisit.
Related reading: Best Email API: Choose With a Production Acceptance Test.
Use a queue with deliberate fairness
Imagine two fictional tenants: one imports several thousand contacts while another sends occasional account invitations. A single first-in-first-out queue may delay the invitations behind the import. Provider capacity alone does not solve this product-level fairness problem.
Separate message classes where justified, then allocate bounded capacity across them. Critical access messages and optional digests can have different deadlines. Avoid claiming that priority gives a guarantee the provider has not made; it only controls how your application uses available submission capacity.
Track oldest pending age by class and tenant. A global average can hide starvation. Also cap how much work a single tenant or request can enqueue, so a bug does not create an unbounded backlog before the provider has a chance to reject it.
Honor delay guidance safely
HTTP 429 describes a rate-limiting response, and Retry-After can provide a time or interval for a later attempt. Implement the documented response handling and use a bounded fallback when guidance is absent or malformed. HTTP 429, Retry-After.
Persist a next-attempt timestamp instead of sleeping inside a scarce worker for a long interval. Add a maximum age or deadline for time-sensitive messages. A verification message that is no longer useful should not continue circulating merely because the retry count is low.
Consider jitter when many jobs become eligible at once. The goal is to avoid synchronized bursts, not to bypass the provider's restrictions. Do not rotate credentials or source addresses to evade limits; resolve the workload or account-capacity mismatch through supported means.
Keep quotas separate from transient rate rejection
An account allowance exhausted over a longer window may require waiting, reducing unnecessary work or changing the plan. Repeating the same request every few seconds does not create new capacity. Surface the condition to operators and, when appropriate, to the product workflow.
A quota rejection can also affect inbound and outbound usage differently depending on the provider's terms. Read the actual accounting definition. Do not assume that one HTTP batch consumes one unit of sending volume or that testing activity is excluded automatically.
For SendDart, the SDK exposes structured quota information documented in its README. Preserve that information in the adapter instead of parsing a human-readable message. It can help an operator distinguish an allowance issue from malformed content or sender configuration. SendDart SDK errors.
Preserve operation identity when retrying
A known pre-processing rate rejection is different from an uncertain submission. Inspect the full documented response and any progress evidence before retrying. A request that includes an original message ID or partial results should not be treated as untouched work merely because an error occurred.
For supported idempotent operations, retain the same operation key and payload during recovery. Do not rebuild a batch with changed ordering under the same identity, and do not assign a new key simply to avoid a conflict. Reconcile partial progress before selecting known-unattempted items.
The scheduler should carry these recovery requirements with the job. An operator resuming a queue after a limit change must not accidentally erase the distinction between rejected, accepted and uncertain work.
Test capacity changes before a launch
Use a fake provider boundary to test rate responses, long delay guidance and a queue that resumes after a pause. Confirm that critical work is not starved and that retries stay within their budget. Test multiple workers sharing one simulated allowance.
Run a controlled live canary within the provider's permitted limits to validate configuration. Do not discover limits by flooding the service or sending test mail to real customers. A load test should have an explicit recipient and resource plan.
Finally, define an operational trigger for capacity review: sustained queue growth, increasing oldest age or a forecast change. Good rate-limit handling makes demand visible and orderly. It turns a temporary submission constraint into managed work rather than a retry storm.