Email Sending Load Test Checklist: Measure the Queue Without Spamming Inboxes
Load-test email infrastructure using controlled transports, bounded canaries and clear metrics for queue age, recovery and provider capacity.
TL;DR
- Decide the test goal up front and avoid sending real customer mail: separate internal capacity checks from external integration and use controlled canaries or fake providers instead of uncontrolled high-volume sends.
- Model realistic workloads and instrument thoroughly: shape fictional data to production distributions, measure queue and rendering metrics, and use a controlled provider simulator with saved scenarios.
- Define limits and success checks: increase input gradually to find the unsafe boundary, set alerts on predictive metrics, and choose production concurrency below the observed unsafe limit so tests produce a capacity-and-recovery story.
Define what the test is meant to prove
A load test can measure application queue throughput, template-rendering cost, database contention or behavior under provider rate limits. It should not be an uncontrolled attempt to send as much mail as possible. State the question before choosing the test volume.
Separate internal capacity testing from external integration testing. Most high-volume experiments can use a fake provider boundary that records requests and returns controlled responses. A smaller canary can validate the actual service contract within permitted limits using addresses your team controls.
Do not use real customers, purchased lists or random addresses as load-test recipients. The test should have an explicit data, recipient and resource plan, including how it stops if behavior differs from expectations.
Build a realistic workload model
Describe the message mix: small account notices, longer receipts, attachment-bearing messages and any batches. Include the expected busiest interval, not only the monthly total. A thousand evenly spaced messages and a thousand simultaneous jobs stress different boundaries.
Use fictional data shaped like production records. Preserve relevant size distributions and template complexity without copying private customer content. Include long fields and optional-data combinations that affect rendering cost.
Model priorities and tenants if the production system has them. A single uniform stream cannot show whether a large import starves reset messages or whether one tenant consumes all worker capacity.
Instrument before increasing volume
Measure queue depth, oldest pending age, claim latency, render duration, provider-attempt duration and result-persistence time. Keep the dimensions bounded and avoid recipient addresses as metric labels.
Track uncertain operations and duplicate external attempts separately. A test can appear fast if it silently drops work or submits duplicates. Reconcile the count of intended notifications with accepted, rejected, skipped and unresolved outcomes.
Record CPU, memory and database pressure using your normal operational tools. The limiting component may be template rendering or database claims rather than the provider. A larger worker pool can make contention worse without improving useful throughput.
Use a controlled provider simulator
A local fake transport can emulate immediate acceptance, delayed responses, rate rejection, malformed responses and connection loss. Configure the behavior deliberately and save the scenario with the test results.
Inject a shared simulated rate allowance across workers to test coordination. If each worker independently assumes it owns the full allowance, the test should expose the resulting burst and retry pattern.
SendDart also documents reserved simulator recipients for integration outcomes. Those addresses help exercise service-specific behavior without ordinary recipient delivery, but they are not a license to run unbounded traffic or a measure of real inbox placement. Follow current account limits and testing guidance. SendDart quickstart.
Related reading: Best Email API: Choose With a Production Acceptance Test.
Test backpressure and fairness
Increase input gradually while watching queue age and recovery. Identify the point where the system no longer keeps up, then confirm that it responds predictably. Producers may need bounded admission, delayed work or a clear capacity error instead of unlimited enqueueing.
Mix critical and non-urgent messages. Verify that the scheduling policy meets your stated priorities without permanently starving lower-priority work. Test multiple tenants and observe whether one workload monopolizes database claims or provider capacity.
Do not interpret an empty in-memory channel as proof that all work completed. The durable notification records and external-attempt ledger should provide the authoritative reconciliation.
Exercise failure during sustained work
Restart a worker while sends are active. Simulate a response loss after submission and a database failure before result persistence. Confirm that recovery preserves original operation identities and does not replay every message under new keys.
For supported SendDart send operations, use the documented idempotency contract and partial-progress evidence. A batch interrupted after several items requires item-level reconciliation. SendDart SDK recovery documentation.
Also stop the webhook consumer temporarily, then resume it with duplicate or delayed events. The system should retain authenticated evidence and process it idempotently. Sending throughput is only half the operational path.
Related reading: Email Domain Warmup: Build a Measured Sending Plan.
Keep live testing small and authorized
A live canary should validate credentials, sender configuration and real response correlation with a bounded number of messages. Use a separate test identity or environment where appropriate, and avoid mixing results into customer analytics.
Confirm the service's current rate and quota rules before the test. Do not evade restrictions by rotating keys or source addresses. If you need a larger coordinated test, use the provider's supported process and an explicit capacity plan.
Label the findings accurately. A local simulator sustaining a particular throughput demonstrates your application under that simulator's behavior. It does not prove that the provider will deliver the same volume at the same speed or place messages in every inbox.
Turn results into operating limits
Document the workload, environment, software version and bottleneck. Choose an initial production concurrency below an observed unsafe boundary with room for normal variation. The exact margin is an engineering decision, not a universal percentage from a blog.
Set alerts on the measurements that predicted trouble, especially oldest age and unresolved attempts. Keep the test repeatable so a worker, template or SDK change can be assessed against the same scenario.
A useful load test produces a capacity and recovery story, not just a messages-per-second headline. It shows what the system accepts, how it slows down, how it recovers and whether every intended notification remains explainable.