Engineering · September 21, 2026
Designing Email Webhooks for At-Least-Once Delivery
Learn how to build resilient webhook consumers for email events using retries, idempotency keys, and signature verification to ensure no delivery event is ever missed.
The Challenge of Reliable Event Delivery
To achieve at-least-once delivery for email webhooks, you must implement a system where the sender retries failed requests with exponential backoff and the receiver ensures idempotency. Because networks are unreliable and servers crash, you cannot assume a single HTTP 200 OK guarantees the event was processed. Reliability is built by combining a persistent retry queue on the sender side with a deduplication layer on the receiver side.
When you integrate an email API like SendHQ, your application needs to know when an email was delivered, bounced, or marked as spam. These events are asynchronous. If your webhook endpoint is down for five minutes during a traffic spike, you could lose thousands of critical delivery signals. This creates a data gap in your analytics and prevents your system from reacting to bounces (which is critical for maintaining sender reputation).
The Anatomy of a Reliable Webhook
A robust webhook architecture consists of three primary pillars: signature verification, idempotent processing, and a retry strategy.
1. Signature Verification
Never trust a POST request to your webhook endpoint based solely on the IP address or the presence of an API key in the body. Attackers can spoof these. Instead, use a HMAC (Hash-based Message Authentication Code) signature.
The sender signs the payload using a shared secret and attaches the signature to a header (e.g., X-SendHQ-Signature). The receiver recalculates the hash using the same secret and compares it to the header.
const crypto = require('crypto');
function verifySignature(payload, signature, secret) {
const expectedSignature = crypto
.createHmac('sha256', secret)
.update(payload)
.digest('hex');
// Use timingSafeEqual to prevent timing attacks
return crypto.timingSafeEqual(Buffer.from(signature), Buffer.from(expectedSignature));
}
2. Idempotency and Deduplication
At-least-once delivery means the sender will keep sending the event until it receives a success response. If your server processes the event but crashes before sending the 200 OK, the sender will send the event again. Without idempotency, you might count a single delivery as two deliveries in your database.
Every event must have a unique event_id. You should use an idempotency key pattern to track processed events.
The Workflow:
- Receive the webhook payload.
- Check if the
event_idexists in yourprocessed_eventstable. - If it exists, return 200 OK immediately and ignore the body.
- If it does not exist, process the event and record the
event_idin a single transaction.
3. The Retry Strategy
From the sender's perspective, a retry policy is mandatory. A standard pattern is exponential backoff with jitter. For example: retry after 1 minute, 5 minutes, 30 minutes, 2 hours, and 12 hours.
If the receiver returns a 4xx error (except 429), it usually indicates a client error (like a bad signature), and retrying will not help. A 5xx error or a timeout indicates a transient failure where retries are essential.
Concrete Payload Example
Here is a typical delivery event payload you might receive from SendHQ:
{
"event_id": "evt_12345abcde",
"event_type": "delivered",
"timestamp": "2026-09-15T10:00:00Z",
"message_id": "msg_98765xyz",
"recipient": "user@example.com",
"metadata": {
"order_id": "ord_5544"
}
}
Handling Failures and Edge Cases
The "Slow Consumer" Problem
If your webhook handler performs heavy database writes or calls other external APIs synchronously, your endpoint will time out. This triggers the sender's retry logic, leading to a "retry storm" that can crash your server.
The Solution: Decouple acceptance from processing.
- Receive the webhook.
- Verify the signature.
- Push the raw payload into a message queue (like RabbitMQ, SQS, or Redis).
- Return 200 OK immediately.
- A separate worker process consumes the queue and updates your database.
The Agent Readiness Problem
When AI agents are triggered by webhooks, the risk of infinite loops increases. If an agent receives a "delivered" event and responds by sending another email, which then triggers another "delivered" event, you have a loop.
Treat sending email as an external side effect. Agents should never send emails automatically based on a webhook without a human-in-the-loop approval or a strict state-machine check to ensure the action is necessary.
Comparing the Ecosystem
When choosing a provider, reliability is often tied to how they handle these events and what they charge for the volume of mail that generates these events.
For high-volume transactional mail, the cost difference is stark. According to Amazon SES pricing, a la carte sending costs 0.10 USD per 1,000 emails. In contrast, Postmark pricing starts at 15 USD per month for 10,000 emails, with overages between 1.20 and 1.80 USD per 1,000. For a volume of 50,000 emails, SES a la carte costs roughly 5 USD, while Postmark tiers would cost approximately 66 USD.
Other options include Resend, which offers a free tier of 3,000 emails per month (capped at 100 per day) and a Pro plan at 20 USD per month for 50,000 emails. Mailgun starts at 15 USD per month for 10,000 emails. SendGrid has moved its free tier to a 60-day trial, with Essentials starting at 19.95 USD per month.
Regardless of the provider, the reliability of your consumption of these events is what determines your data integrity.
Implementation Checklist for Engineers
- Signature Verification: Is the payload verified using a shared secret and a constant-time comparison function?
- Asynchronous Processing: Does the endpoint return 200 OK before performing heavy business logic?
- Idempotency: Is there a unique constraint on
event_idto prevent duplicate processing? - Timeout Management: Is the timeout set lower than the provider's timeout to avoid overlapping retries?
- Monitoring: Do you have alerts for a spike in 5xx responses on your webhook endpoint?
- DNS Health: Are your receiving servers configured correctly? Use tools like the SendHQ Email DNS Checker to ensure your infrastructure is reachable and correctly configured.
- Authentication Standards: Have you implemented DKIM, SPF, and DMARC to ensure your outbound mail is accepted, reducing the number of "bounce" webhooks you have to handle?
Summary of Tradeoffs
Approach | Pros | Cons
Synchronous Processing | Simple to implement, immediate consistency | High risk of timeouts, prone to retry storms
Queue-based Processing | Highly scalable, resilient to spikes | Increased infrastructure complexity, eventual consistency
Simple Logging | Low overhead | No way to recover from missed events without manual logs
Idempotency Table | Guaranteed data integrity | Extra database write per event
Final Thoughts
Reliability in email webhooks is not about preventing failures, but about designing for them. By assuming that the network will fail and that events will be delivered more than once, you build a system that is truly resilient. Whether you are managing SPF records for a small project or scaling a massive transactional system, the patterns of signature verification and idempotency remain the gold standard.
Build your email infrastructure with SendHQ.