Engineering · September 21, 2026

Designing Email Webhooks for At-Least-Once Delivery

Learn how to build resilient webhook consumers for email events using retries, idempotency keys, and signature verification to ensure no delivery event is ever missed.

The Challenge of Reliable Event Delivery

To achieve at-least-once delivery for email webhooks, you must implement a system where the sender retries failed requests with exponential backoff and the receiver ensures idempotency. Because networks are unreliable and servers crash, you cannot assume a single HTTP 200 OK guarantees the event was processed. Reliability is built by combining a persistent retry queue on the sender side with a deduplication layer on the receiver side.

When you integrate an email API like SendHQ, your application needs to know when an email was delivered, bounced, or marked as spam. These events are asynchronous. If your webhook endpoint is down for five minutes during a traffic spike, you could lose thousands of critical delivery signals. This creates a data gap in your analytics and prevents your system from reacting to bounces (which is critical for maintaining sender reputation).

The Anatomy of a Reliable Webhook

A robust webhook architecture consists of three primary pillars: signature verification, idempotent processing, and a retry strategy.

1. Signature Verification

Never trust a POST request to your webhook endpoint based solely on the IP address or the presence of an API key in the body. Attackers can spoof these. Instead, use a HMAC (Hash-based Message Authentication Code) signature.

The sender signs the payload using a shared secret and attaches the signature to a header (e.g., X-SendHQ-Signature). The receiver recalculates the hash using the same secret and compares it to the header.

const crypto = require('crypto'); function verifySignature(payload, signature, secret) { const expectedSignature = crypto .createHmac('sha256', secret) .update(payload) .digest('hex'); // Use timingSafeEqual to prevent timing attacks return crypto.timingSafeEqual(Buffer.from(signature), Buffer.from(expectedSignature)); }

2. Idempotency and Deduplication

At-least-once delivery means the sender will keep sending the event until it receives a success response. If your server processes the event but crashes before sending the 200 OK, the sender will send the event again. Without idempotency, you might count a single delivery as two deliveries in your database.

Every event must have a unique event_id. You should use an idempotency key pattern to track processed events.

The Workflow:

  1. Receive the webhook payload.
  2. Check if the event_id exists in your processed_events table.
  3. If it exists, return 200 OK immediately and ignore the body.
  4. If it does not exist, process the event and record the event_id in a single transaction.

3. The Retry Strategy

From the sender's perspective, a retry policy is mandatory. A standard pattern is exponential backoff with jitter. For example: retry after 1 minute, 5 minutes, 30 minutes, 2 hours, and 12 hours.

If the receiver returns a 4xx error (except 429), it usually indicates a client error (like a bad signature), and retrying will not help. A 5xx error or a timeout indicates a transient failure where retries are essential.

Concrete Payload Example

Here is a typical delivery event payload you might receive from SendHQ:

{ "event_id": "evt_12345abcde", "event_type": "delivered", "timestamp": "2026-09-15T10:00:00Z", "message_id": "msg_98765xyz", "recipient": "user@example.com", "metadata": { "order_id": "ord_5544" } }

Handling Failures and Edge Cases

The "Slow Consumer" Problem

If your webhook handler performs heavy database writes or calls other external APIs synchronously, your endpoint will time out. This triggers the sender's retry logic, leading to a "retry storm" that can crash your server.

The Solution: Decouple acceptance from processing.

  1. Receive the webhook.
  2. Verify the signature.
  3. Push the raw payload into a message queue (like RabbitMQ, SQS, or Redis).
  4. Return 200 OK immediately.
  5. A separate worker process consumes the queue and updates your database.

The Agent Readiness Problem

When AI agents are triggered by webhooks, the risk of infinite loops increases. If an agent receives a "delivered" event and responds by sending another email, which then triggers another "delivered" event, you have a loop.

Treat sending email as an external side effect. Agents should never send emails automatically based on a webhook without a human-in-the-loop approval or a strict state-machine check to ensure the action is necessary.

Comparing the Ecosystem

When choosing a provider, reliability is often tied to how they handle these events and what they charge for the volume of mail that generates these events.

For high-volume transactional mail, the cost difference is stark. According to Amazon SES pricing, a la carte sending costs 0.10 USD per 1,000 emails. In contrast, Postmark pricing starts at 15 USD per month for 10,000 emails, with overages between 1.20 and 1.80 USD per 1,000. For a volume of 50,000 emails, SES a la carte costs roughly 5 USD, while Postmark tiers would cost approximately 66 USD.

Other options include Resend, which offers a free tier of 3,000 emails per month (capped at 100 per day) and a Pro plan at 20 USD per month for 50,000 emails. Mailgun starts at 15 USD per month for 10,000 emails. SendGrid has moved its free tier to a 60-day trial, with Essentials starting at 19.95 USD per month.

Regardless of the provider, the reliability of your consumption of these events is what determines your data integrity.

Implementation Checklist for Engineers

  • Signature Verification: Is the payload verified using a shared secret and a constant-time comparison function?
  • Asynchronous Processing: Does the endpoint return 200 OK before performing heavy business logic?
  • Idempotency: Is there a unique constraint on event_id to prevent duplicate processing?
  • Timeout Management: Is the timeout set lower than the provider's timeout to avoid overlapping retries?
  • Monitoring: Do you have alerts for a spike in 5xx responses on your webhook endpoint?
  • DNS Health: Are your receiving servers configured correctly? Use tools like the SendHQ Email DNS Checker to ensure your infrastructure is reachable and correctly configured.
  • Authentication Standards: Have you implemented DKIM, SPF, and DMARC to ensure your outbound mail is accepted, reducing the number of "bounce" webhooks you have to handle?

Summary of Tradeoffs

Approach | Pros | Cons

Synchronous Processing | Simple to implement, immediate consistency | High risk of timeouts, prone to retry storms

Queue-based Processing | Highly scalable, resilient to spikes | Increased infrastructure complexity, eventual consistency

Simple Logging | Low overhead | No way to recover from missed events without manual logs

Idempotency Table | Guaranteed data integrity | Extra database write per event

Final Thoughts

Reliability in email webhooks is not about preventing failures, but about designing for them. By assuming that the network will fail and that events will be delivered more than once, you build a system that is truly resilient. Whether you are managing SPF records for a small project or scaling a massive transactional system, the patterns of signature verification and idempotency remain the gold standard.

Build your email infrastructure with SendHQ.