Your Webhook Isn't Reliable Until You Stop Trusting the Gateway
Most backend engineers treat a webhook as a single, trustworthy event. It isn't. Gateways retry on timeout, on 5xx, on their own internal hiccups — and they will send the same event twice, sometimes minutes apart, sometimes days apart. This post walks through why relying on the gateway's event ID alone is not enough, and why idempotency has to be owned and enforced by your own backend, not borrowed from someone else's infrastructure.
Your Webhook Isn't Reliable Until You Stop Trusting the Gateway
Every payment integration guide walks you through the happy path: gateway sends webhook, you verify the signature, you update the database, you return 200. It reads like a solved problem. It is not.
The part that gets glossed over is what happens on the unhappy path — the one that actually shows up in production. Your server takes 4 seconds to respond because the database was under load. The gateway's client times out at 3 seconds and assumes failure. It retries. Now you have two identical webhook deliveries for the same event, and your code has already processed the first one and is about to process the second.
If your subscription activation logic isn't built to expect this, you activate the subscription twice, send two confirmation emails, or worse, credit a wallet twice. None of this is hypothetical. It's the default behavior of every major payment gateway, because retry-on-uncertainty is the correct thing for the gateway to do. The correctness burden shifts entirely to you.
Why "just check the event ID" isn't enough
The instinctive fix is to store the gateway's event ID and skip anything you've already seen. That works until it doesn't, for a few reasons that only show up under real traffic.
First, checking and inserting the event ID as two separate steps is itself a race condition. If two webhook deliveries arrive close enough together, both can pass the "have I seen this?" check before either has written its record. You've re-created the exact bug you were trying to prevent, just one layer down.
Second, gateways don't always guarantee that the event ID is stable across every retry in every failure mode. Treating a third-party identifier as your sole source of truth for a financial state transition means your correctness now depends on a system you don't control and can't fully verify from the outside.
Third, event-ID deduplication only protects against exact duplicate deliveries. It does nothing for the more subtle case: two different events that should logically resolve to the same outcome, arriving out of order, where processing them in the wrong sequence leaves your data in a state that doesn't match reality.
Owning the idempotency key instead of borrowing one
The fix that actually holds up is to generate your own idempotency key on your side, before you ever call out to the gateway, and to make your database enforce uniqueness on it — not your application code, the database constraint itself. When the webhook comes back, you're not asking "have I seen this event ID before," you're asking "does a record with this key already exist," and you let a unique constraint do the deciding under concurrent access, not a conditional check that can race against itself.
This also changes how you think about the webhook itself. Instead of treating it as the trigger for the action, you treat it as confirmation of an action whose intent you already recorded. The subscription attempt, the order, the payment intent — all of that gets written to your database with your own key before the gateway is even called. The webhook's job is just to update the status of something that already exists, not to create something new. That single shift eliminates most of the double-processing bugs I've seen, because there's no window where "does this exist yet" is ambiguous.
Treat the webhook as the source of truth, not your own polling
A related mistake is trying to shortcut the webhook by trusting the response of the initial payment API call to mark something as complete. Payment gateways are explicit about this for a reason: the synchronous response can indicate the request was accepted, not that the money has actually settled. Settlement is asynchronous, and the webhook is the only event that reflects the gateway's actual final state. If you activate anything based on the initial call response instead of waiting for and verifying the webhook, you'll eventually hit a case where the initial call succeeded but the payment later failed or reversed, and now your system disagrees with reality until someone notices manually.
This means the webhook has to be the single writer for state transitions like "subscription is active" or "order is paid." Any other code path that's tempted to flip that same flag is a second writer, and two writers racing against each other is how you end up debugging inconsistent state at 2 a.m. with no clear story of how it happened.
Signature verification is necessary, not sufficient
It's worth saying plainly: verifying the webhook signature is table stakes, not a substitute for idempotency. Signature verification tells you the event genuinely came from the gateway. It says nothing about whether you've already processed this exact event, or whether processing it again is safe. Teams sometimes stop at signature verification because it feels like the security-critical part, and it is — but idempotency is the correctness-critical part, and skipping it doesn't fail loudly. It fails quietly, as a support ticket from a customer who got charged twice, weeks after you shipped.
The actual takeaway
None of this is exotic. It's a handful of disciplined decisions: generate your own key before calling the gateway, enforce uniqueness at the database level rather than in application logic, let the webhook be the only writer for financial state, and never let the synchronous API response stand in for settlement confirmation. Individually each of these sounds obvious. The failure mode is skipping one of them under deadline pressure and assuming the gateway's retry behavior will be gentle. It won't be. It's designed to retry aggressively precisely because the gateway would rather over-deliver than risk you missing a critical event — which means the responsibility for handling that over-delivery correctly is, and always was, yours.