Webhooks you can trust: queue, batch and dead-letter
Keith Pillay · 2 October 2026 · 3 min read

Webhooks look simple: Shopify calls your URL when something changes. In practice, three things make them harder than they appear: you must answer quickly, delivery isn't guaranteed, and changes arrive in bursts.
This is the design I settled on for the Product Health Scanner, which needs to stay current when products are created, edited or deleted, without re-scanning the whole catalogue every time.
The rules Shopify gives you
Per Shopify's documentation:
- Your endpoint must return a success status quickly. Shopify applies tight connection and total timeouts.
- Failed deliveries are retried a limited number of times over several hours.
- After repeated failures, subscriptions can be removed and the app's contact gets a warning.
- And importantly, delivery isn't guaranteed.
The takeaway: do almost nothing inside the webhook handler.
Step 1: the handler only enqueues
The handler does two things. It verifies the request signature, which rejects forged calls, and then it writes one row into a queue table identifying the product and the kind of change.
No fetching, no classifying, no heavy work, so there's nothing in there that can time out.
Two details matter:
- The queue row is keyed by product. If a product is updated three times before I process it, there is still one row. If it's updated and then deleted, the row ends up as a delete. The queue always holds the latest intent, which is all I need.
- Processing doesn't depend on the payload's contents. The webhook says which product changed. When I process it, I fetch the product's current state, so out-of-order or duplicate deliveries can't leave me with stale data.
Step 2: a manager drains the queue in batches
A separate function processes up to 200 pending rows at a time:
- Claim the rows first, so a concurrent run doesn't process the same ones. This is best-effort, not a perfect lock, and that's deliberate. Classification is a pure function of a product's current state, so the worst case from a race is redundant work, never wrong data.
- Batch the deletes into one transaction.
- Fetch all the changed products in one GraphQL call, using the
nodesquery with a list of IDs, instead of one call per product. Results come back in input order, and a missing ID comes back as null. - Classify each product in memory, individually, so one bad product can't take down the batch.
- Write all the results in one transaction.
The performance mistake
My first version opened one database transaction per product. For a 200-item batch it took over a minute, because every round trip to a remote database costs real latency.
Writing the whole batch in a single transaction brought it to about 17 seconds. Round trips, not computation, were the cost. It's an easy thing to miss when everything is fast against a local database.
Step 3: bounded retries, then quarantine
Items fail. A product might be deleted between enqueue and fetch, or the API might hiccup. So each item has a retry count, and after five failures it moves to a separate dead-letter table.
That gives three properties I care about:
- Nothing retries forever.
- Nothing disappears silently.
- A failure is always inspectable later.
The boundary I documented instead of hiding
One check, duplicate SKUs, depends on other products: a product's flag can change because a different product was edited. Doing that incrementally and correctly needs a persisted index with cascading invalidation, which is real complexity for one check. So that check stays full-scan-only, and the incremental path is written so it preserves an existing flag rather than clearing it.
I'd rather ship a clear, documented limit than a subtly wrong shortcut. A full scan is always the source of truth.
When it runs
Today the manager runs on page loads, as a cheap no-op when the queue is empty. A scheduled job can call the same function on a fixed interval, so data stays fresh even when nobody has the app open.
What to take from this
- Acknowledge fast, work later.
- Treat webhooks as hints, not data. Fetch current state.
- Batch your database writes.
- Give failures somewhere to go.
- Write down what you chose not to solve.
Part of the build story.
