← All writing
Platform & APIs

Webhooks you can trust: queue, batch and dead-letter

Keith Pillay · 2 October 2026 · 3 min read

Webhooks look simple: Shopify calls your URL when something changes. In practice, three things make them harder than they appear: you must answer quickly, delivery isn't guaranteed, and changes arrive in bursts.

This is the design I settled on for the Product Health Scanner, which needs to stay current when products are created, edited or deleted, without re-scanning the whole catalogue every time.

The rules Shopify gives you

Per Shopify's documentation:

  • Your endpoint must return a success status quickly. Shopify applies tight connection and total timeouts.
  • Failed deliveries are retried a limited number of times over several hours.
  • After repeated failures, subscriptions can be removed and the app's contact gets a warning.
  • And importantly, delivery isn't guaranteed.

The takeaway: do almost nothing inside the webhook handler.

Step 1: the handler only enqueues

The handler does two things. It verifies the request signature, which rejects forged calls, and then it writes one row into a queue table identifying the product and the kind of change.

No fetching, no classifying, no heavy work, so there's nothing in there that can time out.

Two details matter:

  • The queue row is keyed by product. If a product is updated three times before I process it, there is still one row. If it's updated and then deleted, the row ends up as a delete. The queue always holds the latest intent, which is all I need.
  • Processing doesn't depend on the payload's contents. The webhook says which product changed. When I process it, I fetch the product's current state, so out-of-order or duplicate deliveries can't leave me with stale data.

Step 2: a manager drains the queue in batches

A separate function processes up to 200 pending rows at a time:

  1. Claim the rows first, so a concurrent run doesn't process the same ones. This is best-effort, not a perfect lock, and that's deliberate. Classification is a pure function of a product's current state, so the worst case from a race is redundant work, never wrong data.
  2. Batch the deletes into one transaction.
  3. Fetch all the changed products in one GraphQL call, using the nodes query with a list of IDs, instead of one call per product. Results come back in input order, and a missing ID comes back as null.
  4. Classify each product in memory, individually, so one bad product can't take down the batch.
  5. Write all the results in one transaction.

The performance mistake

My first version opened one database transaction per product. For a 200-item batch it took over a minute, because every round trip to a remote database costs real latency.

Writing the whole batch in a single transaction brought it to about 17 seconds. Round trips, not computation, were the cost. It's an easy thing to miss when everything is fast against a local database.

Step 3: bounded retries, then quarantine

Items fail. A product might be deleted between enqueue and fetch, or the API might hiccup. So each item has a retry count, and after five failures it moves to a separate dead-letter table.

That gives three properties I care about:

  • Nothing retries forever.
  • Nothing disappears silently.
  • A failure is always inspectable later.

The boundary I documented instead of hiding

One check, duplicate SKUs, depends on other products: a product's flag can change because a different product was edited. Doing that incrementally and correctly needs a persisted index with cascading invalidation, which is real complexity for one check. So that check stays full-scan-only, and the incremental path is written so it preserves an existing flag rather than clearing it.

I'd rather ship a clear, documented limit than a subtly wrong shortcut. A full scan is always the source of truth.

When it runs

Today the manager runs on page loads, as a cheap no-op when the queue is empty. A scheduled job can call the same function on a fixed interval, so data stays fresh even when nobody has the app open.

What to take from this

  1. Acknowledge fast, work later.
  2. Treat webhooks as hints, not data. Fetch current state.
  3. Batch your database writes.
  4. Give failures somewhere to go.
  5. Write down what you chose not to solve.

Part of the build story.

Hiring a senior Shopify developer?

I'm open to remote roles worldwide. Send a message and I'll reply within a day.