Skip to main content

Webhook Retries

When a webhook delivery fails, Von Payments retries it on a fixed schedule.

Retry schedule — 8 attempts over ~79 hours​

AttemptDelay before this attemptCumulative time (approx.)
1immediate0s
230s30s
32 min~2.5 min
410 min~12 min
51 hour~1h 12 min
66 hours~7h 12 min
724 hours~31h
848 hours~79h
(none)dead—

After attempt 8 fails, the delivery is marked dead. The original event row stays in the dashboard's delivery history with status: dead and a final response code; no further attempts are made.

Full-jitter delays​

The "delay before this attempt" column is the base delay; the actual delay is uniform in [0, base] (full jitter), so the cumulative column is an upper bound.

How response codes drive retry​

OutcomeWhat it meansWhat happens next
200–299DeliveredTerminal. No further attempts.
410 GoneEndpoint declared the URL permanently deadSubscription is disabled; row marked dead. No new events are enqueued for it.
404, 408, 425, 429The route is missing (for example mid-deploy), the endpoint timed out, or it is shedding loadRetry with backoff per the schedule above. A Retry-After header can move the next attempt later (up to 1 hour), never earlier.
Other 4xx (400, 401, 403, 409, 422, …)Endpoint rejected the requestRow marked dead immediately. No retry: the same request gets the same answer.
5xx (500, 502, 503, 504)Endpoint had a transient errorRetry with backoff per the schedule above.
Network error / timeout (no response)Connection refused, DNS failed, or your handler took longer than 10 secondsRetry with backoff per the schedule above.

The 10-second timeout is hard: persist the event id, return 200, and do the work in a background job.

Any other 4xx is dead on the first attempt: a rejected signature or a body your handler refuses will not change between attempts. A 401 you return during a signing-secret switch-over is lost the same way; see Rotating a signing secret.

Per-endpoint circuit breaker​

Layered on top of the retry schedule is a per-endpoint circuit breaker that pauses delivery to an endpoint repeatedly returning 5xx.

State machine​

closed ──5 failures within a 60s window──▶ open
▲ │
│ │ cooldown elapsed
│ success ▼
└──────────────────────────── half-open ──5xx──▶ open (longer cooldown)
  • Closed — normal delivery. Every attempt goes through.
  • Open — delivery to this endpoint is paused. No HTTP attempts are made; rows wait. After a cooldown, the breaker moves to half-open.
  • Half-open — one probe attempt goes through. If it succeeds, the breaker closes. If it fails, the breaker re-opens with a longer cooldown.

Cooldown progression​

The cooldown doubles on each re-open until it caps at 5 minutes:

Re-open countCooldown
1 (first open)30s
260s
3120s (2 min)
4240s (4 min)
5+300s (5 min)

A streak of successful deliveries resets the counter; the next open starts over at 30s.

What this looks like in the delivery log​

The circuit breaker is per-subscription, not global. A breaker opening on Subscription A has no effect on Subscription B, even on the same merchant.

Automatic pause after repeated failures​

Separate from the circuit breaker, a subscription is paused automatically after 10 failed delivery attempts in a row, counted across all its events. Retries of one event count: one event's 8 attempts can carry the count most of the way, and a failure on the next event finishes it. A successful delivery before the tenth resets the count.

A paused subscription does not resume by itself. Nothing is sent to it while it is paused, so no delivery can succeed and clear the pause. It clears only when you set the subscription back to active: click Resume on the webhooks page of your dashboard, or send PATCH /v1/webhook_subscriptions/{id} with {"status": "active"}. Fix the endpoint first, or the next 10 failures pause it again.

Its status still reads active while it is paused, and the API does not report the pause. The dashboard does: the endpoint is marked Auto-paused. If your endpoint stops receiving events, check the webhooks page.

Events raised during the pause are held, not dropped. Resume within an event's retry window (about 79 hours from when it was raised) and it is delivered; after that it is dead-lettered like any other undeliverable event. Every delivery, including a held one, goes to the endpoint address the subscription has when it is sent, so you can change the address while paused and resume.

Dead-letter queue​

When a delivery hits a terminal state — all 8 attempts exhausted, a 4xx that is not retried marked it dead, or a 410 Gone disabled the subscription — the row moves to the dead-letter queue (DLQ). DLQ rows are kept for at least 30 days, with the final response code, the error-message excerpt (your response body or the network error) and the full request payload.

Your endpoint's response body is stored

The error-message excerpt is the response body your endpoint returned on the failing attempt, retained for 30 days and visible to anyone with dashboard access to your account. Keep error responses short and free of PII.

Where to see delivery state​

/dashboard/developers/events lists every event with its delivery attempts — each attempt's timestamp, response code, latency and response-body excerpt — and a Resend button for the whole event or one subscription. Dead deliveries are resent from there once you have fixed the handler; the endpoint card at /dashboard/developers/webhooks can bulk-resend an endpoint's dead deliveries. A test event sent from the endpoint card or the CLI (Test your handler) is a single signed POST — it is not retried and never enters the DLQ.

Two invariants to code against​

  • The same id can arrive more than once — a previous attempt delivered but you answered 5xx, or a manual DLQ retry. Guard on id and return 200 for an id you have already processed.
  • Order is not guaranteed. Of two events emitted close together, the second can arrive first if the first is retrying. Where ordering matters, read state from the API by id rather than from the payload alone.