Rate Limits
Rate limits are sliding-window counters over 60 seconds, counted per IP or per API key (the table says which). Two routes are metered on both axes at once (session create, and the public session read), so a request can be refused by either.
Buckets
| Bucket | Endpoint(s) | Limit |
|---|---|---|
sessionsPerKey | POST /v1/sessions, any key | 30 / 60s per key |
sessionsSecretKeyPerIp | POST /v1/sessions with a secret key (vp_sk_*) | 300 / 60s per IP |
sessions | POST /v1/sessions with a publishable key or no key | 10 / 60s per IP |
paymentOpsPerKey | The payment routes with a secret key: POST /v1/payment_intents (+ /{id} sub-routes: retrieve, capture, void), POST /v1/refunds and GET /v1/refunds, POST /v1/tokens (+ /{id}), /v1/buyers + /v1/buyers/{id}, GET /v1/payment_methods | 300 / 60s per key, shared across these routes |
paymentOpsSecretKeyPerIp | The same routes with a secret key (per-IP axis) | 600 / 60s per IP |
paymentOps | The same routes without a secret key, and POST /v1/public/tokens | 100 / 60s per IP |
cardTokenizePerSession | Saving a card through POST /v1/public/tokens: the embedded card form's save, and the save step of save-then-charge (elements or setup-mode sessions); both share one budget | 6 / 60s per checkout session |
cardTokenizePerIp | The same card saves (per-IP axis) | 20 / 60s per IP |
sessionRead | GET /v1/sessions/:id, GET /v1/capabilities | 30 / 60s per IP |
publicSessionRead | GET /v1/public/sessions/:id, POST /v1/public/binder-load | 1000 / 60s per key |
publicSessionReadPerIp | the same two routes (per-IP axis) | 80 / 60s per IP |
publicWebhookRead | GET /v1/webhook_subscriptions + /{id}, GET /v1/webhook_events/{id} | 100 / 60s per key |
webhook_subscription_writes | POST / PATCH / DELETE /v1/webhook_subscriptions (+ sub-routes) | 30 / 60s per key |
webhookListenSseConnections | GET /v1/webhook_listen | 30 / 60s per key |
sdkTelemetry | POST /v1/sdk-telemetry | 30 / 60s per key |
browserTelemetry | POST /v1/public/telemetry/browser | 30 / 60s per IP |
healthDeep | GET /api/health?deep=true | 5 / 60s per IP |
Shallow GET /api/health (no deep param) is intentionally unmetered — it's what uptime monitors hit.
Create sessions from your server with your secret key. A session create sent with a secret key (vp_sk_*, or a legacy vp_key_*) counts against your key (30/min) and a loose 300/min ceiling for its IP address. It does not count against the strict 10/min per-IP limit, which still applies to calls with a publishable key or no key. This matters on serverless and edge platforms (Cloudflare Workers, for example), where many customers' traffic leaves from one shared address: with a vp_sk_* key you do not share a quota with them.
The per-IP rejection emits rate_limit_exceeded; the per-key rejection emits rate_limit_exceeded_per_key, so SDKs can tell them apart.
The public session read is metered the same way: 1000/min against your publishable key, but only 80/min from a single IP — the number a single browser or a one-machine load test hits first. POST /v1/public/binder-load shares both buckets, so binder-load calls and session reads from one browser draw down the same 80/min allowance.
Response headers
X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset are emitted on both successful responses and 429s.
Retry-After is not exclusive to rate limiting. It also accompanies 503 service_unavailable — during a maintenance write-freeze, where mutating money routes are refused so no new row lands (Retry-After: 30), and on a disabled /v1/webhook_listen. Branch on the status code, not on the presence of this header: a 429 means slow down, a 503 write-freeze means the write did not happen at all.
| Header | When | Description |
|---|---|---|
X-RateLimit-Limit | success + 429 | Maximum requests allowed in the current window |
X-RateLimit-Remaining | success + 429 | Requests remaining in the current window (0 on a 429) |
X-RateLimit-Reset | success + 429 | Unix epoch seconds when the window resets |
Retry-After | 429, and also 503 | Seconds to wait before retrying: on a 429, 1 to 60, the time until a slot frees in the window. A 503 (maintenance write-freeze, disabled webhook stream) sends it too; see the note above |
Handling 429
{
"error": "Too many requests",
"code": "rate_limit_exceeded",
"fix": "Too many requests from this caller — wait the number of seconds in the Retry-After header (1–60), then retry with jittered exponential backoff",
"docs": "https://docs.vonpay.com/reference/api#rate-limits"
}
Wait at least Retry-After seconds (the same number is in selfHeal.actions[].retryAfterSeconds), then retry with jittered exponential backoff. Do not retry in a tight loop.
The error envelope is flat (not nested). The example above shows the fields you'll branch on (error, code, fix, docs); every error response also carries a selfHeal object — see Error Codes for the full envelope.
SDK auto-retry
The Node and Python SDKs automatically retry on 429 and 5xx responses with exponential backoff:
- The SDK reads
Retry-Afterwhen present - Retry delay capped at 60 seconds
- Default
maxRetriesis 2 — configurable via the constructor
On a money-moving call (paymentIntents.create / .capture / .void, refunds.create) an ambiguous response — a 5xx or a timeout — is retried only if you supplied an Idempotency-Key; without one the SDK stops and raises (Node: VonPayError with retryWithheld: true) because a repeat could charge the buyer twice. A 429 is retried either way. Full contract: Node SDK → Auto-retry.
Hitting the per-key ceiling
If you need more than 30 session creates/minute per API key (high-volume batch fulfilment, say), contact support — the limit can be raised.