API rate limiting for frontend developers: 429s and Retry-After
Understand API rate limiting from the browser: what 429 and Retry-After mean, how to read rate-limit headers cross-origin, and how to retry with backoff and jitter.
Every serious API limits how fast you can call it. On the server, API rate limiting is someone else’s problem. In the browser, it becomes yours: a busy page, a retry loop or an eager search box can burn through a quota in seconds, and the error your users see is rarely friendly.
This guide explains the status codes and headers involved, why your JavaScript often cannot see those headers, and how to retry correctly with backoff and jitter.
What API rate limiting looks like on the wire
When you exceed a limit, a well-behaved API responds with HTTP 429 Too Many Requests, defined in RFC 6585. The body usually explains which limit you hit, and the response carries headers that tell you what to do next.
HTTP/2 429
Retry-After: 30
Content-Type: application/json
{"error": "Rate limit exceeded"}
Some APIs use other statuses for related situations, such as 403 for an exhausted quota or 503 when the service itself is overloaded. Read the API’s documentation, but treat 429 as the standard signal.
Limits come in a few common shapes: a fixed number of requests per window (per second, minute or day), a token bucket that allows short bursts above the average rate, and a usage quota per billing period. A single API often combines several.
The headers that matter
Retry-After
Retry-After is defined in RFC 9110. Its value is either a number of seconds to wait or an HTTP date to wait until:
Retry-After: 120
Retry-After: Wed, 07 Oct 2026 14:30:00 GMT
When a server sends it, it is the most reliable answer to “when can I try again?”. Prefer it over any guess of your own.
X-RateLimit-* and RateLimit headers
Many APIs also report your budget on every response, so you can slow down before you hit the wall. The most common style is a family of X-RateLimit-* headers. GitHub, for example, sends:
| Header | Meaning |
|---|---|
x-ratelimit-limit | Requests allowed in the window |
x-ratelimit-remaining | Requests left |
x-ratelimit-used | Requests made |
x-ratelimit-reset | When the window resets, as a Unix timestamp in seconds |
These names are a convention, not a standard, so details vary. Some APIs report the reset as seconds remaining rather than a timestamp. The IETF HTTP APIs working group has a draft specification for standard RateLimit and RateLimit-Policy fields, and some APIs already send them. Check what each API you use actually returns.
Why your code can’t see the headers
Here is the frontend-specific trap. You call an API from the browser, the response clearly includes x-ratelimit-remaining in DevTools, and yet response.headers.get('x-ratelimit-remaining') returns null.
That is CORS at work. For a cross-origin response, the browser only lets script read the CORS-safelisted response headers: Cache-Control, Content-Language, Content-Length, Content-Type, Expires, Last-Modified and Pragma. Every other header is hidden unless the server lists it in Access-Control-Expose-Headers:
Access-Control-Expose-Headers: Retry-After, X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset
If you control the API, add that header. If you don’t, your code can only react to the status code, or you need something in between that exposes the headers. The no Access-Control-Allow-Origin header guide covers the related case where the whole response is blocked.
Retrying correctly
A retry is the right response to a 429 or a transient 503. A bad retry makes things worse: hundreds of clients retrying in lockstep is exactly the traffic spike that caused the limit.
Exponential backoff with jitter
The standard approach is exponential backoff: wait longer after each failure, up to a cap. Add jitter, a random component, so that clients that failed together do not retry together. A simple, effective version is “full jitter”: wait a random time between zero and the current backoff ceiling.
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
function retryDelay(response, attempt) {
const header = response?.headers.get('Retry-After');
if (header) {
const seconds = Number(header);
if (Number.isFinite(seconds)) return seconds * 1000;
const date = Date.parse(header);
if (!Number.isNaN(date)) return Math.max(0, date - Date.now());
}
const ceiling = Math.min(30_000, 500 * 2 ** attempt); // 0.5s, 1s, 2s, ... up to 30s
return Math.random() * ceiling;
}
export async function fetchWithRetry(url, options = {}, { retries = 4 } = {}) {
for (let attempt = 0; ; attempt++) {
let response;
try {
response = await fetch(url, options);
if (response.status !== 429 && response.status !== 503) return response;
} catch (error) {
if (attempt >= retries) throw error; // network failure
}
if (attempt >= retries) return response;
await sleep(retryDelay(response, attempt));
}
}
Only retry what is safe to repeat
Retrying a GET is harmless. Retrying a POST that created an order, when the first attempt actually succeeded and only the response was lost, creates a second order.
- Retry freely for safe and idempotent methods such as
GET,HEAD,PUTandDELETE. - For
POST, retry only if the API supports idempotency keys: a unique value you send with the request so the server can recognize a repeat. Many payment APIs offer this through anIdempotency-Keyheader. - Never retry a
4xxother than429automatically. A400will fail the same way every time.
Use fewer requests in the first place
The best rate-limit handling is not hitting the limit. In browser apps, a few habits remove most of the pressure:
- Debounce input-driven requests. A search box should wait for a pause in typing, not fire on every keystroke.
- Deduplicate in-flight requests. If three components ask for the same resource at once, share one promise.
- Cache responses that do not change often, in memory or with HTTP caching.
- Batch several small calls into one where the API allows it. See Batching API requests.
- Respect
remaining. When a header says you are nearly out, slow down before the429arrives.
Tell users what is happening
Retries should be invisible when they succeed quickly, and honest when they don’t. A few patterns keep rate limits from feeling like bugs:
- Show progress, not a spinner forever. If a retry is scheduled more than a second or two out, say so: “Busy, retrying in 8 seconds”.
Retry-Aftergives you the number. - Disable the button that caused the burst. A user who clicks “Generate” five times in frustration sends five requests.
- Give up gracefully. After the last retry, show a clear message and a way to try again manually, instead of an empty screen.
- Separate quota from bursts. A per-second limit clears in a moment; a monthly quota does not. If the error body tells you which one you hit, show a different message for each.
Rate limits through ProxifyEdge
When you call an API through ProxifyEdge, two sets of limits can apply: your ProxifyEdge plan and key limits, and the upstream API’s own limits. The headers keep them apart (quotas and rate limits):
| Header | From | Meaning |
|---|---|---|
X-Proxify-RateLimit-Limit | ProxifyEdge | Requests allowed in the current window |
X-Proxify-RateLimit-Remaining | ProxifyEdge | Requests left in it |
X-Proxify-RateLimit-Reset | ProxifyEdge | Reset time, as a Unix timestamp in seconds |
X-Quota-Warning | ProxifyEdge | Present once a window passes the warning threshold |
X-RateLimit-* | The upstream API | Passed through untouched |
Retry-After | Either | Seconds to wait, on a 429 |
All of these are exposed to script, so response.headers.get() works from the browser. When ProxifyEdge itself refuses a request, the JSON body names the reason, such as QUOTA_EXCEEDED, RPS_LIMIT_EXCEEDED or RPM_LIMIT_EXCEEDED, so your retry logic can tell a per-second burst from an exhausted monthly quota (error codes). You can see the headers on a live response in the playground.
Key takeaways
429 Too Many Requestsis the standard rate-limit signal;Retry-Aftertells you when to try again, in seconds or as a date.- Rate-limit headers such as
X-RateLimit-Remainingare conventions; check what each API sends. - Cross-origin, your script can only read rate-limit headers the server lists in
Access-Control-Expose-Headers. - Retry with exponential backoff and jitter, honor
Retry-After, and only retry requests that are safe to repeat. - Debounce, deduplicate, cache and batch to stay under limits in the first place.