Skip to content

API rate limiting for frontend developers: 429s and Retry-After

Understand API rate limiting from the browser: what 429 and Retry-After mean, how to read rate-limit headers cross-origin, and how to retry with backoff and jitter.

The ProxifyEdge team 6 min read

Every serious API limits how fast you can call it. On the server, API rate limiting is someone else’s problem. In the browser, it becomes yours: a busy page, a retry loop or an eager search box can burn through a quota in seconds, and the error your users see is rarely friendly.

This guide explains the status codes and headers involved, why your JavaScript often cannot see those headers, and how to retry correctly with backoff and jitter.

What API rate limiting looks like on the wire

When you exceed a limit, a well-behaved API responds with HTTP 429 Too Many Requests, defined in RFC 6585. The body usually explains which limit you hit, and the response carries headers that tell you what to do next.

HTTP/2 429
Retry-After: 30
Content-Type: application/json

{"error": "Rate limit exceeded"}

Some APIs use other statuses for related situations, such as 403 for an exhausted quota or 503 when the service itself is overloaded. Read the API’s documentation, but treat 429 as the standard signal.

Limits come in a few common shapes: a fixed number of requests per window (per second, minute or day), a token bucket that allows short bursts above the average rate, and a usage quota per billing period. A single API often combines several.

The headers that matter

Retry-After

Retry-After is defined in RFC 9110. Its value is either a number of seconds to wait or an HTTP date to wait until:

Retry-After: 120
Retry-After: Wed, 07 Oct 2026 14:30:00 GMT

When a server sends it, it is the most reliable answer to “when can I try again?”. Prefer it over any guess of your own.

X-RateLimit-* and RateLimit headers

Many APIs also report your budget on every response, so you can slow down before you hit the wall. The most common style is a family of X-RateLimit-* headers. GitHub, for example, sends:

HeaderMeaning
x-ratelimit-limitRequests allowed in the window
x-ratelimit-remainingRequests left
x-ratelimit-usedRequests made
x-ratelimit-resetWhen the window resets, as a Unix timestamp in seconds

These names are a convention, not a standard, so details vary. Some APIs report the reset as seconds remaining rather than a timestamp. The IETF HTTP APIs working group has a draft specification for standard RateLimit and RateLimit-Policy fields, and some APIs already send them. Check what each API you use actually returns.

Why your code can’t see the headers

Here is the frontend-specific trap. You call an API from the browser, the response clearly includes x-ratelimit-remaining in DevTools, and yet response.headers.get('x-ratelimit-remaining') returns null.

That is CORS at work. For a cross-origin response, the browser only lets script read the CORS-safelisted response headers: Cache-Control, Content-Language, Content-Length, Content-Type, Expires, Last-Modified and Pragma. Every other header is hidden unless the server lists it in Access-Control-Expose-Headers:

Access-Control-Expose-Headers: Retry-After, X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset

If you control the API, add that header. If you don’t, your code can only react to the status code, or you need something in between that exposes the headers. The no Access-Control-Allow-Origin header guide covers the related case where the whole response is blocked.

Retrying correctly

A retry is the right response to a 429 or a transient 503. A bad retry makes things worse: hundreds of clients retrying in lockstep is exactly the traffic spike that caused the limit.

Exponential backoff with jitter

The standard approach is exponential backoff: wait longer after each failure, up to a cap. Add jitter, a random component, so that clients that failed together do not retry together. A simple, effective version is “full jitter”: wait a random time between zero and the current backoff ceiling.

const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));

function retryDelay(response, attempt) {
  const header = response?.headers.get('Retry-After');
  if (header) {
    const seconds = Number(header);
    if (Number.isFinite(seconds)) return seconds * 1000;
    const date = Date.parse(header);
    if (!Number.isNaN(date)) return Math.max(0, date - Date.now());
  }
  const ceiling = Math.min(30_000, 500 * 2 ** attempt); // 0.5s, 1s, 2s, ... up to 30s
  return Math.random() * ceiling;
}

export async function fetchWithRetry(url, options = {}, { retries = 4 } = {}) {
  for (let attempt = 0; ; attempt++) {
    let response;
    try {
      response = await fetch(url, options);
      if (response.status !== 429 && response.status !== 503) return response;
    } catch (error) {
      if (attempt >= retries) throw error; // network failure
    }
    if (attempt >= retries) return response;
    await sleep(retryDelay(response, attempt));
  }
}

Only retry what is safe to repeat

Retrying a GET is harmless. Retrying a POST that created an order, when the first attempt actually succeeded and only the response was lost, creates a second order.

  • Retry freely for safe and idempotent methods such as GET, HEAD, PUT and DELETE.
  • For POST, retry only if the API supports idempotency keys: a unique value you send with the request so the server can recognize a repeat. Many payment APIs offer this through an Idempotency-Key header.
  • Never retry a 4xx other than 429 automatically. A 400 will fail the same way every time.

Use fewer requests in the first place

The best rate-limit handling is not hitting the limit. In browser apps, a few habits remove most of the pressure:

  • Debounce input-driven requests. A search box should wait for a pause in typing, not fire on every keystroke.
  • Deduplicate in-flight requests. If three components ask for the same resource at once, share one promise.
  • Cache responses that do not change often, in memory or with HTTP caching.
  • Batch several small calls into one where the API allows it. See Batching API requests.
  • Respect remaining. When a header says you are nearly out, slow down before the 429 arrives.

Tell users what is happening

Retries should be invisible when they succeed quickly, and honest when they don’t. A few patterns keep rate limits from feeling like bugs:

  • Show progress, not a spinner forever. If a retry is scheduled more than a second or two out, say so: “Busy, retrying in 8 seconds”. Retry-After gives you the number.
  • Disable the button that caused the burst. A user who clicks “Generate” five times in frustration sends five requests.
  • Give up gracefully. After the last retry, show a clear message and a way to try again manually, instead of an empty screen.
  • Separate quota from bursts. A per-second limit clears in a moment; a monthly quota does not. If the error body tells you which one you hit, show a different message for each.

Rate limits through ProxifyEdge

When you call an API through ProxifyEdge, two sets of limits can apply: your ProxifyEdge plan and key limits, and the upstream API’s own limits. The headers keep them apart (quotas and rate limits):

HeaderFromMeaning
X-Proxify-RateLimit-LimitProxifyEdgeRequests allowed in the current window
X-Proxify-RateLimit-RemainingProxifyEdgeRequests left in it
X-Proxify-RateLimit-ResetProxifyEdgeReset time, as a Unix timestamp in seconds
X-Quota-WarningProxifyEdgePresent once a window passes the warning threshold
X-RateLimit-*The upstream APIPassed through untouched
Retry-AfterEitherSeconds to wait, on a 429

All of these are exposed to script, so response.headers.get() works from the browser. When ProxifyEdge itself refuses a request, the JSON body names the reason, such as QUOTA_EXCEEDED, RPS_LIMIT_EXCEEDED or RPM_LIMIT_EXCEEDED, so your retry logic can tell a per-second burst from an exhausted monthly quota (error codes). You can see the headers on a live response in the playground.

Key takeaways

  • 429 Too Many Requests is the standard rate-limit signal; Retry-After tells you when to try again, in seconds or as a date.
  • Rate-limit headers such as X-RateLimit-Remaining are conventions; check what each API sends.
  • Cross-origin, your script can only read rate-limit headers the server lists in Access-Control-Expose-Headers.
  • Retry with exponential backoff and jitter, honor Retry-After, and only retry requests that are safe to repeat.
  • Debounce, deduplicate, cache and batch to stay under limits in the first place.

Are you sure?