Batching API requests from the browser to cut round trips
When batching API requests helps, how to express dependencies between calls, how to handle partial failures, and the limits that keep a batch from hurting you.
A dashboard that needs a user, their projects, their usage and their billing status can easily fire six or eight requests before it renders anything useful. On a fast connection that is fine. On a phone with 300 ms of latency, each sequential round trip is felt. Batching API requests sends several upstream calls in one HTTP request and returns all the results together, trading many round trips for one.
This guide covers when batching is worth it, how to express calls that depend on each other, how to deal with partial failure, and which limits keep a batch endpoint from becoming a liability.
Why round trips, not bandwidth, are usually the problem
Most API responses a page needs are small. What makes them slow is latency: each request pays for DNS, connection setup (unless a connection is reused), TLS, and at least one network round trip before the first byte arrives. HTTP/2 lets a browser multiplex many requests over one connection, which helps a lot, but two costs remain:
- Sequential dependencies. If request B needs data from request A, the browser cannot start B until A returns. Two dependent calls cost two full round trips no matter how good the transport is.
- Per-request overhead on the other side. Every request through a gateway or proxy is authenticated, rate-limited and logged separately. When those calls are going to the same place anyway, doing that work once is cheaper.
Batching addresses both: the browser makes one trip to a server close to the APIs, and that server fans out to the upstreams in parallel, over connections it already has warm.
When batching helps, and when it does not
Batching is a good fit when:
- A view needs several independent resources before it can render, and they are small.
- Calls are naturally sequential (“fetch the user, then their orders”) and the browser is far from the API.
- The calls target third-party APIs that the browser cannot reach directly because of CORS, so they go through a proxy anyway.
It is a poor fit when:
- One response is large or slow. A batch returns when its slowest member finishes, so one heavy call holds the rest hostage. Fetch it separately and let the fast data render first.
- Responses should be cached independently by the browser or a CDN. A batch response is one opaque blob to an HTTP cache.
- The calls are user actions that should fail or succeed on their own.
A useful rule: batch what the page needs to render its first useful state, and leave everything else to normal requests.
Expressing dependencies between calls
The interesting part of a batch is ordering. Some calls can start immediately; others only make sense after another succeeds. A clean way to express this is to give every item an id and let an item name one it depends_on:
{
"requests": [
{ "id": "user", "url": "https://api.example.com/me" },
{ "id": "projects", "url": "https://api.example.com/projects", "depends_on": "user" },
{ "id": "status", "url": "https://status.example.com/summary.json" }
]
}
Here user and status start together; projects waits for user. Two design decisions matter:
- A failed dependency skips its dependents. If
userreturns401, sendingprojectsanyway wastes quota and may produce a confusing second error. Mark the dependent as skipped instead, so “never attempted” is distinguishable from “attempted and failed”. - Reject cycles up front.
adepends onbandbonawould wait forever. Validate the graph before running anything and return a clear error.
Note that a dependency is about ordering, not data flow. Passing a value from one response into the next request’s URL turns a batch endpoint into a scripting language, with all the injection and debugging problems that brings. If you genuinely need that, it is usually a sign the logic belongs in your own backend.
Handling partial failure
A batch is not a transaction. Some items succeed and some fail, and the response has to say which, item by item. A practical result shape:
{
"results": [
{ "id": "user", "status": 200, "body": { "name": "Ada" }, "duration_ms": 142 },
{ "id": "status", "status": 503, "text": "unavailable", "duration_ms": 88 },
{ "id": "projects", "status": 0, "skipped": true, "error": "skipped: \"user\" did not succeed", "duration_ms": 0 }
],
"summary": { "total": 3, "succeeded": 1, "failed": 1, "skipped": 1, "duration_ms": 150 }
}
A few conventions make this easy to consume:
- Match results by
id, never by position. Items finish in different orders, and an implementation may return them in completion order. - Separate “the upstream answered with an error” from “the upstream could not be reached”. A
503from the API is a status; a DNS failure or timeout is anerror. - Return JSON bodies as JSON and everything else as text, in different fields, so the caller never has to guess.
- Include a summary, so the caller can branch on “everything succeeded” without walking every result.
On the client, treat each item as its own request:
const byId = Object.fromEntries(results.map((result) => [result.id, result]));
if (byId.user?.status === 200) renderUser(byId.user.body);
if (byId.status?.error || (byId.status?.status ?? 0) >= 400) showStatusUnavailable();
Limits that keep a batch endpoint healthy
A batch endpoint multiplies whatever it receives, which makes it an attractive tool for abuse and an easy way to overload an upstream. Sensible limits:
- A maximum number of items per batch, so one request cannot become a thousand.
- A maximum envelope size, because request bodies inside the batch count too.
- A concurrency cap, so a batch of twenty does not open twenty simultaneous connections to one upstream.
- A deadline for the whole batch, after which remaining items are reported as timed out rather than left hanging.
- Per-item accounting. Each item should count against rate limits and quotas exactly as a single request would, or batching becomes a way around them. The rate limiting guide covers how those limits are usually reported.
Batching through ProxifyEdge
ProxifyEdge exposes a batch endpoint at /proxy/batch that applies these rules to proxied requests. You POST a JSON body with your key; each item needs a unique id and a url, and may set method, headers, body, response transforms, and depends_on:
const response = await fetch('https://api.proxifyedge.com/proxy/batch', {
method: 'POST',
headers: { 'X-API-Key': 'pk_your_public_key', 'Content-Type': 'application/json' },
body: JSON.stringify({
requests: [
{ id: 'user', url: 'https://api.github.com/users/octocat' },
{
id: 'repos',
url: 'https://api.github.com/users/octocat/repos',
select: 'name,stargazers_count',
depends_on: 'user',
},
],
}),
});
const { results, summary } = await response.json();
A batch holds up to 20 items in a 1 MiB envelope, runs at most 6 at once, and has 30 seconds in total. Each item counts as one request against your quota; the envelope itself does not. Dependents of a failed item are returned with skipped set. When a key requires signed URLs, each item carries its own signature for its own target. The batch section of the docs has the full reference.
Key takeaways
- Batching saves round trips, which matter most for dependent calls and high-latency connections.
- Batch only what a view needs to render; a batch is as slow as its slowest item.
- Express ordering with
idanddepends_on, skip dependents of failures, and reject cycles. - Report results per item, matched by
id, with errors and statuses kept distinct. - Cap items, size, concurrency and time, and count every item against quotas.