Call the OpenAI API from the browser without leaking your key
How to call the OpenAI API and other LLM APIs from a web app without exposing your key: a server route, streaming, cost limits and a gateway.
Adding a chat box or a “summarize this” button is often a few lines of fetch. The catch is the API key. If you call the OpenAI API from the browser directly, your key travels with every request, and anyone who opens DevTools can copy it and spend your budget.
This guide covers why direct browser calls are risky, how to build a small server route that keeps the key safe (including streaming), how to bound costs, and how a gateway can do the same job without a server of your own.
Why you should not call the OpenAI API from the browser directly
LLM providers authenticate with a bearer token, usually Authorization: Bearer <key>. To send that header from the browser, the key has to be in your JavaScript. As we cover in You can’t hide an API key in frontend code, there is no way to keep it there and keep it secret.
The official OpenAI Node SDK makes the risk explicit. It refuses to run in a browser unless you pass an option called dangerouslyAllowBrowser: true. The name is the documentation.
A leaked LLM key is especially costly:
- Usage is billed per token, so abuse turns directly into charges.
- Rate limits are shared across everything using the key, so a scraper can make your product unusable.
- Keys often have broad scope, covering every model and endpoint on the project.
The baseline: a server route that holds the key
The safe default is a route on your own server. The browser sends only what the user typed; the server adds the key, chooses the model and limits, and calls the provider.
// app/api/chat/route.ts in a Next.js App Router project.
export async function POST(request: Request) {
const { message, stream } = await request.json();
if (typeof message !== 'string' || message.length === 0 || message.length > 4000) {
return Response.json({ error: 'Invalid message' }, { status: 400 });
}
const upstream = await fetch('https://api.openai.com/v1/chat/completions', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: process.env.OPENAI_MODEL, // chosen on the server, not by the client
max_tokens: 500,
stream: stream === true,
messages: [{ role: 'user', content: message }],
}),
});
return new Response(upstream.body, {
status: upstream.status,
headers: { 'Content-Type': upstream.headers.get('Content-Type') ?? 'application/json' },
});
}
Notice what the client does not control: the model, the token ceiling and any system prompt. If you let the browser send the full request body, anyone can ask for the most expensive model with the largest output on your key.
Authenticate your own users first
A route like this is a public endpoint unless you protect it. Check your session or token before calling the provider, and keep a per-user count so a single account cannot run unlimited prompts.
Streaming responses to the browser
Users expect tokens to appear as they are generated. Providers support this with server-sent events (SSE): set stream: true and the response arrives as a series of data: lines.
Your route can pass the stream straight through, as the example above already does by returning upstream.body. Two details matter:
- Use
fetch, notEventSource, in the browser.EventSourceonly supportsGETrequests and cannot send a JSON body. - Make sure nothing buffers the stream. Compression proxies, some serverless platforms and response middleware can hold the whole body until it finishes, which defeats streaming.
Reading the stream on the client looks like this:
const response = await fetch('/api/chat', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ message, stream: true }),
});
const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();
for (;;) {
const { value, done } = await reader.read();
if (done) break;
// Each chunk holds one or more "data: {...}" lines; parse and render them.
renderChunk(value);
}
The server route above forwards the client’s stream flag and nothing else from the request body, so streaming stays opt-in while the model and limits stay under your control.
Bounding cost and abuse
Keeping the key secret is half the job. The other half is making sure that even legitimate access cannot run away with your budget.
| Control | Where | Why |
|---|---|---|
Fixed model and max_tokens | Your route | Stops clients asking for expensive output |
| Input length limit | Your route | Prompt tokens are billed too |
| Per-user quotas | Your route or gateway | One account cannot exhaust everything |
| Budget or usage limit | Provider dashboard, where offered | A hard ceiling if everything else fails |
| Separate project or key per app | Provider dashboard | Limits the blast radius of a leak |
Rate-limit responses from the provider come back as HTTP 429, usually with a Retry-After header. Handle them gracefully rather than retrying in a tight loop; API rate limiting for frontend developers shows how.
Handling provider errors
LLM calls fail in more ways than most API calls, and they take longer, so plan for errors from the start:
| Status | Usual cause | What to do |
|---|---|---|
401 | The key is wrong, revoked or missing | Check configuration; never retry automatically |
400 | The request is invalid, for example too many tokens | Show a clear message; retrying will fail the same way |
429 | A rate limit or a usage limit was hit | Wait for Retry-After, then retry with backoff |
5xx | The provider is overloaded or failing | Retry a small number of times with backoff |
Generations can take tens of seconds, so set a generous but finite timeout on your server route, and let users cancel. In the browser, an AbortController stops reading the stream when the user navigates away or presses “Stop”:
const controller = new AbortController();
stopButton.onclick = () => controller.abort();
const response = await fetch('/api/chat', { method: 'POST', body, signal: controller.signal });
Aborting the browser request does not always stop the generation upstream, so keep server-side limits on output length regardless.
Using a gateway instead of writing the route
If you are building a prototype, a static site or a no-code app, writing and hosting a server route may be more work than the feature. A gateway can add the key for you instead.
With ProxifyEdge, you store the provider key in the secrets vault and bind it to the provider’s host, for example api.openai.com. The browser sends a reference, and the real value is inserted on the way out:
const target = 'https://api.openai.com/v1/chat/completions';
const response = await fetch(`https://api.proxifyedge.com/proxy?url=${encodeURIComponent(target)}`, {
method: 'POST',
headers: {
'X-API-Key': 'pk_your_public_key',
'X-Proxify-Upstream-Authorization': 'Bearer {{secret.OPENAI_KEY}}',
'Content-Type': 'application/json',
},
body: JSON.stringify({ model: 'your-chosen-model', max_tokens: 300, messages }),
});
What you get:
- The provider key never reaches the browser, and the vault refuses to send it to any host other than the one you bound it to.
- Your public key is locked to your site’s origins, with per-key quotas and request-rate limits (quotas and rate limits).
- Streamed responses are flushed to the browser as they arrive.
Know what a gateway does not decide for you
With a pure gateway, the request body still comes from the browser, so a determined user can change the model or token count in their own requests, within your quotas. Plan for that:
- Use a provider key limited to the models and spend you are comfortable exposing, and set a budget at the provider.
- Keep quotas on the public key tight, and watch usage in the dashboard.
- For stricter control, require signed URLs, so only your server can mint the requests the browser may make, or keep a server route for the expensive calls.
Other LLM APIs
The same patterns apply to any provider that authenticates with a bearer token, which includes many OpenAI-compatible APIs. Check each provider’s documentation for the exact header it expects, and remember that a gateway can only add credentials in the places it supports.
If you are still seeing CORS errors when calling an AI API from the browser, start with What is CORS? and the no Access-Control-Allow-Origin header fix.
Key takeaways
- Calling the OpenAI API from the browser directly exposes your key; the SDK’s
dangerouslyAllowBrowserflag exists to warn you. - A small server route that fixes the model, limits tokens and authenticates users is the safe baseline.
- Stream with
fetchand aReadableStream, and make sure nothing between you and the provider buffers the response. - Bound cost with input limits, per-user quotas and a provider-side budget.
- A gateway with a secrets vault keeps the key off the client, but the request body is still client-controlled: pair it with tight quotas or signed URLs.