Skip to content

Call the OpenAI API from the browser without leaking your key

How to call the OpenAI API and other LLM APIs from a web app without exposing your key: a server route, streaming, cost limits and a gateway.

The ProxifyEdge team 6 min read

Adding a chat box or a “summarize this” button is often a few lines of fetch. The catch is the API key. If you call the OpenAI API from the browser directly, your key travels with every request, and anyone who opens DevTools can copy it and spend your budget.

This guide covers why direct browser calls are risky, how to build a small server route that keeps the key safe (including streaming), how to bound costs, and how a gateway can do the same job without a server of your own.

Why you should not call the OpenAI API from the browser directly

LLM providers authenticate with a bearer token, usually Authorization: Bearer <key>. To send that header from the browser, the key has to be in your JavaScript. As we cover in You can’t hide an API key in frontend code, there is no way to keep it there and keep it secret.

The official OpenAI Node SDK makes the risk explicit. It refuses to run in a browser unless you pass an option called dangerouslyAllowBrowser: true. The name is the documentation.

A leaked LLM key is especially costly:

  • Usage is billed per token, so abuse turns directly into charges.
  • Rate limits are shared across everything using the key, so a scraper can make your product unusable.
  • Keys often have broad scope, covering every model and endpoint on the project.

The baseline: a server route that holds the key

The safe default is a route on your own server. The browser sends only what the user typed; the server adds the key, chooses the model and limits, and calls the provider.

// app/api/chat/route.ts in a Next.js App Router project.
export async function POST(request: Request) {
  const { message, stream } = await request.json();
  if (typeof message !== 'string' || message.length === 0 || message.length > 4000) {
    return Response.json({ error: 'Invalid message' }, { status: 400 });
  }

  const upstream = await fetch('https://api.openai.com/v1/chat/completions', {
    method: 'POST',
    headers: {
      Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
      'Content-Type': 'application/json',
    },
    body: JSON.stringify({
      model: process.env.OPENAI_MODEL, // chosen on the server, not by the client
      max_tokens: 500,
      stream: stream === true,
      messages: [{ role: 'user', content: message }],
    }),
  });

  return new Response(upstream.body, {
    status: upstream.status,
    headers: { 'Content-Type': upstream.headers.get('Content-Type') ?? 'application/json' },
  });
}

Notice what the client does not control: the model, the token ceiling and any system prompt. If you let the browser send the full request body, anyone can ask for the most expensive model with the largest output on your key.

Authenticate your own users first

A route like this is a public endpoint unless you protect it. Check your session or token before calling the provider, and keep a per-user count so a single account cannot run unlimited prompts.

Streaming responses to the browser

Users expect tokens to appear as they are generated. Providers support this with server-sent events (SSE): set stream: true and the response arrives as a series of data: lines.

Your route can pass the stream straight through, as the example above already does by returning upstream.body. Two details matter:

  • Use fetch, not EventSource, in the browser. EventSource only supports GET requests and cannot send a JSON body.
  • Make sure nothing buffers the stream. Compression proxies, some serverless platforms and response middleware can hold the whole body until it finishes, which defeats streaming.

Reading the stream on the client looks like this:

const response = await fetch('/api/chat', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({ message, stream: true }),
});

const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();
for (;;) {
  const { value, done } = await reader.read();
  if (done) break;
  // Each chunk holds one or more "data: {...}" lines; parse and render them.
  renderChunk(value);
}

The server route above forwards the client’s stream flag and nothing else from the request body, so streaming stays opt-in while the model and limits stay under your control.

Bounding cost and abuse

Keeping the key secret is half the job. The other half is making sure that even legitimate access cannot run away with your budget.

ControlWhereWhy
Fixed model and max_tokensYour routeStops clients asking for expensive output
Input length limitYour routePrompt tokens are billed too
Per-user quotasYour route or gatewayOne account cannot exhaust everything
Budget or usage limitProvider dashboard, where offeredA hard ceiling if everything else fails
Separate project or key per appProvider dashboardLimits the blast radius of a leak

Rate-limit responses from the provider come back as HTTP 429, usually with a Retry-After header. Handle them gracefully rather than retrying in a tight loop; API rate limiting for frontend developers shows how.

Handling provider errors

LLM calls fail in more ways than most API calls, and they take longer, so plan for errors from the start:

StatusUsual causeWhat to do
401The key is wrong, revoked or missingCheck configuration; never retry automatically
400The request is invalid, for example too many tokensShow a clear message; retrying will fail the same way
429A rate limit or a usage limit was hitWait for Retry-After, then retry with backoff
5xxThe provider is overloaded or failingRetry a small number of times with backoff

Generations can take tens of seconds, so set a generous but finite timeout on your server route, and let users cancel. In the browser, an AbortController stops reading the stream when the user navigates away or presses “Stop”:

const controller = new AbortController();
stopButton.onclick = () => controller.abort();
const response = await fetch('/api/chat', { method: 'POST', body, signal: controller.signal });

Aborting the browser request does not always stop the generation upstream, so keep server-side limits on output length regardless.

Using a gateway instead of writing the route

If you are building a prototype, a static site or a no-code app, writing and hosting a server route may be more work than the feature. A gateway can add the key for you instead.

With ProxifyEdge, you store the provider key in the secrets vault and bind it to the provider’s host, for example api.openai.com. The browser sends a reference, and the real value is inserted on the way out:

const target = 'https://api.openai.com/v1/chat/completions';

const response = await fetch(`https://api.proxifyedge.com/proxy?url=${encodeURIComponent(target)}`, {
  method: 'POST',
  headers: {
    'X-API-Key': 'pk_your_public_key',
    'X-Proxify-Upstream-Authorization': 'Bearer {{secret.OPENAI_KEY}}',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({ model: 'your-chosen-model', max_tokens: 300, messages }),
});

What you get:

  • The provider key never reaches the browser, and the vault refuses to send it to any host other than the one you bound it to.
  • Your public key is locked to your site’s origins, with per-key quotas and request-rate limits (quotas and rate limits).
  • Streamed responses are flushed to the browser as they arrive.

Know what a gateway does not decide for you

With a pure gateway, the request body still comes from the browser, so a determined user can change the model or token count in their own requests, within your quotas. Plan for that:

  • Use a provider key limited to the models and spend you are comfortable exposing, and set a budget at the provider.
  • Keep quotas on the public key tight, and watch usage in the dashboard.
  • For stricter control, require signed URLs, so only your server can mint the requests the browser may make, or keep a server route for the expensive calls.

Other LLM APIs

The same patterns apply to any provider that authenticates with a bearer token, which includes many OpenAI-compatible APIs. Check each provider’s documentation for the exact header it expects, and remember that a gateway can only add credentials in the places it supports.

If you are still seeing CORS errors when calling an AI API from the browser, start with What is CORS? and the no Access-Control-Allow-Origin header fix.

Key takeaways

  • Calling the OpenAI API from the browser directly exposes your key; the SDK’s dangerouslyAllowBrowser flag exists to warn you.
  • A small server route that fixes the model, limits tokens and authenticates users is the safe baseline.
  • Stream with fetch and a ReadableStream, and make sure nothing between you and the provider buffers the response.
  • Bound cost with input limits, per-user quotas and a provider-side budget.
  • A gateway with a secrets vault keeps the key off the client, but the request body is still client-controlled: pair it with tight quotas or signed URLs.

Are you sure?