Guide for platform teams

AI Gateway vs API Gateway: Do You Need Both?

An API gateway governs requests: who may call which service, how often, and where the call goes. An AI gateway governs model traffic, where the token is the unit that matters: it holds provider keys, meters tokens and spend, fails over between providers, and inspects prompts and streamed replies. You need both functions, but not necessarily two gateways.

This guide comes from the team building Tygress, a self-hosted gateway that does both jobs in one binary, so weigh it accordingly. It covers where the two differ and why, four ways to run them, when one gateway is enough, and a checklist for evaluating either.

By the Tygress team · Published · Updated

Side by side

The differences that matter

Both gateways sit between callers and upstreams, and both authenticate, route, limit and log. They part ways when the upstream is a model, because a model call is priced by the token, streamed over minutes, and rewritten in transit.

How an AI gateway differs from an API gateway, most decisive differences first
ConcernAPI gatewayAI gateway
Unit of limitingRequests per second, minute or hourTokens per window (prompt, completion, reasoning), plus requests
Unit of costNot tracked; calls cost about the same to forwardUSD per token, per model, including cached and hidden reasoning tokens
Payload inspectionHeaders and path; bodies buffered, if read at allPrompts, and every chunk of a streamed SSE reply as it passes
Upstream timeoutSeconds; a slow backend is usually a failing oneMinutes; a non-streamed completion sends no headers until it finishes
Failover triggerHealth checks, connect errors, 5xx responsesAlso provider 429s, exhausted key quotas, and targets that cannot serve the request
CachingHTTP semantics: method, URL, Vary headersPrompt similarity: a close-enough prompt reuses a stored reply, at one embedding call per lookup
CredentialsVerifies the caller; the backend trusts the gatewayAlso holds the provider keys; callers get virtual keys bound to them
Signing after mutationRare; bodies usually pass through unchangedNeeded for signed providers such as Bedrock (SigV4): redaction and translation change the body, so sign after them, on every attempt
ProtocolsREST, gRPC, WebSocket, TCP/UDPChat Completions, Responses and Messages, translated between provider dialects
Agent trafficMCP and A2A look like POSTs to one JSON-RPC endpointPer-tool allowlists and limits; human approval for sensitive calls

The columns describe jobs, not product categories: some API gateways add AI plugins, and some AI gateways forward plain HTTP.

Why it differs

Why model traffic breaks API gateway assumptions

Each row in the table comes from a property of model traffic that request-oriented gateways were not designed around.

The cost is in the response, not the request

An API gateway's limiter decides before it forwards: one request, one unit. A model call costs whatever tokens it consumes, most of them generated after the gateway has let it through. Usage arrives with the reply, and for a stream the output count comes only at the end, so token limits are enforced after the fact: the request that crosses the line completes, and the next is refused. Dollars also need a price list: cached and cache-write tokens have their own rates, and hidden reasoning tokens are billed even though the caller never sees them.

Streaming changes what inspection means

Most chat clients stream: the reply arrives as server-sent events, a small delta at a time, over seconds or minutes. A filter that needs the whole body must either buffer the stream, so the user waits, or skip it, so streamed replies go unchecked. Scanning in flight has a catch: a key or card number can be split across two events, so the scanner must hold back a short window of text. And bytes that have left the gateway cannot be recalled, so a match in a reply can be masked but not cleanly refused.

The body changes, so the signature has to follow

AI gateways edit request bodies: DLP redacts an email address, a template adds a system message, a bridge rewrites Anthropic Messages as Chat Completions. Providers that authenticate with a request signature, such as AWS Bedrock with SigV4, sign a hash of the exact body. Sign before the edit and either the signature no longer matches or the unredacted body is sent. Failover adds a twist: the next target may speak another dialect with other credentials, so translation and signing must run again on every attempt.

Timeouts and retries need model-sized defaults

A generic HTTP backend that takes 30 seconds to send headers is usually broken. A model asked for a long answer without streaming sends none until it finishes, which can take minutes; the official OpenAI and Anthropic SDKs wait up to ten minutes. A gateway that gives up first turns one slow request into several abandoned ones the provider may still be generating, and each retry can be billed again.

Failover is about quotas as much as health

Backends fail by going down; model providers more often fail by saying no. A 429 means a key has hit its rate limit, not that the service is down, so rest that key and try another key or provider. Some providers report remaining quota in response headers, so traffic can move before the 429s start. A failover target may be a different model or API dialect, so the gateway must know which targets can serve the request and, if all fail, say which were tried and why.

Agents call tools through one endpoint

MCP and A2A typically carry JSON-RPC over HTTP. To a path-based API gateway, every MCP tool call is a POST to the same URL, so its rules cannot tell a harmless read from a destructive write. Governing agent traffic means parsing the envelope: allowing tools per caller, limiting per tool, and holding sensitive calls for a person to approve, with a timeout that fails closed.

Architecture

Four ways to run an API gateway and an AI gateway

Most teams land on one of these. None is wrong; each puts the same policies in a different place.

  1. Pattern 01

    Chained: an API gateway in front of an AI gateway

    1. Client
    2. API gateway
    3. AI gateway
    4. Model provider

    Your existing API gateway keeps authentication, request limits and routing, and forwards model routes to a separate AI gateway or LLM proxy holding provider keys, budgets and guardrails.

    Works well when
    You have a mature API gateway, and AI belongs to another team or timeline.
    Trade-offs
    Two policy systems, and two places to look in an incident. The AI layer must trust an identity the first one passes on, and timeouts and retries need tuning in both.
  2. Pattern 02

    An API gateway with AI plugins

    1. Client
    2. API gateway + AI plugins
    3. Model provider

    Model features run as plugins inside the API gateway, next to the identity and request policy already there.

    Works well when
    Your gateway's AI features cover the rows in the table, in the edition you run.
    Trade-offs
    AI features are only as deep as the plugin model allows. Check whether guardrails see streamed replies, whether budgets roll up to teams, and which AI plugins need another license.
    How Tygress compares with Kong →
  3. Pattern 03

    A standalone LLM proxy

    1. Application
    2. LLM proxy
    3. Model providers

    A dedicated proxy fronts model APIs, and in some products MCP servers and agents too, but it is not a general gateway for your REST and gRPC services. Applications point their SDK at it and get one endpoint for many providers, with keys, budgets and fallbacks.

    Works well when
    Model access is the main problem, and provider breadth matters more than consolidation.
    Trade-offs
    Limits live in the proxy's own key system unless you connect it to your identity provider, and your REST and gRPC APIs still need a separate gateway with its own identity and rate-limit policy.
    How Tygress compares with LiteLLM →
  4. Pattern 04

    One gateway for both

    1. Client or agent
    2. One gateway
    3. Services, tools and models

    One data plane carries API, model, tool and agent traffic, with one identity model, one plugin chain, and one set of logs, limits and budgets.

    Works well when
    The same callers use APIs and models, agents mix model and tool calls, or you want fewer components to run.
    Trade-offs
    A larger blast radius, so config validation and staged rollouts matter more. One team owns both, the product must do both jobs well, and replacing an established API gateway is a migration of its own.

Decision guide

When one gateway is enough, and when to keep two

One product need not mean one deployment: the same gateway can run as separate fleets for inbound API traffic and outbound model traffic, scaled independently, with one policy language. The real question is how many policy systems you want to maintain.

One gateway is enough when

  • Agents call models, MCP tools and internal APIs in the same workflow, and one policy should cover all three.
  • The same identities, from the same identity provider, call APIs and models, and budgets should follow the person or team rather than a shared key.
  • You are choosing a gateway for new services anyway, or the current one is due for replacement.
  • You self-host and want fewer components to deploy, patch and monitor.

Separate gateways are the better choice when

  • Your API gateway is mature, widely deployed, and owned by a team whose change process you cannot move.
  • You need a capability today that only a specialized product offers, such as very broad provider coverage or a managed service.
  • Security or procurement policy requires a separate product for model egress, not just a separate deployment.
  • You need generally available, vendor-supported software now. Tygress is pre-release and cannot fill that role yet.

How Tygress does it

Both jobs in one binary

Tygress is a self-hosted gateway in Rust, on its own async transport. REST, gRPC, WebSocket, LLM, MCP and A2A routes all run through the same plugin chain, so one set of identity, limit and logging plugins can cover them all. Planned pricing puts every gateway capability in the $0 tier, so identity and AI governance are not upsells.

Clients call nine provider types through Chat Completions, Responses or Messages: natively where the provider offers that API, bridged otherwise. Bridges refuse features they cannot carry, such as image input, rather than dropping them.

Status: pre-release. The gateway is built and in active pre-1.0 development, but not publicly available. There is no published container image or Helm chart yet, the OpenAI Realtime API returns 501, and benchmarks will be published with their methodology. An open-source core is planned at launch.

01

Identity once, limits per person

OIDC, JWT with remote JWKS, LDAP and Active Directory, mTLS, OAuth2 introspection and JWE are built in, with OPA and ACLs for authorization. Verified JWT claims can key limits and budgets per user.

02

Cost in tokens and dollars

Token limits over fixed, sliding, rolling-anchor or token-bucket windows. USD budgets per consumer and per team, with calendar-day, calendar-month or lifetime caps and threshold alerts, priced from a bundled 80-model catalog or your own per-model rates.

03

Guardrails on every streamed chunk

With reply scanning on, DLP and the prompt guard check each chunk of a streamed SSE reply as it passes, and DLP holds back a configurable window of recent text, so a match split across chunks is still caught when it fits inside that window. Prompts can be redacted, blocked or audited; matches in replies are redacted. DLP uses regex detectors with Luhn and entropy checks.

04

Signed per attempt, after every edit

Translation and request signing, such as SigV4 for Bedrock, run inside each retry and failover attempt, after plugins edit the body, so a signed request covers the exact bytes that attempt sends.

05

Failover that explains itself

Key pools rest a key on a 429 and read quota headers from OpenAI, Azure OpenAI, Anthropic and Groq. If every target fails, the 502 lists each one and why. Model providers get a configurable response-header timeout, 600 seconds by default.

06

Agent calls held at the gateway

Per-consumer MCP tool allowlists, per-tool limits, and human approval for MCP tool calls and A2A methods, denied after 45 seconds by default. REST and OpenAPI upstreams can be served as MCP tools.

Evaluation

Evaluation checklist

Ask these of any gateway, or pair of gateways, before you commit. Each maps to a row in the comparison table, and a vague answer usually means the feature stops at buffered, request-shaped traffic.

  1. 01

    Can token limits and USD budgets key on the identity your API policies use, such as a team or an identity-provider claim?

  2. 02

    Are cached and reasoning tokens counted and priced correctly, and can you override the price list?

  3. 03

    Do guardrails inspect streamed replies chunk by chunk, or buffer or skip them? What about a match split across two events?

  4. 04

    After a plugin redacts or translates a body, is the request signed over the new bytes, on every retry?

  5. 05

    What triggers failover, and does the final error say which targets were tried and why?

  6. 06

    What is the default upstream timeout on model routes, and is it separate from API routes?

  7. 07

    Which client APIs work on which providers, and are features a translation cannot carry refused or silently dropped?

  8. 08

    Can you allowlist MCP tools per caller and hold a call for approval? What happens when nobody answers?

  9. 09

    Can a caller's key be capped or revoked without rotating the provider key behind it?

  10. 10

    Which of these features are in the edition you would run, and which need another license?

FAQ

AI gateway and API gateway questions

Is an LLM gateway the same as an AI gateway?

Mostly; the terms are used loosely. An LLM gateway usually means a proxy in front of model APIs: one endpoint, provider keys, routing, failover and token accounting. An AI gateway often adds guardrails on prompts and replies, and governance of what models act through, such as MCP tool servers and agent-to-agent (A2A) calls. Some vendors use the name for an API gateway with AI plugins. Judge a product by the jobs it covers, not by its label.

Can an API gateway rate-limit by tokens?

Only if it can read the model's reply, because the token count arrives with the response, at the very end of a stream. A token limit therefore admits requests while the caller is under budget and counts each reply afterwards, so one request can overshoot before the next is refused. Estimating from the prompt and max_tokens narrows the gap but rejects some requests that would have fit. Key token counters on the same identity as request limits, or the two policies drift apart.

Do I need an AI gateway if I only use one model provider?

Not always. With one provider and one application, the provider's own keys, dashboards and spend limits may be enough. A gateway pays off when several teams share the provider and you want per-team or per-user budgets instead of a shared key, keys you can revoke without rotating the provider's, prompts checked for secrets before they leave, and usage attributed to whoever made each call.

Can I put an AI gateway behind my existing API gateway?

Yes. That is the chained pattern, and the least disruptive way to add model governance. Decide first which layer owns each policy. The API gateway usually keeps authentication and request limits; the AI gateway needs the caller's identity to key token limits and budgets, so the first layer must pass on a verified identity, such as a signed JWT. Raise timeouts on model routes, do not buffer streamed replies, and do not retry in both layers.

Does a self-hosted AI gateway keep prompts inside my network?

It keeps policy inside your network, not inference. A self-hosted gateway decides where each request goes and what it may contain, and can redact secrets and personal data before a request leaves. A prompt sent to an external model API still leaves your network. Keeping prompts in-house also needs model endpoints you run yourself, and control of the gateway's logs, caches and telemetry, which can hold prompt content.

Early access

Follow Tygress toward release

If one gateway for API and AI traffic fits your platform, join the waitlist and follow the build toward 1.0.

Join the early-access waitlist

Get notified when the beta opens, plus development and launch updates. Tygress is pre-release: the gateway is built but not publicly available yet.

Just your email. No payment required. Unsubscribe anytime.