What's new

Tygress changelog: what has been built in the gateway

Tygress is a self-hosted API gateway and AI gateway written in Rust. One binary routes REST, gRPC and WebSocket traffic, LLM requests to nine provider types, MCP tools and A2A agents under shared identity, rate-limit, budget and guardrail policies. It is pre-release: the gateway is built and in active pre-1.0 development, but not publicly available yet. This changelog lists the notable changes that have landed on the gateway's main branch since June 2026, newest first. Notes describe merged work in plain terms, including known limits such as opt-in build features and requests an API bridge refuses.

Updated 40 dated updates since June 2026Latest:

Join the early-access waitlist

Get notified when the beta opens, plus development and launch updates. Tygress is pre-release: the gateway is built but not publicly available yet.

Just your email. No payment required. Unsubscribe anytime.

October 2026

1 update

    • AI gateway

    Failover errors that explain themselves

    When every target on a failover route fails, the 502 now lists each target the route passed over and why: the upstream status, a connect failure, or a refusal of the request body. Streamed tool calls stay separate on the Anthropic Messages bridge, and Bedrock and Azure OpenAI keep a base URL path prefix, so they work behind API-management layers and re-signing proxies.

    Related:Enterprise AI gateway

September 2026

7 updates

    • AI gateway
    • Agents & MCP

    Function calling on /v1/responses, on every provider

    Clients using the OpenAI Responses API can complete multi-turn function-calling round trips on any configured provider whose model supports tool calling, buffered or streaming. Providers without a native Responses API are bridged through chat completions, including upstreams that stream tool calls without indices. On a route with a selection strategy, a request one target cannot accept moves to a target that can.

    Related:Enterprise AI gateway MCP and agent governance

    • Security
    • AI gateway

    Prompt guard catches look-alike-letter injections

    The prompt guard's built-in instruction-override, role-override, and system-prompt-probe rules also scan a Unicode-normalized copy of each prompt, with invisible characters removed and look-alike letters, mostly Greek and Cyrillic, folded to ASCII, so “ignore all previous instructions” written with swapped letters is still caught. Custom patterns still match the raw text.

    Related:Enterprise AI gateway security Enterprise AI gateway

    • AI gateway

    Long non-streaming completions no longer cut off at 30 seconds

    Each AI provider gets its own transport with a configurable response-header timeout. The 600-second default matches the wait the official OpenAI and Anthropic SDKs use, and transport timeouts now return 504 instead of a misleading 502. A route's total timeout still caps how long a request may take.

    Related:Enterprise AI gateway

    • Security
    • Operations

    Pre-release hardening pass

    An end-to-end smoke pass fixed defects in boot, drain, the zero-downtime upgrade handoff and the response path, and closed the four security defects it found. Admin API reads now redact stored keys and secrets; a key minted through the key-issuance endpoint is shown once, when it is issued. On SIGTERM, /healthz now reports draining while the listeners keep accepting, so load balancers deregister the node before connections close.

    Related:Enterprise AI gateway security Self-hosted deployment

    • Engine

    Transport keep-alive repair and worker-thread control

    The transport now resends an idempotent request without a body once on a fresh connection when a pooled HTTP/1 keep-alive connection the peer had already closed fails before any response byte, closing a rare spurious-502 race. A single-worker acceptance run afterwards served 694,128 requests without an error. Operators can also pin the HTTP runtime's worker-thread count.

    Related:Why a Rust API gateway

    • Engine

    Tygress now runs on its own transport and request engine

    The vendored fork of Cloudflare's Pingora is gone. Tygress now runs on tygress-proxy, an async transport written for this gateway on hyper and rustls, and a new request engine, and all 64 built-in plugins were ported. Provider translation and request signing now run inside each retry and failover attempt, after every plugin edit; DLP and prompt guards see every chunk of a streamed reply; the zero-downtime upgrade handoff runs over a Unix socket on macOS as well as Linux; and L4, gRPC-Web, WebSocket and CONNECT traffic share one transport. Dynamically loaded native plugins were retired; sandboxed proxy-wasm plugins are the extension path.

    Related:Why a Rust API gateway

    • Operations
    • Security

    Release readiness: CI, dependency audit and safer defaults

    CI covering lint, tests, the minimum Rust version, release builds and WebAssembly guests; a daily dependency-advisory audit; a security policy for reporting vulnerabilities; generated third-party notices in the container image; and safer deployment defaults, such as a Compose stack that refuses to start without an admin token and keeps NATS and the admin API on loopback.

    Related:Self-hosted deployment Enterprise AI gateway security

August 2026

11 updates

    • API gateway
    • Engine

    proxy-wasm plugins that keep per-request state

    proxy-wasm plugins get one sandboxed guest per request, so state carries from the request hook to the response hook as the SDKs assume, and header iteration works. Plugin configuration is checked against the plugin's declared schema at load, and changed modules are recompiled on the next admin reload, without a restart. HTTP callouts, shared data, queues and timers are not implemented yet.

    Related:Enterprise API gateway Why a Rust API gateway

    • AI gateway
    • Agents & MCP

    Tool calling on /v1/messages across providers

    The Anthropic Messages bridge translates tools, tool_choice, tool_use and tool_result in both directions, including streaming, so tool-using clients written against the Anthropic format can run on non-Anthropic providers. Image and document blocks, non-text tool results, extended thinking and vendor-executed tools still need a native Anthropic endpoint.

    Related:Enterprise AI gateway MCP and agent governance

    • AI gateway

    Anthropic Messages API (/v1/messages) on any provider

    Clients that speak the Anthropic Messages format can reach every configured provider: natively on Anthropic, and bridged through chat completions everywhere else, including streaming; tool calling followed on 20 August. DLP and the prompt guard run on bridged traffic too. Image and document blocks, extended thinking and vendor-executed tools need a native Anthropic endpoint, and the bridge refuses them by name rather than dropping them.

    Related:Enterprise AI gateway

    • AI gateway

    Images, audio, moderations, files and batches

    AI routes can carry OpenAI's images, audio, moderations, files and batches endpoints to providers of the OpenAI type, including multipart uploads; Azure OpenAI and OpenAI-compatible endpoints are not included. These calls are not cost-accounted yet, so spend budgets do not apply to them.

    Related:Enterprise AI gateway

    • API gateway
    • Operations

    Kubernetes EndpointSlice service discovery

    An upstream can follow a Kubernetes Service's ready endpoints by watching its EndpointSlices, so pod churn during a rolling deploy lands in well under a second instead of on the next DNS refresh. If the API server becomes unreachable, the last known endpoints keep serving. Available as an opt-in build feature.

    Related:Validate, deploy and operate the gateway

    • API gateway

    Expressive route predicates

    Route matches can use exact, regex, prefix and exists operators on headers, query parameters and cookies, and any condition can be negated.

    Related:Enterprise API gateway

    • Security
    • AI gateway

    Per-user limits for IdP-authenticated callers

    Verified JWT claims or forward-auth headers can key rate limits, token limits and budgets, so each user your identity provider authenticates gets their own bucket instead of sharing one.

    Related:Enterprise AI gateway security Enterprise AI gateway

    • AI gateway

    Google Vertex AI provider

    Vertex AI joins the provider list through its OpenAI-compatible endpoint for Gemini models, authenticating with a GCP service account and refreshing tokens in the background; a rotated key is picked up on hot reload. Embeddings are not available on Vertex. DLP and the prompt guard now also scan Responses API replies, including streamed ones.

    Related:Enterprise AI gateway

    • AI gateway

    OpenAI Responses API (/v1/responses) on every provider

    /v1/responses is a first-class operation: native on OpenAI and Azure OpenAI, and bridged through chat completions, including streaming, on every other provider. DLP, the prompt guard and budgets engage on it; scanning of streamed replies followed on 5 August. The bridge refuses what it cannot carry, such as stored conversations, hosted tools and multimodal input, with a 400 that names the field.

    Related:Enterprise AI gateway

    • Security
    • Operations

    Single sign-on for the operator console

    The operator console supports OIDC sign-in (authorization code with PKCE) against any compliant identity provider, mapping an IdP claim to the admin or read-only role. Users who match no mapping are denied; admin tokens remain as a break-glass path.

    Related:Enterprise AI gateway security Self-hosted deployment

    • Observability

    OpenTelemetry GenAI semantic conventions

    With the OpenTelemetry plugin attached, AI requests emit a child span that follows the OpenTelemetry gen_ai.* conventions: provider, requested and resolved model, token counts, finish reasons, time to first chunk and cost, so LLM calls appear as LLM calls in GenAI-aware tracing tools.

    Related:LLM telemetry and observability

July 2026

7 updates

    • Security

    Internal security review: 27 verified vulnerabilities fixed

    An internal two-round review across the proxy, authentication, admin API, AI and MCP layers. Every finding was verified against the code before it was fixed, then re-reviewed after the fix.

    Related:Enterprise AI gateway security

    • Observability

    Langfuse export and a canonical LLM record

    Each AI request produces one record with timing including time to first token, the provider and model failover trail, and, for metered calls, token counts and realized cost, plus optional DLP-scrubbed prompt content and, for non-streamed replies, completion content. A Langfuse plugin exports it as traces and generations.

    Related:LLM telemetry and observability

    • AI gateway

    Teams with shared budgets, and governed keys in one call

    Consumers can belong to a team with a shared spend budget, enforced alongside each member's own, plus per-team usage rollups. One admin call mints a key with its spend budget, request and token limits, allowed models and expiry.

    Related:Enterprise AI gateway

    • AI gateway

    Daily, monthly and lifetime spend caps with alerts

    Cost budgets can reset each calendar day or month, or run for a lifetime. Spend survives restarts with PostgreSQL and is shared across replicas with Redis, and threshold crossings fire Slack-compatible webhook alerts.

    Related:Enterprise AI gateway

    • AI gateway

    Built-in model pricing and context-window catalog

    A compiled-in catalog prices mainstream models per token, so budgets, usage analytics and cost-aware routing work without hand-written pricing. Operator pricing always takes precedence.

    Related:Enterprise AI gateway

    • Operations

    Validate configuration before you apply it

    tygress validate (since 1 July) checks a config offline, and the admin API's dry-run mode runs the same pipeline a real apply uses, so CI and GitOps pipelines can check a change before it goes live.

    Related:Validate, deploy and operate the gateway

June 2026

14 updates

    • Operations

    Operator console

    An embedded web console on the admin port, with admin and read-only roles, covering routes, services, consumers, upstreams, AI providers, A2A agents and TLS; live upstream health and MCP toolset pages followed in early July. Available as an opt-in build feature.

    Related:Self-hosted deployment

    • Operations
    • Agents & MCP

    Multi-replica state and Bedrock tool calling

    A2A task affinity, provider quota state, human-approval queues and the autopilot leader lease can live in Redis, so several gateway replicas behave as one. OpenAI-style tool calling is translated end to end for Bedrock Converse, including streaming, and a route can inject an MCP toolset's tools into LLM requests.

    Related:Self-hosted deployment MCP and agent governance

    • Security
    • Operations

    Automatic TLS certificates via ACME

    In-process ACME, using Let's Encrypt by default, issues, renews and hot-swaps certificates over HTTP-01 without a restart. DNS-01, for wildcard certificates, followed on 28 June; it publishes the TXT record through an operator-supplied command or webhook rather than built-in DNS-provider integrations.

    Related:Validate, deploy and operate the gateway

    • Security
    • Agents & MCP

    Hybrid post-quantum TLS and an OAuth front door for MCP

    Hybrid X25519MLKEM768 key exchange on listeners and upstreams with observe, prefer or require policies, and metrics on the group that inbound connections negotiate. A crypto-posture report followed on 19 June, and a pqc-guard plugin that requires post-quantum handshakes on chosen routes on 27 June. MCP servers gain the 2025 MCP authorization pattern: protected-resource metadata, token validation and optional dynamic client registration.

    Related:Post-quantum TLS

    • Operations
    • Security

    Dockerfile, Compose stack, autopilot and incident response

    A multi-stage Dockerfile on a distroless, non-root runtime, and a Compose topology with NATS, a control plane and two data planes. Opt-in autopilot tunes weighted provider routing and semantic-cache settings within operator limits, with automatic rollback or human approval; it runs in the single-process role, not the split control-plane/data-plane topology. An incident-response plugin bans abusive clients or opens circuits on upstream outages.

    Related:Self-hosted deployment Enterprise AI gateway security

    • Security

    OPA, LDAP, JWE, OAuth2 introspection and mTLS identity

    Between 15 and 17 June: OPA external authorization, LDAP with simple or search-then-bind, JWE decryption, RFC 7662 token introspection, mTLS client-certificate-to-consumer mapping, and JWT validation against remote JWKS with required scopes.

    Related:Identity and authorization plugins

    • AI gateway

    Virtual keys

    A consumer's gateway key can be bound to a provider key the operator holds, so consumers never see real provider credentials and upstream capacity is isolated and attributed per consumer.

    Related:Enterprise AI gateway

    • API gateway

    gRPC-JSON transcoding and first-class gRPC routing

    REST and JSON clients can call native gRPC backends through a protobuf descriptor loaded at runtime, with no code generation. The same week added gRPC service and method route matching, gRPC health-check probes, and streaming gRPC-Web responses.

    Related:Enterprise API gateway

    • Agents & MCP

    The gateway as an MCP server

    Serve REST upstreams as MCP tools. From 14 June, a toolset can also merge several upstream MCP servers into one virtual MCP endpoint with namespaced tools and per-consumer tool filtering; aggregation supports stateless JSON-mode Streamable HTTP servers, not SSE sessions, resources or prompts.

    Related:MCP and agent governance

    • Observability

    Usage and spend analytics API

    Admin usage endpoints report requests, tokens and realized cost per consumer, provider and model without Prometheus. With PostgreSQL, the data survives restarts and aggregates across data planes.

    Related:LLM telemetry and observability

    • Security
    • AI gateway

    Prompt-injection and jailbreak guardrail

    The prompt-guard plugin blocks or audits prompts using built-in rules for instruction overrides, role-play jailbreaks, system-prompt probes and encoded payloads, plus custom patterns and an optional external classifier, which lets requests through if it is unavailable unless set to block. It also scans responses, including streamed ones, for system-prompt leakage.

    Related:Enterprise AI gateway security Enterprise AI gateway

    • AI gateway
    • Operations

    Quota-aware provider key pools, graceful drain and binary upgrades

    A provider can hold a pool of API keys: the gateway rotates through them, rests a key after a 429 or near its quota where the provider reports rate-limit headers, and fails over to the next provider when the pool is exhausted. Graceful shutdown with connection draining and a zero-downtime binary upgrade (--upgrade) also arrived. Both were rebuilt on Tygress's own transport in September, and the drain window that lets load balancers deregister a node before it stops accepting has worked as documented since 18 September.

    Related:Provider routing and failover Validate, deploy and operate the gateway

    • AI gateway

    Cost-aware semantic model routing

    A semantic strategy classifies each prompt into a complexity tier and can send it to that tier's cheaper model. Shadow mode, the default, serves the requested model and records the choice it would have made and the projected saving.

    Related:Enterprise AI gateway

    • AI gateway

    Azure OpenAI, Mistral, Groq and AWS Bedrock

    Four more providers on a shared backbone, including Azure deployment paths and Bedrock with SigV4 signing, the Converse API and AWS event-stream responses (Converse followed a day later).

    Related:Enterprise AI gateway