Tygress changelog: what has been built in the gateway
Tygress is a self-hosted API gateway and AI gateway written in Rust. One binary routes REST, gRPC and WebSocket traffic, LLM requests to nine provider types, MCP tools and A2A agents under shared identity, rate-limit, budget and guardrail policies. It is pre-release: the gateway is built and in active pre-1.0 development, but not publicly available yet. This changelog lists the notable changes that have landed on the gateway's main branch since June 2026, newest first. Notes describe merged work in plain terms, including known limits such as opt-in build features and requests an API bridge refuses.
Updated 40 dated updates since June 2026Latest:
Join the early-access waitlist
Get notified when the beta opens, plus development and launch updates. Tygress is pre-release: the gateway is built but not publicly available yet.
Just your email. No payment required. Unsubscribe anytime.
October 2026
1 update
AI gateway
Failover errors that explain themselves
When every target on a failover route fails, the 502 now lists each target the route passed over and why: the upstream status, a connect failure, or a refusal of the request body. Streamed tool calls stay separate on the Anthropic Messages bridge, and Bedrock and Azure OpenAI keep a base URL path prefix, so they work behind API-management layers and re-signing proxies.
Function calling on /v1/responses, on every provider
Clients using the OpenAI Responses API can complete multi-turn function-calling round trips on any configured provider whose model supports tool calling, buffered or streaming. Providers without a native Responses API are bridged through chat completions, including upstreams that stream tool calls without indices. On a route with a selection strategy, a request one target cannot accept moves to a target that can.
The prompt guard's built-in instruction-override, role-override, and system-prompt-probe rules also scan a Unicode-normalized copy of each prompt, with invisible characters removed and look-alike letters, mostly Greek and Cyrillic, folded to ASCII, so “ignore all previous instructions” written with swapped letters is still caught. Custom patterns still match the raw text.
Long non-streaming completions no longer cut off at 30 seconds
Each AI provider gets its own transport with a configurable response-header timeout. The 600-second default matches the wait the official OpenAI and Anthropic SDKs use, and transport timeouts now return 504 instead of a misleading 502. A route's total timeout still caps how long a request may take.
An end-to-end smoke pass fixed defects in boot, drain, the zero-downtime upgrade handoff and the response path, and closed the four security defects it found. Admin API reads now redact stored keys and secrets; a key minted through the key-issuance endpoint is shown once, when it is issued. On SIGTERM, /healthz now reports draining while the listeners keep accepting, so load balancers deregister the node before connections close.
Transport keep-alive repair and worker-thread control
The transport now resends an idempotent request without a body once on a fresh connection when a pooled HTTP/1 keep-alive connection the peer had already closed fails before any response byte, closing a rare spurious-502 race. A single-worker acceptance run afterwards served 694,128 requests without an error. Operators can also pin the HTTP runtime's worker-thread count.
Tygress now runs on its own transport and request engine
The vendored fork of Cloudflare's Pingora is gone. Tygress now runs on tygress-proxy, an async transport written for this gateway on hyper and rustls, and a new request engine, and all 64 built-in plugins were ported. Provider translation and request signing now run inside each retry and failover attempt, after every plugin edit; DLP and prompt guards see every chunk of a streamed reply; the zero-downtime upgrade handoff runs over a Unix socket on macOS as well as Linux; and L4, gRPC-Web, WebSocket and CONNECT traffic share one transport. Dynamically loaded native plugins were retired; sandboxed proxy-wasm plugins are the extension path.
Release readiness: CI, dependency audit and safer defaults
CI covering lint, tests, the minimum Rust version, release builds and WebAssembly guests; a daily dependency-advisory audit; a security policy for reporting vulnerabilities; generated third-party notices in the container image; and safer deployment defaults, such as a Compose stack that refuses to start without an admin token and keeps NATS and the admin API on loopback.
proxy-wasm plugins get one sandboxed guest per request, so state carries from the request hook to the response hook as the SDKs assume, and header iteration works. Plugin configuration is checked against the plugin's declared schema at load, and changed modules are recompiled on the next admin reload, without a restart. HTTP callouts, shared data, queues and timers are not implemented yet.
The Anthropic Messages bridge translates tools, tool_choice, tool_use and tool_result in both directions, including streaming, so tool-using clients written against the Anthropic format can run on non-Anthropic providers. Image and document blocks, non-text tool results, extended thinking and vendor-executed tools still need a native Anthropic endpoint.
Anthropic Messages API (/v1/messages) on any provider
Clients that speak the Anthropic Messages format can reach every configured provider: natively on Anthropic, and bridged through chat completions everywhere else, including streaming; tool calling followed on 20 August. DLP and the prompt guard run on bridged traffic too. Image and document blocks, extended thinking and vendor-executed tools need a native Anthropic endpoint, and the bridge refuses them by name rather than dropping them.
AI routes can carry OpenAI's images, audio, moderations, files and batches endpoints to providers of the OpenAI type, including multipart uploads; Azure OpenAI and OpenAI-compatible endpoints are not included. These calls are not cost-accounted yet, so spend budgets do not apply to them.
An upstream can follow a Kubernetes Service's ready endpoints by watching its EndpointSlices, so pod churn during a rolling deploy lands in well under a second instead of on the next DNS refresh. If the API server becomes unreachable, the last known endpoints keep serving. Available as an opt-in build feature.
Verified JWT claims or forward-auth headers can key rate limits, token limits and budgets, so each user your identity provider authenticates gets their own bucket instead of sharing one.
Vertex AI joins the provider list through its OpenAI-compatible endpoint for Gemini models, authenticating with a GCP service account and refreshing tokens in the background; a rotated key is picked up on hot reload. Embeddings are not available on Vertex. DLP and the prompt guard now also scan Responses API replies, including streamed ones.
OpenAI Responses API (/v1/responses) on every provider
/v1/responses is a first-class operation: native on OpenAI and Azure OpenAI, and bridged through chat completions, including streaming, on every other provider. DLP, the prompt guard and budgets engage on it; scanning of streamed replies followed on 5 August. The bridge refuses what it cannot carry, such as stored conversations, hosted tools and multimodal input, with a 400 that names the field.
The operator console supports OIDC sign-in (authorization code with PKCE) against any compliant identity provider, mapping an IdP claim to the admin or read-only role. Users who match no mapping are denied; admin tokens remain as a break-glass path.
With the OpenTelemetry plugin attached, AI requests emit a child span that follows the OpenTelemetry gen_ai.* conventions: provider, requested and resolved model, token counts, finish reasons, time to first chunk and cost, so LLM calls appear as LLM calls in GenAI-aware tracing tools.
An internal two-round review across the proxy, authentication, admin API, AI and MCP layers. Every finding was verified against the code before it was fixed, then re-reviewed after the fix.
Each AI request produces one record with timing including time to first token, the provider and model failover trail, and, for metered calls, token counts and realized cost, plus optional DLP-scrubbed prompt content and, for non-streamed replies, completion content. A Langfuse plugin exports it as traces and generations.
Teams with shared budgets, and governed keys in one call
Consumers can belong to a team with a shared spend budget, enforced alongside each member's own, plus per-team usage rollups. One admin call mints a key with its spend budget, request and token limits, allowed models and expiry.
Daily, monthly and lifetime spend caps with alerts
Cost budgets can reset each calendar day or month, or run for a lifetime. Spend survives restarts with PostgreSQL and is shared across replicas with Redis, and threshold crossings fire Slack-compatible webhook alerts.
A compiled-in catalog prices mainstream models per token, so budgets, usage analytics and cost-aware routing work without hand-written pricing. Operator pricing always takes precedence.
tygress validate (since 1 July) checks a config offline, and the admin API's dry-run mode runs the same pipeline a real apply uses, so CI and GitOps pipelines can check a change before it goes live.
An embedded web console on the admin port, with admin and read-only roles, covering routes, services, consumers, upstreams, AI providers, A2A agents and TLS; live upstream health and MCP toolset pages followed in early July. Available as an opt-in build feature.
A2A task affinity, provider quota state, human-approval queues and the autopilot leader lease can live in Redis, so several gateway replicas behave as one. OpenAI-style tool calling is translated end to end for Bedrock Converse, including streaming, and a route can inject an MCP toolset's tools into LLM requests.
In-process ACME, using Let's Encrypt by default, issues, renews and hot-swaps certificates over HTTP-01 without a restart. DNS-01, for wildcard certificates, followed on 28 June; it publishes the TXT record through an operator-supplied command or webhook rather than built-in DNS-provider integrations.
Hybrid post-quantum TLS and an OAuth front door for MCP
Hybrid X25519MLKEM768 key exchange on listeners and upstreams with observe, prefer or require policies, and metrics on the group that inbound connections negotiate. A crypto-posture report followed on 19 June, and a pqc-guard plugin that requires post-quantum handshakes on chosen routes on 27 June. MCP servers gain the 2025 MCP authorization pattern: protected-resource metadata, token validation and optional dynamic client registration.
Dockerfile, Compose stack, autopilot and incident response
A multi-stage Dockerfile on a distroless, non-root runtime, and a Compose topology with NATS, a control plane and two data planes. Opt-in autopilot tunes weighted provider routing and semantic-cache settings within operator limits, with automatic rollback or human approval; it runs in the single-process role, not the split control-plane/data-plane topology. An incident-response plugin bans abusive clients or opens circuits on upstream outages.
OPA, LDAP, JWE, OAuth2 introspection and mTLS identity
Between 15 and 17 June: OPA external authorization, LDAP with simple or search-then-bind, JWE decryption, RFC 7662 token introspection, mTLS client-certificate-to-consumer mapping, and JWT validation against remote JWKS with required scopes.
A consumer's gateway key can be bound to a provider key the operator holds, so consumers never see real provider credentials and upstream capacity is isolated and attributed per consumer.
gRPC-JSON transcoding and first-class gRPC routing
REST and JSON clients can call native gRPC backends through a protobuf descriptor loaded at runtime, with no code generation. The same week added gRPC service and method route matching, gRPC health-check probes, and streaming gRPC-Web responses.
Serve REST upstreams as MCP tools. From 14 June, a toolset can also merge several upstream MCP servers into one virtual MCP endpoint with namespaced tools and per-consumer tool filtering; aggregation supports stateless JSON-mode Streamable HTTP servers, not SSE sessions, resources or prompts.
Admin usage endpoints report requests, tokens and realized cost per consumer, provider and model without Prometheus. With PostgreSQL, the data survives restarts and aggregates across data planes.
The prompt-guard plugin blocks or audits prompts using built-in rules for instruction overrides, role-play jailbreaks, system-prompt probes and encoded payloads, plus custom patterns and an optional external classifier, which lets requests through if it is unavailable unless set to block. It also scans responses, including streamed ones, for system-prompt leakage.
Quota-aware provider key pools, graceful drain and binary upgrades
A provider can hold a pool of API keys: the gateway rotates through them, rests a key after a 429 or near its quota where the provider reports rate-limit headers, and fails over to the next provider when the pool is exhausted. Graceful shutdown with connection draining and a zero-downtime binary upgrade (--upgrade) also arrived. Both were rebuilt on Tygress's own transport in September, and the drain window that lets load balancers deregister a node before it stops accepting has worked as documented since 18 September.
A semantic strategy classifies each prompt into a complexity tier and can send it to that tier's cheaper model. Shadow mode, the default, serves the requested model and records the choice it would have made and the projected saving.
Four more providers on a shared backbone, including Azure deployment paths and Bedrock with SigV4 signing, the Converse API and AWS event-stream responses (Converse followed a day later).