AI Gateway, API Gateway & Agent Routing
One gateway for your AI providers, your own APIs, and your AI agents - with guardrails, caching, analytics and automatic failover built in.
A production system that calls an AI provider directly carries a single point of failure it can rarely do anything about - a rate limit, a regional outage, a slow response - and no shared place to put the controls every team eventually needs: redaction, spending limits, caching, an audit trail. Airelay is a gateway that sits in front of your AI providers, your own APIs, and your AI agents, so routing and policy stop being application code.
Connect your own keys for any of 45 providers and reach 2,376 models through one OpenAI-compatible endpoint. Point your existing SDK at one base URL, and name a model - or name a routing policy instead of a model, and let the gateway choose.
Seven ways to route
A routing policy is configured in the dashboard and referenced as model: "policy/<name>":
- Fallback chains - an ordered list of models, each with its own retry count, tried until one succeeds.
- Weighted load balancing - split traffic across models or accounts by weight.
- Lowest latency, lowest cost, lowest usage - chosen from what recent traffic actually did.
- Priority tiers - exhaust your preferred group of models before falling to the next.
- Semantic routing - the gateway reads the meaning of the prompt and picks the model you described for that kind of work.
A per-key circuit breaker skips a provider that has failed repeatedly for a cooldown window, instead of every request rediscovering the same outage the slow way. A structurally empty response counts as a failure, and your own application can report a bad answer through the feedback API, which feeds the same breaker.
Policies that run inside the gateway
Fourteen AI policies apply to every matching request, with no change to your application:
- PII redaction, including UK identifiers - NI number, NHS number, postcode, sort code, phone - with the real values optionally restored in the reply.
- Semantic caching - a question close enough to one already answered is served from cache instead of being paid for again.
- Guardrails from AWS Bedrock, Azure AI Content Safety, Google Model Armor, Lakera, or your own endpoint, plus prompt and response guards that work by meaning rather than keywords.
- Retrieval from documents you upload, injected into prompts automatically.
- Token and cost budgets per organisation, key, route or model, and prompt compression to cut what each call costs.
- Response-quality scoring on a sample of traffic, so quality regressions surface on their own.
Your own APIs, through the same gateway
Publish your services at gw.noviqent.co.uk/your-org/... and configure what sits in front of them: API keys, JWT, OAuth2, basic and HMAC authentication, ACLs and consumer groups, rate limiting, request and response transformation, caching, CORS, IP restriction, logging to your own collector, load-balanced upstreams with health checks, certificates and vault-backed secrets - 48 policies in total, each configured rather than built.
An audited door for AI agents
Connect an MCP server or an agent-to-agent tool from a directory of 15,063 MCP servers and 418 agents, or add your own by URL. Your team reaches it through one authenticated endpoint, with the upstream credential held by the gateway rather than copied onto every developer machine.
Billed the way it is actually used
Every connected provider key is billed directly by that provider at their own published rate - Airelay never holds a balance on your behalf. Each request is metered and a flat 4% is added on top, visible in the dashboard as it accrues. A curated set of models with a genuine free provider tier works with no key connected at all, capped per organisation per day, so an account can try the platform before connecting anything.
Keys stay where you choose
A connected key is encrypted at rest, or held in your organisation's own Noviqent Vault space and resolved only at the moment a request needs it. Keys are never stored inside the gateway engine itself - they are attached to each request as it passes through. Analytics, a searchable request log, an audit trail of every change, team roles (owner, admin, developer, viewer), optional two-factor authentication and a developer portal for your API documentation come with it. Airelay runs on Noviqent's UK-hosted infrastructure.
How it works
- 45 AI providers and 2,376 models through one OpenAI-compatible endpoint, including your own self-hosted models
- Seven routing strategies - fallback, weighted load balancing, lowest-latency, lowest-cost, lowest-usage, priority tiers, and semantic routing by prompt meaning
- A per-key circuit breaker with exponential backoff, skipping a repeatedly-failing provider proactively
- PII redaction with UK identifiers - NI number, NHS number, postcode, sort code, phone - optionally restored in the reply
- Semantic caching, so a repeated question is answered without paying for it twice
- Guardrails from AWS Bedrock, Azure AI Content Safety, Google Model Armor, Lakera, or your own endpoint
- Prompt and response guards that work by meaning, plus prompt templates and request/response transformation
- Retrieval from your own uploaded documents, injected into prompts automatically
- Token and cost budgets per organisation, key, route or model, with prompt compression
- Automatic response-quality scoring on a sample of traffic
- A full API gateway for your own services - authentication, rate limiting, caching, transformations, upstreams with health checks, certificates and vault-backed secrets (48 policies in total)
- An agent gateway for MCP and agent-to-agent tools, with 15,063 MCP servers and 418 agents to connect from
- Analytics, a searchable request log, an audit trail, team roles and a developer portal
- Bring your own provider keys - encrypted at rest or held in your own Noviqent Vault space, never stored in the gateway engine
- A flat 4% on your own provider cost, metered per request, plus a curated free tier that needs no key at all
Want to see it against your own data?
Talk to us about ai gateway, api gateway & agent routing, or get started directly on Airelay.