Wes Ellis./ a personal notebook
Technology. Stories. Side projects.
A few things worth writing down.
← Back to G8KEPR

G8KEPR

One Gateway, Many Models

Many railway tracks curving together into one junction below overhead wires.

Part 15 of the thread Building G8KEPR

THE SHORT VERSION4 points
  • The AI Gateway puts many LLM providers (the homepage says 14) behind a single front door.
  • It can route by round-robin, lowest latency, lowest cost or failover, with a circuit breaker for each provider.
  • It's bring-your-own-key: your provider key, your account, your billing relationship.
  • The bigger reason for a gateway is security: one checkpoint every request passes through beats each app rolling its own.

A lot of companies started their AI journey with one model from one provider. Then a second team picked a different one. Then someone needed a cheaper model for a high-volume job, and someone else needed one that's better at code. Before long, there are four apps talking to three providers in five slightly different ways.

The AI Gateway is the G8KEPR pillar that sits in the middle of all that. This post is about what it does with traffic, and why I think having one gateway matters more for security than for convenience.

What a gateway is

A gateway is a single point your apps send their model requests to. Instead of each app talking directly to each provider, they all talk to the gateway, and the gateway talks to the providers. G8KEPR's routes many LLM providers behind that one front door. The homepage says 14.

From the app's point of view, it asks for an answer. The gateway decides how to get it.

Four ways to route

"How to get it" is where routing comes in. The gateway supports four strategies:

  • Round-robin. Spread requests evenly across the providers you've set up, one after another. Simple, and it keeps any single provider from carrying all the load.
  • Lowest latency. Send the request wherever it'll come back fastest. Good for anything a person is sitting there waiting on.
  • Lowest cost. Send it wherever it's cheapest. Good for background jobs where nobody's staring at a spinner.
  • Failover. Use your preferred provider, and if it's having a bad day, fall back to another one.

Which one you want depends on the job, and that's sort of the point. The choice lives in the gateway, not hard-coded into every app.

Circuit breakers

Every provider has outages. When one starts failing, the worst thing you can do is keep sending it traffic. Requests pile up, users wait, and each retry adds to the pile.

A circuit breaker borrows the idea from the box in your basement. When a circuit is overloaded, the breaker trips and cuts it off before something catches fire. Later, you reset it and see if things are okay.

The gateway keeps one per provider. When a provider starts failing, its breaker trips and the gateway stops piling traffic onto it, the same basic idea as the one in your basement. Pair that with failover routing and a provider outage can turn into a blip instead of an incident.

Bring your own key

G8KEPR is bring-your-own-key. You use your own provider keys, on your own accounts. Your billing relationship with each provider stays yours, along with whatever terms you've negotiated.

That fits the rest of the design. G8KEPR runs inside your VPC, and it doesn't resell model access or sit between you and your provider bill. It just routes, and checks, the requests you were already making. The reasoning behind the self-hosted setup is in Why It Runs in Your VPC.

The real reason: one checkpoint

Routing is useful. But if I'm honest, the reason the gateway exists in a security product is simpler than any of that. It's the one place every model request passes through.

Think about the alternative. Every app that uses a model has to do its own prompt-injection checking, its own secret filtering, its own spend limits. Some teams do it carefully. Some copy the code from another team and never update it. One team is new, didn't know it was needed, and ships without it. Now your security is only as good as the least careful app.

With a gateway, the checks live in one place:

  • Every request gets the same prompt-injection guard, PII and secret filtering, and cost budget, no matter which app sent it or which model it lands on.
  • When a check improves, it improves for everything at once. No hunting down eleven copies.
  • A brand-new app gets the protection on day one, just by pointing at the gateway.
  • And because everything flows through one place, the signals from it can be lined up with what the other pillars see, which is what the correlation engine is built on.

That's the same logic behind putting a firewall at the edge of a network instead of asking every server to defend itself. Nothing new, just applied to a new kind of traffic.

For the other gateway checks in more detail, see Prompt Injection, Explained Without the Jargon, Keeping Secrets Out of Prompts and A Budget for Your AI Bill.