Wes Ellis./ a personal notebook
Technology. Stories. Side projects.
A few things worth writing down.
← Back to G8KEPR

G8KEPR

Four Checkpoints on One Request: The G8KEPR Pillars

Part 2 of the thread Building G8KEPR

THE SHORT VERSION4 points
  • A request in an AI app passes four checkpoints: the API, the MCP tools, the model provider, and the model's answer.
  • G8KEPR puts a check at each one, all in the same request path.
  • Each pillar catches a different kind of trouble, from ordinary web attacks to made-up answers.
  • The real point is that one system sees all four, so it can connect signals no single check would notice.

In the overview I said G8KEPR watches four places at once. This post walks through them one at a time, in the order a request usually meets them. I'm sticking to what the product actually claims on g8kepr.com, plus plain explanations of the terms.

Picture one request. A user types something into your AI app. Your API receives it, hands it to a model, the model calls a tool or two, and an answer comes back. Four checkpoints.

1. API Security

This is the front door, and it's the most familiar ground.

The claim: "A WAF with nine built-in block rules, four layers of rate limiting, CSRF and an identity check on every route."

  • A WAF (web application firewall) inspects incoming web requests and blocks the ones that match known attack patterns. Nine rules ship built in.
  • Rate limiting caps how fast someone can hit you, which blunts scraping, brute-forcing, and the "let's see what happens if I send ten thousand of these" crowd. There are four layers of it.
  • CSRF protection (cross-site request forgery) stops another website from quietly making a logged-in user's browser send requests to your app.
  • An identity check on every route means no endpoint gets waved through because somebody forgot to protect it.

None of this is exotic. It's here because AI apps are still web apps, and skipping the basics to chase the shiny stuff is how people get burned.

2. MCP Security

MCP, the Model Context Protocol, is how a model gets hooked up to tools: a database lookup, a file reader, a ticketing system, whatever. Each tool describes itself to the model in text, and the model reads that description to decide how to use it.

That's the soft spot. If someone can change what a tool says about itself, they can slip instructions to the model through a channel nobody's watching. That's tool poisoning.

G8KEPR catches it with a "SHA-256 fingerprint of the full tool definition, re-checked on every call." If the definition changes, the fingerprint changes, and that gets noticed. This one gets its own post, because it's the pillar people most often haven't thought about.

3. AI Gateway

The gateway sits between your app and the model providers. It "routes 14 LLM providers with checks inside the request path: a prompt-injection guard, a cost budget, and PII and secret filtering."

  • Prompt injection is text crafted to override the instructions your app gave the model. "Ignore everything above and…" is the cartoon version; real ones are sneakier. The guard looks for it before the prompt goes out.
  • A cost budget keeps a runaway loop or an abusive user from turning into a surprise bill.
  • PII and secret filtering watches for personal information and things like API keys, so they don't end up sent somewhere they shouldn't go.

Because it routes across 14 providers, you get the same checks no matter which model a request lands on.

4. Verification Engine

The last checkpoint is the answer itself. The Verification Engine "checks the model's answer against the evidence it was actually given — flags ungrounded and fabricated output."

Put simply: if your app hands the model three documents and asks a question, the answer should be supported by those three documents. When the model states something the evidence doesn't back up, that's ungrounded. When it invents a detail outright, that's fabricated. Both get flagged.

This is the one that surprises people. We spend a lot of energy on what goes into a model and not much on whether what comes out is true to its sources.

Why one request plane matters

Here's the reason these four live in one product instead of four.

Each checkpoint on its own produces signals that are often ambiguous. A slightly strange prompt. A tool definition that changed. An answer with a claim nothing supports. Any one of those might be innocent.

G8KEPR's cross-pillar correlation engine scores signals that co-occur across all four checkpoints. When several of them light up on the same request, that's worth a lot more attention than any of them alone. You can only do that if one system sits in the path of the whole request, which is the entire bet behind the product.

Note

All of this runs self-hosted in your own VPC, in-path, and alongside an existing API gateway if you have one. The why is in Why It Runs in Your VPC.

The detection behind these checks is regex plus classical machine learning, not another AI model grading the first one. That was a deliberate choice, and I explain the trade-offs in No LLM-as-Judge.