G8KEPR
Checking the Answer, Not Just the Question

Part 12 of the thread Building G8KEPR
- The Verification Engine checks every AI-generated answer after it's produced, before your app trusts it.
- It runs four kinds of check: integrity, grounding, structure and constraints.
- Grounding asks a simple question: is this answer actually backed by the material the model was given?
- It also watches for trouble that builds up over a whole conversation, like poisoned memory or a system prompt being subverted.
Most of the conversation about AI security is about the way in. Bad prompts, poisoned tools, leaked secrets. All real, and G8KEPR handles those in other pillars. But there's a second half that gets a lot less attention: what comes back out.
Your app asks a model a question, the model writes an answer, and very often the app just takes that answer and runs with it. Shows it to a customer. Saves it to a record. Hands it to the next step in an automated workflow. That's a lot of trust to place in something nobody checked.
The Verification Engine is the pillar that checks. It looks at every AI-generated answer after it's produced, and it sorts what it's looking for into four kinds of check.
The four kinds of check
I'll describe these as the questions each one asks, in plain terms, rather than walking through how they work inside.
- Integrity. Is the answer sound as a piece of output? Is it what it appears to be, without anything in it that doesn't belong?
- Grounding. Is what the answer says actually supported by the evidence the model was given? More on this one below, because it's the heart of it.
- Structure. Is the answer in the shape your app expects? If the next step needs a particular format, an answer that wanders off that format can break things or be misread downstream.
- Constraints. Does the answer stay inside the rules your app set? If you told the model to stay on one topic or never do a certain thing, did it listen?
None of these is exotic on its own. The point is having all four run on every answer, automatically, instead of hoping someone notices.
Grounding, with a lease
Grounding is the easiest to explain with something that has nothing to do with computers.
Say you're about to rent an apartment, and you ask a friend to read the twelve-page lease and tell you what's in it. They come back and say three things. The rent goes up after a year. You're responsible for snow removal. And pets are allowed.
You flip through the lease. The rent increase is on page four. The snow removal clause is on page nine. The lease doesn't say a single word about pets.
Your friend wasn't lying on purpose. Most leases they've seen probably mention pets, and they filled the gap with what seemed likely. But that third claim isn't grounded. It's not in the thing you asked them to read. And it's the one that'd bite you, because you'd show up with a dog.
That's exactly what language models do, and they do it confidently. When an app hands a model some documents and asks a question, the answer ought to be backed by those documents. G8KEPR's homepage puts it as checking "the model's answer against the evidence it was actually given," and flagging output that's ungrounded (a claim the evidence doesn't support) or fabricated (a detail invented outright).
Why is this security and not just quality? Because an app that acts on made-up facts is an app that can be steered. If someone can nudge a model into asserting something false, and nothing checks, that false thing flows straight into whatever happens next.
Problems that build up over a conversation
Some trouble doesn't show up in any single answer. It accumulates.
Two examples the Verification Engine watches for. The first is poisoned memory: models that remember things across turns or sessions can be fed a false "fact" early on that shapes everything they say later. The second is a system prompt being subverted: the instructions your app gave the model get worn down or overridden bit by bit, so that by turn twenty it's behaving in a way you never intended, even though no single message looked alarming.
Catching these means looking at the conversation as a whole, not just the latest reply. There's more on the memory side of this in Agents That Turn on You.
Why this pillar matters to the rest
The Verification Engine is also one of the four sources the correlation engine draws on. An ungrounded answer might be a harmless slip. An ungrounded answer on the same request where the prompt guard saw something slightly off, and a tool behaved strangely? That's a pattern, and the output side is often where it finally becomes visible.
The detection behind all of this is regex plus classical machine learning rather than a second model grading the first, for reasons I went into in No LLM-as-Judge. And like everything else in G8KEPR, it runs inside your own VPC.
I think "check the answer" is going to feel obvious in a few years. Right now it mostly isn't happening.