G8KEPR
No LLM-as-Judge: Why G8KEPR Uses Regex and Classical ML
Part 4 of the thread Building G8KEPR
- LLM-as-judge means asking a second AI model whether the first one's input or output looks bad. G8KEPR doesn't do that.
- It uses deterministic regex plus classical machine learning instead: a 4.6 ms median detection time, the same answer every time, no per-token bill, nothing sent out.
- The published numbers are 0.788 recall and 0.857 precision on held-out data. Roughly: it catches about four out of five, and most of what it flags is real.
- That means it misses some. I'd rather publish that than pretend otherwise.
When people see what G8KEPR does, a fair question comes up pretty quickly: how does it decide what's an attack? The answer on the site is short. "Deterministic regex plus classical ML, no LLM-as-judge." This post is the long version.
What LLM-as-judge means
An LLM is a large language model, the kind of AI behind chatbots. LLM-as-judge is a popular pattern where you take a prompt, or a model's answer, and hand it to a second model with instructions like "Is this a prompt-injection attempt? Answer yes or no."
It's appealing. Language models are flexible, they understand phrasing and context, and you can get a working prototype in an afternoon. I understand why so many tools go that way.
I didn't, for four reasons.
Reason 1: speed
G8KEPR sits in-path. Every request goes through it before it goes anywhere else. Whatever it spends deciding, your user spends waiting.
The median detection time is 4.6 ms. Regex (pattern matching on text) and classical machine learning models (the smaller, older family of statistical models that came before today's giant ones) are fast. Calling another language model on every request, and sometimes on every step of a request, is a different order of cost in time.
Reason 2: determinism
"Deterministic" means the same input gives the same result, every time.
Language models generally don't promise that. Ask the same question twice and you can get two answers. For a security control, that's awkward. If a request was blocked, you want to be able to explain why, reproduce it, and test that a fix actually fixed it. A judge that might rule differently tomorrow on the identical case makes that hard.
There's a subtler problem too. A judge that reads text and follows instructions can itself be talked to by that text. Using a language model to catch prompt injection means pointing the thing that's vulnerable to prompt injection at the prompt injection. Regex doesn't read instructions. It just matches.
Reason 3: cost
Metered inference means paying per token, per chunk of text a model processes. If every request in your app triggers one or more judge calls, your security bill grows with your traffic, and it grows fastest exactly when someone's hammering you.
G8KEPR's detection runs locally: "$0 per token." The cost doesn't climb with the number of words going through it.
Reason 4: nothing leaves
If the judge is a hosted model, every prompt and every answer you want checked gets sent to someone else to be checked. That's often the very data you were trying to protect.
G8KEPR is built so that "0 bytes leave your VPC." A local, classical approach is what makes that possible. More on that in Why It Runs in Your VPC.
The numbers, in plain terms
Here's what's published, and what it means.
| Metric | Published | In plain English |
|---|---|---|
| Median detection time | 4.6 ms | Half of checks finish faster than this |
| Recall (held-out) | 0.788 | Of the real attacks in the test set, it caught about 79 in 100 |
| Precision | 0.857 | Of the things it flagged, about 86 in 100 were real attacks |
A few words on those.
Held-out means the test examples were kept separate from the ones used to build the models. It's a check on whether the detector learned something general, not just memorized its homework.
Recall is about misses. A recall of 0.788 means roughly one real attack in five, in that test set, got through without being flagged. That's not nothing, and I'm not going to pretend it is.
Precision is about false alarms. At 0.857, roughly one flag in seven was something harmless.
The two pull against each other. Tune a detector to catch everything and it starts crying wolf. Tune it to never cry wolf and it misses more. These numbers are where it sits today, not a ceiling.
Note
No single detector catches everything, including the LLM-based ones. That's part of why G8KEPR has four pillars and a cross-pillar correlation engine that scores signals co-occurring across them. An attack that slips past one checkpoint still has three more places to show up.
What I'm trading away
Honestly? Some flexibility. A language model judge can sometimes recognize a clever, never-before-seen phrasing that a pattern or a small classifier won't. That's a real advantage, and it's the reason the approach is popular.
I decided fast, repeatable, cheap and private mattered more for something sitting in the path of every request. You can disagree with that call, and some people will. I'd just rather make it out loud, with the numbers published, than hide the trade behind a vague "AI-powered" label.
If you want to see how it holds up on your own traffic, g8kepr.com has a free trial.