G8KEPR
A Budget for Your AI Bill

Part 14 of the thread Building G8KEPR
- Model providers charge by usage, so anything that makes an app call a model more than intended costs real money.
- Agents stuck in loops, bugs, abusive users and leaked keys can all drive sudden spend spikes.
- G8KEPR's cost controls include per-organization budget caps, alerts on token-spend spikes, and the cost of each response reported back.
- Dashboards break spend down by model and provider, so you can see where the money actually goes.
Cost doesn't sound like a security topic. It sounds like something for the finance team to worry about at the end of the month. I put it in the security product anyway, and this post is about why.
How AI bills work
Most model providers charge by the token. A token is a chunk of text, roughly a short word or a piece of a longer one. You pay for the tokens you send in and the tokens the model writes back, and the price varies a lot between models.
That's a perfectly fair way to price things. It also means your bill is directly tied to how many times your app talks to a model and how much it says each time. Anything that makes that number jump makes your bill jump with it, and it happens in real time, not at the end of the month.
How it gets away from you
Here are the usual ways spend runs off the rails. None of them require anybody to be especially clever.
Runaway agents. An AI agent is a model that's allowed to take steps on its own: call a tool, look at the result, decide what to do next, repeat. That's what makes agents useful. It's also what lets them keep going long after they should have stopped.
Loops. A close cousin. The agent calls a tool, the tool returns an error, the agent tries again, gets the same error, tries again. Or two automated steps keep handing work back and forth. Each lap costs money, and a loop doesn't get tired.
Bugs. A change ships that accidentally sends a giant document with every request instead of a short summary. Nothing breaks. The answers look fine. The bill is just much bigger.
Abuse. Someone finds your public AI feature and decides to use it as a free model for their own purposes. Or they hammer it on purpose, because running up your costs is its own kind of attack.
Leaked keys. If a provider key escapes, somebody else is now spending on your account.
In every one of these, the first sign is often a number going up faster than it should. That's why I think of cost as a security signal and not just an accounting one. It's frequently the earliest evidence that something is misbehaving.
What G8KEPR does about it
The cost controls live in the AI Gateway, in the request path, so they see every call on its way to a provider. At the level the product describes them:
- Per-organization budget caps. You set how much an organization is allowed to spend, and the gateway holds that line instead of letting spend run open-ended.
- Alerts on token-spend spikes. When spend jumps in a way that doesn't look normal, you hear about it while it's happening, not when the invoice lands.
- Cost per response. The cost of each completion is reported back, so you can tie spend to the specific requests and features causing it.
- Dashboards. Spend broken down by model and by provider, so "we spent a lot on AI this month" turns into "this model, on this provider, for this feature."
That third point is my favorite, honestly. A monthly total tells you something went wrong. A cost attached to each response tells you where.
Two different bills
One thing worth separating out. The cost controls are about your spend with model providers. G8KEPR itself doesn't meter anything: there's no per-token charge for running traffic through it, because it runs inside your own VPC and does its detection there. I explain that choice in Why It Runs in Your VPC. It would be a little odd to sell a tool for controlling a per-token bill that added a second per-token bill on top.
The same instinct, different decade
This isn't a new idea, just a new place for it. People have been writing scripts to catch cloud spend creeping up for years. I have one on this site for monitoring Azure resource costs. The difference with AI is speed. A forgotten VM costs you a little every hour. A looping agent can cost you a lot every minute.
Budgets, alerts and per-request visibility are boring. That's the point. They're the kind of boring that stops a Friday afternoon bug from becoming a Monday morning conversation.
If you want to see how the cost pieces fit alongside the other gateway checks, the pillars post has the overview.
- One Gateway, Many ModelsG8KEPR