The Rental Half of the Invoice: What Renting AI Actually Costs
By LumaVista Team
Here’s one month of one coding agent. Not a team, not a company — one agent, running on rented intelligence. The total at the bottom is about $4,700. It’s a modelled bill, not a screenshot — but every line on it is priced off a published rate card, and you’ll recognise all of them by the end of this article.
Nobody bought a GPU. Nobody called an electrician. Every line on that bill felt small when it happened — a few dollars of tokens here, an hour of rented compute there. That’s not an accident. That’s the entire business model.
Last time, we priced the machine: two million euros of silicon to run a frontier open model, plus the power bill, the cooling, and the depreciation nobody says out loud. And that was only half the invoice. This is the other half — the one you rent. Renting feels cheap the way leases feel cheap, so let’s do the math the pricing page is quietly hoping you’ll skip.
The one formula
Part one had a single formula: parameters times bytes equals memory. Part two has one too, and it’s even simpler.
A machine you own costs the same asleep or thinking. A machine you rent only costs when it thinks.
That asymmetry is the entire cloud economy. Every subscription, every per-token price, every “serverless” pitch is that one sentence wearing a costume. And it means every rent-or-buy question collapses to a single number: how much of the time is the meter actually running? Below some threshold, renting wins — genuinely, not as a trick. Above it, the machine you own is the cheap option and everything else is paying someone else’s margin.


Utilization isn’t one factor in the price. Utilization is the price. Hold onto that — the rest of this article is just following it down four rungs.
The rental ladder
The hardware ladder in part one had four rungs, a thousand-fold apart. The rental ladder mirrors it, rung for rung — and the rungs line up more than the vendors would like.
Rung one: the subscription. Around $20 a month for a chat window — ChatGPT Plus, Claude Pro, Google AI Pro all cluster there. For a human asking questions, it’s honestly the best deal in the history of computing. Remember that sentence. It stops being true in about two rungs.
Rung two: the token meter. The API. You pay for exactly what the model reads and writes, by the million tokens. A mid-size open coding model runs about $0.25 per million tokens in and $1.00 out; the open frontier — Kimi K3, the 2.8-trillion-parameter model from part one — costs $3 in and $15 out. Input is cheap, output costs three to five times more, and cached context is nearly free. The meter is honest. It’s also very, very patient.
Rung three: the rented GPU. The same cards from part one, rented by the hour instead of bought by the crate. A single RTX 4090 goes for $0.34 to $0.69 an hour; an H100 lands around $3 an hour on demand — call it $2.50 to $4.30 depending on how good the provider’s marketing is — and roughly $2 on spot if you can tolerate eviction. Instead of thirty thousand euros up front, you pay for the hours you use.
Rung four: the reserved cluster. A yearly commitment with a discount of 20 to 60 percent off the hourly rate — which looks generous right up until you realize you’ve just bought the machine, on their books instead of yours. Reserved capacity is a commitment to pay, not a commitment to use: AWS will tell you plainly that a reservation is billed “whether you run instances in reserved capacity or not.”
Let’s race them
Take one workload — a coding agent doing real work — and run it three ways: the machine you own, the GPU you rent, the tokens you buy. Same silicon, same model, same work. The only thing we’ll change is how many hours a day the meter runs.
Here’s a mid-size coding model on a single 4090, priced per month. Don’t read this as three numbers — read it as three shapes.
The same three numbers, if you want them precisely:
| Hours per day | Own the box | Rent the GPU | Buy tokens |
|---|---|---|---|
| 1 h/day | ~$126 | ~$12 | ~$11 |
| 1½ h/day | ~$129 | ~$24 (2 h billed) | ~$16 |
| 8 h/day | ~$168 | ~$96 | ~$84 |
| 24/7 | ~$264 | ~$288 | ~$252 |
Once you see it as shapes, the argument stops being about numbers at all. Owning starts high and barely tilts: $120 has already left your account before the machine does a single useful thing. Its only moving part is electricity, which is why the line is smooth — a kilowatt-hour doesn’t come in a minimum size. The tilt is your tariff, not your habits, and it’s steeper than people expect: at €0.15 a kilowatt-hour the round-the-clock total is $192, and at €0.45 it’s $336. Same box, same work, a $144 spread that has nothing to do with AI.
Renting is a staircase, and that shape is the whole difference between the brochure and the bill. GPUs sell by the hour. Ask for ten minutes and you buy sixty; the other fifty are billed to you while the card sits there doing nothing. Per-second providers file the steps down but don’t remove them, because you’re also paying while the machine boots and loads the model — minutes you’re charged for and get no tokens out of. Run an hour and a half a day and you don’t pay for ninety minutes, you pay for a hundred and twenty. Tokens are the only line that bills what you actually consumed, which is why that one is a smooth climb from zero.
Now read the crossing. Past nineteen hours a day, renting has to buy a twentieth hour it can’t use, and the owned box wins from there on — earlier than the smooth arithmetic suggested, and for a reason that isn’t about GPUs at all. But look where the tokens line sits: underneath everything, the whole way across. At one hour a day it’s twelve times cheaper than owning; at full time it’s still the cheapest of the three. One honest caveat — that last part is the electricity assumption talking. Drop your power to €0.15 and the owned box slides under the token line somewhere around sixteen hours a day. Which is the real lesson: at this rung the answer isn’t set by the model or the silicon. It’s set by your utility bill and your billing increment.
Hold that, because it’s the setup. At this rung the meter is genuinely the right answer for almost anyone, and the numbers are small enough that nobody gets hurt guessing. Now watch what happens when the customer stops sleeping — and when the model gets expensive.
The twist: your agent never sleeps
Humans are spiky. You ask, you read, you think, you wander off to lunch. A $20 subscription monetizes all the time you’re not using it — you’re paying for a machine you touch in bursts and leave idle for the rest. That idle time is the product.
Agents are the opposite. An agent is a loop. It reads, it writes, it retries, it re-reads its own context — around the clock, if you let it. Anthropic puts the contrast bluntly in its own docs: “one debugging session can consume more than a day of chat.” The workload that made owning hardware make sense in part one is the exact workload that breaks rental math in part two.


You can watch this happening in real time in the pricing pages. In August 2025 Anthropic added weekly rate limits aimed squarely at people running Claude Code “continuously in the background, 24/7.” Cursor tore up its unlimited plan in June 2025 because “the hardest requests cost an order of magnitude more than simple ones.” GitHub moved Copilot to token-based billing in June 2026, admitting outright that “a quick chat question and a multi-hour autonomous coding session can cost the user the same amount,” and that the old flat model “is no longer sustainable.” Every one of those is the same story: the subscription was priced for a human, and the agent isn’t one.
Here’s the part that matters for your own bill. The real crossover isn’t a time on the clock — it’s utilization, how full you keep the box. And here’s the uncomfortable result, on our own modelling: retail token pricing is extremely hard to beat. It wins because you pay only for your slice of a machine the provider keeps saturated across thousands of customers, while owning or renting means paying for the whole box whether you fill it or not. You don’t beat that by being clever. You beat it by supplying the utilization the provider was supplying for you — all of it.
Which is why the honest answer to “when should I own?” is later than you think, and for reasons that aren’t on the invoice.
The rental invoice
Just like the hardware had silent lines, the rental invoice has its own. Four of them.
Egress. On the big clouds, your data checks in for free and checks out at a price — around $0.09 a gigabyte to leave AWS or Azure, more on Google. Worth knowing: the GPU-specialist clouds flip this — RunPod, Lambda, and CoreWeave all charge zero egress. The exit tax is a hyperscaler habit, not a law of nature, and knowing which kind of provider you’re on is real money.
Storage and idle. The pod you paused still bills its disk — on RunPod, a stopped pod’s volume actually doubles from $0.10 to $0.20 per gigabyte-month the moment you stop it. And reserved capacity bills whether you think or not. That one you already know — it’s ownership cosplay.
The retries. Agents fail and try again, and the meter counts every attempt. Worse, an agent re-sends its whole conversation every turn — “a one-line question in a session that has been open all day uses tokens for the whole conversation, not just the one line.” Prompt caching softens this a lot: a cached token costs about a tenth of a fresh one, and on some providers’ own APIs closer to a fiftieth. But caching isn’t free either — writing to the cache costs 1.25× to 2× a normal input token, so a workload that keeps missing pays the premium again and again. And caching only covers the unchanging prefix; the growing history and the retries are still full price.
The reprice. This is the one you cannot lock. Your hardware’s cost was fixed the day you bought it. Your provider’s price is a variable you don’t control, on a bill you’ve built your business on. In April 2026 Anthropic briefly pulled Claude Code from its $20 Pro plan. It was reversed in under a day, and the company later described it as a test on a small slice of new signups — but the reasoning was the part worth keeping. The plans, its head of growth said, “weren’t built for this”; the Max tier had been “designed for heavy chat usage, that’s it,” and shipped before Claude Code existed. Nobody’s bill actually moved. The point is that it could have, overnight, on a page you don’t control.


The verdict: own the floor, rent the peaks
So — rent or buy? Wrong question. The math gives a better answer: own your floor, rent your peaks — but only once your floor is big enough to be worth a building. The workload that runs flat-out every day, at a scale that would fill the machine, belongs on silicon you own or reserve. The spike you hit twice a year belongs on someone else’s. And below that line — which is most people, most of the time — the meter genuinely is the cheap option, and anyone who tells you otherwise is selling hardware.
The frontier makes this vivid. Running one heavy agent on Kimi K3 around the clock costs roughly €2,700 a month in API tokens. Owning the cluster that serves that model — the €2 million of GPUs, the 65-to-80-kilowatt power draw, the cooling, the ops engineer, spread over three years — runs about €77,000 a month, fixed, whether it’s busy or not. For one agent, the API is nearly thirty times cheaper. Owning only catches up once you’re running the equivalent of about thirty constant agents — the “continuous, massive demand” that describes a handful of organizations and almost nobody else. It’s the same reason you rent a concert hall for one night instead of building one: you need the capability occasionally, not the building forever.
That’s the whole invoice. Part one, the machine. Part two, the meter. Run the crossover on your own numbers before the next bill runs it for you.
But there’s a third bill, and it never shows a number. Every prompt you send to a rented machine is read by that machine. What that costs — and the kind of hardware that can prove it isn’t reading — is where this story goes next.
What to do now
- Find your real utilization first. Before you argue rent-vs-buy, measure how many hours a day your workload actually runs the meter. That one number decides the whole thing.
- Match the rung to the job. A human asking questions belongs on the $20 subscription. Don’t reach for a rented GPU to do what a chat window does for pennies.
- Price your agents separately from your people. A coding agent can burn more in a session than a person burns in a day — budget it as compute, not as a seat.
- Turn on prompt caching and clear stale sessions. Cached context is ~10× cheaper to read, and a day-old open session re-bills its whole history every turn. Caching does cost extra to write, so it pays off on stable prefixes you hit repeatedly — not on context that churns.
- Read the silent lines before you commit. Check egress, paused-pod storage, and reserved-idle terms — they’re where “cheap” quietly stops being cheap.
- Own your floor — but check that you have one. Owned or reserved capacity pays off when a workload would keep the machine genuinely full, which at the frontier means many constant agents, not one. Below that, the meter wins on cost, and owning has to earn its keep some other way.
- Assume the price will move. Don’t build a business on a plan you can’t lock. Keep the open-weight exit open, and re-check the numbers every couple of months — they change that fast.