Own LLM Hardware or Cloud? No Gut Feelings. Just Numbers.
€3,800 for your own AI machine pays itself off in 5.6 years. Sounds good. Only true if you leave out half the calculation. We don't leave it out.
It's a tempting idea: instead of renting every token in the cloud, buy the hardware once and stop paying per request. Owning instead of renting. Feels like control, independence, common sense.
Feels like. We prefer to calculate rather than feel. So: spreadsheet open, assumptions in, look honestly.
Note: All prices and model specs below are example assumptions (as of 2026, a ~31B open-weight model on both sides, cloud price ~$0.36 / 1M tokens via OpenRouter). Your numbers will differ. That's exactly what the calculator below is for.
Important: This calculation compares the same model on both sides — local and in the cloud. Anyone comparing a small local model against a significantly more capable cloud model is comparing apples and oranges. That's a different decision requiring different numbers.
The Calculation Everyone Loves to Run
Cloud side. A ~31B open-weight model via OpenRouter, $0.36 per million tokens.
- 8.6M tokens/day → $3.10/day → approx. €2.85/day
- Monthly: ~€85
- Annually: ~€1,040
Hardware side. The same model, running locally on a Ryzen AI box with 128 GB RAM for €3,800. Around 100 tokens/second, giving a theoretical 8.64M tokens/day at full load.
- Power (optimistically estimated): €1/day → €365/year
- Saving: €1,040 − €365 = €675/year
- Break-even: €3,800 / €675 ≈ 5.6 years
On paper the hardware wins. In production it rarely does. Let's look at why.
Where the Calculation Falls Apart
Utilisation. 100% is fiction. No inference machine runs flat out around the clock — maintenance, updates, idle time, quiet nights. Realistically you get 50–80%. You paid for 8.64M tokens/day of capacity and use maybe 5–6M of it. The rest is paid-off silence. That alone stretches break-even towards 10–12 years.
Lifespan. After 3–5 years the box isn't broken, just old. New hardware by then will be twice as fast at half the cost per token. Planning for 5.6-year break-even means betting on a machine that will already be technologically obsolete halfway through its payback period.
Power. This is where it really breaks. €1/day sounds harmless — and assumes an average draw of around 140 watts. A 128 GB machine with AI acceleration under load draws closer to 500–800 watts. Let's calculate honestly:
€0.30/kWh × 0.7 kW × 24 h = €5.04/day ≈ €1,840/year
And now the saving: €1,040 (cloud) − €1,840 (power) = −€800/year.
Negative. The meter eats more than the cloud costs. The box doesn't break even in 5.6 years — it never breaks even. And that's before counting the €3,800 purchase price at all.
Maintenance and overhead. Cooling, fans, the occasional failure, optimisation tools, and — the most expensive item of all — your time. Nobody maintains a cloud bill at three in the morning.
Model switching. If you're running a 31B model locally today and want to move to a better one tomorrow, you buy new hardware or adjust your quantisation. In the cloud you change a parameter in the API URL — and pay the new price. That's not an argument for cloud per se, but a flexibility factor that doesn't appear in a pure cost calculation. Keep in mind: a bigger model means higher costs on both sides.
Scalability. You scale cloud with a flag. You scale hardware with an order, a delivery window, and a new power contract.
The Power Bill Is the Whole Story
If only one number stays with you, make it this: the entire "hardware saves money" thesis hinges on the electricity price and average power draw. The optimistic figure (€1/day) isn't a pessimism offset — it's simply unrealistic for the machine we're talking about. Plug in real watts and real cents per kilowatt-hour, and the business case flips from savings model to loss-maker.
When Own Hardware Still Makes Sense
We're geeks, not cloud salespeople. Own hardware isn't inherently wrong — it's just rarely right for cost reasons. There are good reasons to go local:
- Data sovereignty & compliance. When data cannot leave the building, "per token at a US provider" isn't an option — whatever it costs. Sometimes cloud isn't cheaper, it's simply off-limits.
- Latency & offline. Local is local. If you need milliseconds or network independence, you're buying proximity, not tokens.
- Predictable fixed costs instead of token roulette. A capital investment is fixed. A cloud bill scales with usage — upward too, and uncomfortably so.
- Learning on real hardware. Running your own models teaches you inference, quantisation, and bottlenecks in a way no billing dashboard ever will.
- Genuine sustained load plus cheap power. If the box really runs near full capacity around the clock and draws cheap electricity (own solar? off-peak tariff?), the maths swings back in your favour.
Our Conclusion
For most workloads at European electricity prices, cloud wins on pure cost. Full stop. Buy own hardware for data sovereignty, latency, or learning — not to save money, unless you have cheap power and genuinely constant load.
Above all: don't decide on gut feeling. Decide with your numbers. The optimistic calculation always looks good. The honest one often looks different. We'd rather give you the honest one.
And yes — the answer isn't always 42. Sometimes it's: it depends. But at least you can now calculate it yourself.
Enough talk. Here's the calculator. Put your numbers in.