Start with the concession, because a chief financial officer will find it anyway. At ordinary interactive volume, cloud AI is cheaper than any box in your building. A team that chats with an assistant for a few hours a day costs a few hundred euros a month in tokens. Nobody should pretend otherwise.
The problem is not the size of the bill. It is that three variables in the bill are outside your control, and they change without a signature from you.
Variable one: the price
On 30 June 2026 Anthropic launched Claude Sonnet 5 and wrote: “It is priced at $2 per million input tokens and $10 per million output tokens.” The same post called it introductory pricing, with a rise to $3 and $15 scheduled for 1 September. On 10 August the post was edited: the introductory price “is now permanent”.
That is a vendor behaving well. It announced a rise weeks ahead, then chose not to take it. It is also a 50 % swing in a budget line, decided twice in ten weeks, in San Francisco. A finance team in Cluj read about both decisions afterwards.
Variable two: the volume
Per-token billing means the bill tracks adoption. A calculation, so you can check it: a 50-person team at 100 million tokens a month uses about 95,000 tokens per person per working day, or roughly 24 busy turns of 4,000 tokens. At $2 and $10 per million tokens, with 80 % of the traffic as input, that month costs about USD 360, or roughly EUR 330.
Now the team gets good at it. Agents draft, review and re-draft without a person typing each turn. The same arithmetic at ten times the volume is ten times the bill, and the finance line that was EUR 330 is EUR 3,300, in the same quarter, with no purchase order. A line that doubles when a team becomes productive is a line nobody can budget.
Variable three: the model
The model you priced is not the model you will run. Vendors retire models on their own timetable. OpenAI’s help centre records that ChatGPT retired GPT-4o, GPT-4.1, GPT-4.1 mini and o4-mini on 13 February 2026, and its API deprecations page lists the removal of the chatgpt-4o-latest snapshot on 17 February 2026. Each retirement moves your traffic to a model with a different price and a different behaviour, and your prompts, your tests and your training material follow.
The cloud bill is not expensive. It is unowned. Three people you have never met set it, and each of them is doing their job well.
Where the arithmetic turns
A flat monthly fee is a bad deal for a team that chats, and a good one for a team that automates. Our own calculation puts the crossover for one Reasoning box against a frontier cloud model at roughly 400,000 tokens per user per working day. That is agent and document-processing volume, not conversation. Below it, cloud wins on price. Above it, the flat fee wins, and it wins more every month adoption grows.
With Bastion
What Bastion changes
Bastion is one monthly fee for a box that serves your team. The fee covers the hardware, the model refreshes and the support, and it does not move when usage does. The model on the box changes only when you accept the update, because nothing on the box calls home.
The limit, stated first: if your team’s use of AI is ordinary chat at ordinary volume, the cloud is cheaper and it will stay cheaper. Bastion is bought for the other two reasons on this site, where the data goes and who holds the switch. The cost case starts only at automation volume.
On the record
- 1
- 2
- 3