From Token-Maxing to AI Wallets: Key Takeaways from Human[X] Amsterdam 2026
From vanity metrics to cost-effective engineering: how to structure enterprise AI architectures to protect your bottom line.
I spent three days at the first European edition of Human[X] in Amsterdam. Whatever the topic on stage — legal, cybersecurity, search, or infrastructure — the Q&A kept coming back to the exact same question: How much is AI actually costing us? As Parser’s CTO, leading a software consultancy specialising in AI, I hear this question from clients on a weekly basis. Below, I’ve synthesised my primary insights and key observations on AI cost management gathered from the event.
1. The Token-Maxing Era is Over
Transitioning from vanity consumption to precise token attribution
Until recently, AI transformation was measured by leaderboards of who consumed the most tokens. It felt like déjà vu of measuring engineering productivity in lines of code written instead of functionality delivered. Just because a metric is easy to count does not mean it measures value.
Now, cost is an urgent concern across the enterprise:
- Glean (enterprise AI search) reported that tokens account for roughly a third of total client spend before optimisation.
- Legora (AI legal platform) tracks and assigns token spend directly to individual legal matters so costs can be passed on to clients. This practice — token attribution — is becoming standard.
- Palo Alto Networks (a cybersecurity company) argued that while raw intelligence will become commoditised over time, compute costs remain high today. Crucially, most cost is incurred at inference (OpEx), not training (CapEx).
Key Insight: You can’t optimise what you can’t attribute, and attribution isn’t automatic. In our client work, agents running through cloud model hosts often don’t send usage back to the provider’s dashboards at all. Without an LLM gateway or an OpenTelemetry export, it is hard to say what a team or a customer actually spent.

Place an LLM gateway in front of every model call, tag tokens to specific projects, and track cost per successful task, not cost per raw token.
2. The AI Wallet vs. Automated Routing
Don’t put model selection burden on the end user
Atlassian introduced the concept of an AI Wallet: each person gets a budget and a menu of models, and has to spend it wisely instead of defaulting to the newest, most expensive model. Many other companies use similar monthly individual budgets.
I like the intent, but I think it underestimates what it asks of people. Choosing the right model for each task takes training, experimentation, measurement, knowledge and, above all, time. Most people don’t want to spend that time keeping up with which model was released last Tuesday, and they shouldn’t have to.
My take: Make that decision for them with automatic LLM routers that select the model based on the task, plus an easy way to escalate to a better model when the result isn’t good enough. The pattern we see working is cheap by default, escalating on low confidence. Most routine requests (often 60 to 80% of coding prompts) can go to a small, cheap model, and only the minority that genuinely needs a frontier model goes there.

The wallet sets the budget, and the router spends it well.
One honest caveat: routers aren’t magic either. A router whose confidence signal is poor sends hard tasks to weak models. It saves money on paper and loses it in rework, and the degradation is silent. It needs evals and measurement behind it, and someone accountable for them. The difference is that the effort sits with a few experts rather than with every single employee.
3. Inference is OpEx And Scales With Your Customers
Nebius (AI-native cloud) highlighted, made the economics concrete. Training is CapEx, while inference grows with every user. Context windows have grown in the last few years from about 1.5k tokens to as much as 120k, roughly 20 books or an average-sized GitHub repo. Reasoning traces keep getting longer, and tool outputs often consume more tokens than the answer itself. Agents’ tool calls also run on CPU, not GPU, which is another line on the bill that nobody budgeted for.
Key Takeaway: Every successful AI product gets more expensive as it grows, so architecture is now a cost decision. I would enforce request caching along with using worker sub-agents running cheaper models (for example, open-weight models running locally) to process the bulky information and hand back short summaries of a few thousand tokens to the orchestrating agents on frontier models.

Frontier intelligence should be reserved for high-value judgement, not spent reading raw logs.
4. 100x productivity means 100x database queries
Cockroach Labs (a high-performance database provider) pointed out that if agents make us 100x more productive, they will also query our databases at least 100x more. Nobody wants a database bill that’s 100x bigger, so evaluating consumption costs must be integrated into the design right from the start, rather than catching everyone off guard at quarter-end.
What I learned: In multi-agent, continuously running systems, the model bill is only the visible part of the iceberg. The cost of scaling the supporting infrastructure (databases, CPU, storage, observability) has to go into the ROI calculation from the first business case. The cheapest control is also the least glamorous: bounded agent loops with budget ceilings, iteration limits and kill switches, enforced by the infrastructure rather than written into the prompt.

Prompt-level constraints are merely suggestions. True guardrails must be enforced at the infrastructure level.
5. Cheaper Models Won’t Fix a Badly Framed Problem
Kerem Tomak (author of Learning AutoML) explained that classic AutoML needed a human to frame the problem precisely before it searched for the best model. GenAI can now question that framing, refine it iteratively with you, and then choose the model or simply write a new one. Automation optimises a problem you framed, while autonomy takes part in the framing.
My insight: The most expensive mistake happens before any model runs. If you frame the wrong problem, the perfect solution to it will be built at light speed, delivering results you don’t want.

Tokens spent optimising the wrong problem are the most wasteful tokens of all.
My to-do: spend the first tokens of every project on challenging the question and writing it down as a spec, and add agents only when a single one clearly hits its limit.


