,

Token Spend Is Not a Cost to Cut. It Is a Portfolio to Manage.

Token Spend Is Not a Cost to Cut. It Is a Portfolio to Manage — featured image

Cutting the AI bill misses the point. The organisations getting value from AI treat token spend as a portfolio to allocate, diversify and rebalance, not a number to shrink.

Every organisation running AI at scale eventually has the same conversation.

Finance sees the token bill. It has grown faster than almost any other line in the technology budget — by one recent estimate, AI is now the fastest-growing expense in corporate technology spend, and for some firms it already accounts for up to half of the entire IT budget. Someone asks the obvious question:

How do we bring this down?

It is the wrong question. Not because cost does not matter, but because it treats every token as the same kind of spend, deserving the same scrutiny and the same downward pressure. Some token spend eliminates hours of manual work every week. Some produces a marginally faster first draft nobody needed urgently. Some sits behind a customer-facing product where a cheaper model would quietly damage trust.

Treating all three the same way, as one number to minimise, throws away the information that actually matters: which spend is working.

I have watched a version of this meeting before, with different technology under a different budget line. The mainframe-to-client-server migration had it. So did the move to cloud, and again when waterfall gave way to agile. Someone always looks at the new spend and asks why it costs so much, when the better question is what it is buying and whether that’s the right amount of it.

The organisations getting genuine value from AI are not the ones with the smallest token bill. They are the ones who can explain, tier by tier, why each part of that bill exists, and what happens if it stops. That is portfolio management, not cost control.

That is portfolio management, not cost control.

Tokens are not one asset class

Portfolio managers do not evaluate a single “amount invested.” They look at asset classes — cash, bonds, equities, alternatives — each with a different risk and return profile, held in deliberate proportions for different reasons.

Token spend inside an organisation works the same way, whether anyone has named it or not. It usually breaks into four tiers.

Diagram of four token spend tiers: core operational, decision-support, customer-facing, and experimental

1. Core operational tokens

High volume, low individual risk. Internal search, drafting routine documents, summarising meeting notes, first-pass triage. The value here is efficiency, not judgement.

These tasks rarely need a frontier model. The failure mode is usually invisible: routing high-volume, low-stakes work through an expensive model by default, because nobody revisited the choice after the pilot ended.

2. Decision-support tokens

Spend behind analysis and synthesis that feeds a real decision. A report a leader will act on. A recommendation going to a customer. A first draft of something with commercial consequences attached.

This tier needs a stronger model and more oversight than tier one. The cost of being subtly wrong here usually exceeds the cost of the tokens themselves by an order of magnitude.

3. Customer-facing and high-stakes tokens

Usually the smallest volume, and the highest consequence per token. Anything a customer, regulator or auditor will see directly. Capability, accuracy and governance matter most here, and cutting cost is the most expensive mistake an organisation can make in this tier.

4. Experimental tokens

A deliberate, bounded allocation for trying new use cases and testing what a model can do before it earns a place in tiers one to three. A higher failure rate here isn’t a problem to fix. It’s the research and development line doing its job, and it should be sized and reviewed like one, not judged against the same success bar as production spend.

Allocation is a decision, not a default

Most organisations never consciously choose this split. It emerges from whichever team moved fastest, whichever tool was easiest to plug in, whichever model got pre-approved first. The result is usually the opposite of what a deliberate allocation would produce: expensive models doing cheap work, and cheap models quietly carrying risk they were never evaluated for.

A portfolio approach treats the split itself as a leadership decision. What proportion of token spend should sit in each tier this quarter, and who decided that, and on what basis? If the honest answer is “however it happened to grow,” the allocation isn’t being managed. It’s being observed after the fact.

Diversification protects against a single point of failure

A portfolio concentrated in one asset is fragile. The same is true of an organisation routing every task through a single model or provider.

A portfolio concentrated in one asset is fragile. The same is true of an organisation routing every task through a single model or provider.

Prices change. Capability changes, sometimes overnight, with a new release. Outages happen. A model can be deprecated with a few months’ notice, taking every workflow built on it down at once. Diversifying token spend across models matched to tier — a cheap, fast model for tier one, something stronger reserved for tiers two and three — isn’t just a cost decision. It’s resilience.

Rebalancing: this changes underneath you

Portfolios get reviewed and rebalanced on a cadence, because the ground shifts under a static allocation. Token portfolios shift faster than most I’ve worked with.

Model prices fall. Capability improves. A task that genuinely needed a frontier model six months ago may now run just as well on something far cheaper. A tier-one task that has quietly grown in scope may now carry risk it didn’t carry at launch, and belongs a tier higher than where it started.

Without a review cadence, an allocation that was correct at launch slowly becomes wrong, and nobody notices because nothing broke. Set a rebalancing review — quarterly is a reasonable default — and ask the same two questions each time. What can move down a tier now that it’s cheaper to do so safely? What has to move up because the stakes have grown?

An allocation that was correct at launch slowly becomes wrong, and nobody notices because nothing broke.

Governance: who actually owns this

In most organisations, nobody does. Not really. Finance owns the budget cap, a blunt instrument that can’t distinguish a tier-one saving from a tier-three risk. Engineering owns model routing, optimised for latency and cost, usually without visibility into which decisions downstream actually depend on the output.

The fix isn’t another approval layer. It’s naming an owner: a small forum with finance, engineering and the business represented, that reviews tier definitions, current allocation and rebalancing decisions together, on the same cadence, with something like the seriousness an investment committee applies to capital.

Questions worth asking

  • What tier does each major AI use case actually belong to, and does the model behind it match that tier?
  • What proportion of token spend sits in each tier right now, and was that decided, or did it just happen?
  • Which tier-one tasks are quietly running on a tier-three model, paying for capability nobody needs?
  • Which tier-three tasks are quietly running on a tier-one model, under-provisioned for the risk involved?
  • Who reviews this allocation, and how often?

Token spend is not a cost to cut. It is a portfolio to manage.

Token spend is not a cost to cut. It is a portfolio to manage. A capable organisation does not have the smallest AI bill in its sector. It has the clearest answer to what each part of that bill is buying, and why.

Sources used

  • Deloitte, Tech Trends analysis (January 2026) — AI cited as the fastest-growing expense in corporate technology budgets, reaching up to half of IT spend at some firms.

Comments

Leave a Reply