Knowledge hubCost

I asked a few CFOs what a token is

I've been talking to CFOs about what their AI actually costs.

What keeps coming up is this: AI spend doesn’t behave like most software spend, it moves with usage, often unpredictably. That’s new. And the vocabulary around it was written by engineers, for engineers, with no one stopping to translate. 🈹​

So, translation.

Three things set the bill: how many tokens each request burns, which model those tokens go to, and where the model runs.

None of the three appear on a pricing page. All three are knowable by Friday.

What are you actually paying for?

A model is, essentially, a huge collection of numbers that has learned patterns from data. Two words explain how that turns into cost.

Take a model that can tell whether a photo shows a cat or a dog.

Training 🏃‍♂️‍➡️​(see what I did here?) = show the model millions of examples until it learns to tell cats from dogs.

Inference = give it a photo it hasn’t seen before and ask: cat or dog?

Inference happens every time someone uses the model. That is the part you keep paying for after the model has been trained. It happens every single time anyone asks it anything, and you pay for every one of those times. ​​💰​

Sometimes inference is used more broadly to mean the infrastructure required to run and serve the model. For your purposes, the distinction doesn’t matter much: this is the ongoing cost of actually using AI.

Gartner's August forecast has inference spending passing training for the first time this year: $23.3 billion versus $19 billion.

Models don't read words. They read tokens: pieces of words that vary in size, but average roughly 0.75 words each.

Input is everything you send to the model: instructions, pasted documents, tool definitions, the conversation so far, and the question itself.

Output is everything the model sends back. And output usually costs much more than input, because generating text requires more computation than processing it.

Example: Anthropic prices input and output tokens separately. Here, output tokens cost x5 as much as input across these models.

The model you choose changes the price

A million tokens is not a million tokens when it comes to cost.

Different models charge very different prices for processing the same amount of text:

*Open-weight models. Prices shown are approximate median hosted API prices across providers. If you run the model yourself, there is no per-token vendor price - you pay for the compute instead.

That means the model you route a task to can change its cost dramatically - even when the task itself stays the same.

And those prices move.

When Anthropic launched Claude Sonnet 5, it priced it at $2 per million input tokens and $10 per million output tokens, with plans to raise that to $3/$15 on September 1. On August 10, Anthropic cancelled the increase and made the lower pricing permanent.

Once AI is part of your product, an LLM provider can change your unit economics with a pricing update🤯​

Why the agent tripled the bill

Here’s the property behind a lot of AI cost surprises: a model has no memory between requests.

So if your product needs it to know what was said five minutes ago, that context has to be sent again. And you pay for those tokens again.

A chatbot usually answers a question.

An agent does a job - look something up → decide what to do next → call a tool → inspect the result → try again if needed → keep going until it’s done.

Each step can mean another model call. And each call may carry more context than the one before it.

That adds up quickly. Gartner estimates that agentic tasks can consume 5 to 30 times more tokens than a standard chatbot task!

And the harder part is that the cost isn’t always predictable.

Researchers at Stanford’s Digital Economy Lab ran agents repeatedly on the same tasks and found that token usage could vary by as much as 30x between runs. Same task, same setup - the agent simply took a different path.

That unpredictability is already showing up in company budgets. KPMG surveyed 2,145 senior leaders and found that 49% had scaled back agent deployments because operating costs outweighed the benefits.

One thing worth saying plainly: your engineers probably aren’t withholding this number. In many companies, the number simply hasn’t been built yet.

One API key might serve six products, and the provider sends back one bill. Nobody has historically needed to separate the cost by feature, customer, workflow, or agent.

Finance asking the question is often what causes that visibility to get built.

And that’s really the point.

You don’t need to become an AI engineer to understand AI spend. You need to know what is being used, which model is handling it, where it runs, and what each workload costs.

Once you can see those four things, AI stops being a mysterious bill and starts becoming something you can actually manage.

Subscribe to our blog to get notified about our latest articles.