We can build you the truck. What does the gas cost? It is the fairest question a buyer asks about AI, and most vendors answer it with a shrug. The honest answer is that the gas depends on the commute: two miles a day, or to Canada and back every week. But the commute is not a mystery. It is a set of design choices, and design choices can be measured.
Below are three real tasks from Pisteyo's own work, measured from session logs and priced at Anthropic's published list rates on 1 September 2026. Then comes the one number worth measuring, and the eight dials that move it.
Why is the token price the wrong number to watch?
Because it expires, and it was never the part you control. AI cost has two layers, one that keeps and one that goes stale, and mixing them up is how budgets get caught out.
Ignore the token price. It expires. Rates change every few months, so any figure you remember is already out of date, ours included. On Anthropic's price sheet today, the newest Opus model lists 20% below the Opus these examples were measured on, and the newest Fable model re-reads cached context for a quarter of what its predecessor charged. The price can also hold while the bill moves: Anthropic notes that its Claude 4.7 and later models use a newer tokenizer that “produces approximately 30% more tokens for the same text,” with the exact increase depending on the content (Anthropic pricing documentation, checked 29 September 2026).
Learn how it is built. It lasts. How an agent is built sets the bill, and that knowledge does not go stale. One agentic task is typically ten to twenty model calls, each one re-reading what came before, and the retries, not the first answer, are what drive the cost.
What are you actually paying for when you buy tokens?
Every request bills across three layers, and they are priced very differently.
| Layer | What it is | How it bills | Example: 500 words on Claude Opus 5 |
|---|---|---|---|
| Input | Your prompt, the whole conversation so far, any files, and the definition of every connected tool | Cheapest per token, and it builds up quietly because it is re-sent every turn | About $0.003. 500 words is about 670 tokens: a third of a cent. |
| Reasoning | The model's thinking before it answers | Invisible to you, and billed at the output rate | About $0.05 (moderate thinking): roughly 2,000 thinking tokens. |
| Output | The answer you read | Priciest per token, and usually the smallest of the three | About $0.02. A 500-word answer is about 670 tokens at the output rate. |
Estimated prices at Anthropic's list rates as of 29 September 2026: $5 per million input tokens and $25 per million output tokens for Claude Opus 5. In a month, these numbers will be out of date.
Typing 500 words cost a third of a cent. The same 500 words coming back cost five times as much, and moderate thinking in between cost about fifteen times as much. The whole exchange, 500 words in and 500 out, costs about 7 cents. So say everything you need to say, out loud if that is faster. Save by matching the thinking and the answer to the job, not by trimming your prompt.
Reasoning is the layer that surprises people. A short answer can sit on top of a long internal monologue you never see and still pay for, like the kitchen time on a restaurant bill. Input is where agent bills quietly grow, because an agent re-reads its instructions, its tools and its history every time it takes a step.
What are the three kinds of token spend, and what should you do with each?
Every token an organization buys lands in one of three kinds. The framework comes from Nufar Gaspar on The AI Daily Brief, and it turns “our AI bill is too high” into a to-do list. Here is each kind, with a real task from our own work.
Spin: attack this first
Spin is activity with no output. Jobs that run more often than the data changes, reports nobody opens, agents idling on a schedule. It is the one kind of spend to cut on sight.
Measured: a nightly briefing agent
$675 a year
on automated briefings nobody reads
- One run
- 3.5 million tokens on Claude Opus 5, $5.39
- Per year
- $1,350 across 250 runs
- The waste
- $675, if half the briefings go unread
Nothing broke. The agent did exactly what it was built to do, on schedule. That is the lesson of spin: a working automation can turn into waste without a single error. The fix is rarely a cheaper model. It is running the job only when the data changes, and switching off the ones nobody reads.
Produce: tune this
Produce is work that ships. The deliverable, the analysis, the code. Do not cut it. The job is to make each accepted output cost less.
Measured: a board-level strategy deck
$21
for a deck that replaced about $6,500 of people time
- One build
- 16.4 million tokens on Claude Opus 5
- Per year
- $43 across 2 builds
- People time
- 2 to 3 people for a week, at $65 an hour
This is where fear of the bill does the most damage. A team that watches AI spend on a dashboard and flinches starts rationing the work that pays. Tune it instead: the right model for each step, a strong first draft, and cheap edits after it.
Teach: protect this
Teach is the tries it takes to get a new workflow working. Experiments, failed approaches, the same task run three ways to learn which one holds. On a dashboard it looks like pure waste, because nothing shipped. This is tuition.
Measured: debugging a new AI workflow
$16
118 attempts built a workflow the business runs daily
- One build
- 13.0 million tokens on Claude Fable 5
- Per year
- $32 across 2 retunes
- What it left
- A workflow that has run every day since
Cutting teach spend saves a rounding error and gives up the learning curve. Protect it, and give the people doing it room to try.
The order matters, so say it the same way every time: kill the spin, tune the production, protect the teaching. It is what makes a cost conversation feel like engineering instead of austerity.
What is the one number to measure?
Cost per accepted task, not price per token. Price per token is the sticker. Cost per accepted task is what one output a person actually accepts really costs, with every retry and every minute of review counted.
accepted task =
Run the deck from above through it. The tokens were $21. Say a senior reviewer spends two hours checking it: at $65 an hour, that is $130. The accepted deck really cost about $151, and the review cost six times the tokens. It is still about 2% of the $6,500 of people time it replaced.
Two lessons fall out of that arithmetic. First, a cheap model that needs three tries costs more than a premium one that lands it once, and once review time is counted it is not close. Second, count only the review time the AI adds. If someone would have checked that deck anyway, those minutes are not the AI's cost, and a good CFO will ask.
If a build cannot report its own cost per accepted task, it is not finished.
What are the eight dials that set the bill?
These are the design choices that move a bill by an order of magnitude. Audit any workflow, team or vendor build against them.
- Model fits the task. Match model strength to the task, not its price. Test on five to ten real examples before you commit, because the instinct about which model is cheaper is often backwards.
- Budget the loop. Cap the iterations and define what done looks like, so the work can stop on its own. Most of an agentic task's cost lands after the first answer, in checking and redoing.
- Pay setup tax once. Every run re-sends the agent's setup, its standing instructions and tool definitions, before anyone types a word. Keep it lean and cache what does not change: on Anthropic's current price sheet, a cached read costs a tenth of the normal input price or less.
- Filter at the source. Retrieve twenty rows, not five hundred. Anthropic sizes an average 10 kB web page at about 2,500 tokens and a 500 kB research PDF at about 125,000, so handing an agent the right page instead of the whole paper is a fifty-to-one difference. Precision costs less and answers better.
- Scope the session. New task, new session. History is re-sent on every turn, so a marathon thread costs more with every message, and usually gets worse, not better.
- Right-size thinking. The biggest single dial. Reasoning bills at the output rate, and overthinking a simple question is both pricier and worse. Set the effort per step, not once for everything.
- Instrument the build. Log tokens per run and per accepted output, and sample the minutes people spend reviewing. If a build cannot report its own cost, it is not done.
- Fit the kill switch. Everything on a schedule gets a spend cap, an alert when cost jumps, and a review date. Nufar Gaspar has described an always-on assistant of her own that ran up a four-figure bill in two weeks while she was travelling and not using it at all, reading hundreds of millions of tokens and writing almost nothing, because of scheduled jobs it had created for itself. Autonomous systems fail silently, and upward.
Take the one-page version
The whole framework fits on one page: the three kinds of spend, the formula and the eight dials. Download the one-pager (PDF) and keep it next to the budget.
Framework credit: the three kinds of token spend come from “Everything You Need to Know About AI Tokens,” The AI Daily Brief, 2 August 2026, with Nufar Gaspar. Pisteyo's examples are measured from real session logs and priced at Anthropic's published list rates on 1 September 2026, including prompt-caching rates. Each covers one working session, so read them as a close order of magnitude rather than an invoice. Human time is valued at $65 an hour.
