Note: I’m traveling extensively for work at the moment, so publishing across AI Unfiltered, Intelligent Founder AI and QuantOpinion has been a little or more irregular. I don’t schedule articles in advance, so thank you for your patience when there’s a long gap between posts.
This week, The Wall Street Journal explored why businesses are struggling to budget for AI. One example captured the problem: a more expensive model completed a video-analysis task for roughly $1, while a cheaper model consumed about $14 and still failed. The price of each step mattered less than how many steps the system took, and whether they led anywhere.
But there is another question behind that bill:
who decided the next attempt was worth paying for?
When we ask an AI agent to “check this properly” or “keep trying,” we want persistence. We do not want it to give up at the first missing document or failed search. Yet those instructions leave something important unsaid: how far should it go before coming back to us?
Imagine asking an agent to compare three suppliers. It finds conflicting specifications, searches again, then tries another route. That extra work might resolve the uncertainty. It might also leave the agent exactly where it started, with a larger bill.
The issue is not that agents should never retry. It is that permission to pursue an outcome should not automatically mean permission to keep spending.
A useful agent needs more than a task and access to tools. It needs a boundary:
when to continue,
when to return a partial answer, and
when to ask whether another attempt is worth it.
The cheap model that kept going
The Journal described a video-analysis test using two Google models.
The more expensive model completed the assignment in 85 steps for approximately $1. The cheaper model took 952 steps, consumed approximately $14 worth of tokens and still failed. According to the podcast, the unsuccessful run encountered repeated obstacles and even consulted another chatbot for help.
This was an illustrative test, but its not a proof that expensive models always offer better value. it does however, exposes the weakness of choosing by unit price alone.
The Journal also discussed research published earlier this year involving Stanford, Carnegie Mellon, UC Berkeley and Microsoft. Across more than 6,800 maths, programming and science tasks, the cheaper model incurred higher overall costs in 32% of tested cases.
This week’s coverage brought those earlier findings into the budgeting debate; it was not a newly released study.
There is a straightforward lesson here. You pay for the work a model performs, including work that does not produce a useful result.
A lower-priced model can be the right choice. It can also become an expensive way to discover that the task was beyond it.
A prompt is becoming a process
Consider an illustrative request:
“Check these suppliers, compare their specifications and tell me which three meet our requirements.”
It sounds like one task. An agent might break it into searches, document retrieval, extraction, comparison and verification. A missing specification could trigger another search. Conflicting information could trigger a second check. An inaccessible document could send it down a different route.
Some of that additional work improves the answer. Some may simply repeat an unsuccessful approach.
The user sees one request and, eventually, one response. The system has performed a sequence of billable operations in between.
That is why comparing model prices is only the beginning. The relevant questions include how much work the system does, whether it succeeds and how much checking the result still needs.
The budgeting conversation also extends beyond model selection. On October 9, The Economic Times’ CIO publication reported McKinsey’s warning that expanding usage and multi-step agent workflows could outweigh falling AI costs. Its coverage highlighted the importance of linking spending to measurable outcomes.
Cheaper access makes more uses possible. It does not make those uses equally valuable.
What does “keep trying” mean?
The difficult part is that persistence can look like a feature.
An agent that gives up immediately is not particularly useful. We want it to recover from errors, find alternatives and check its work. Those are precisely the behaviors that make a demonstration feel less like a chatbot and more like an assistant.
But persistence needs a boundary.
Return to the supplier example. If the agent cannot verify a certification, should it search five more sources? Try a different model? Ask the user for the missing document? Return a qualified answer?
Each option involves a different balance of cost, confidence and delay.
The question is no longer just whether the agent can continue. It is whether continuing is worth it- and who gets to make that decision.
A person given this assignment might eventually say, “I can investigate further, but it will take another hour.” The comparable moment for an agent is an escalation point: here is what remains unresolved, here is what another attempt would involve, and here is the decision needed.
Without that moment, “be thorough” can become an open-ended instruction.
A budget needs a boundary
There are two separate questions businesses need to ask.
What did the system spend?
What is it allowed to spend next?
The first is about visibility. The second is about authority.
A record of expenditure helps explain a bill. It does not necessarily prevent the next unsuccessful attempt. A warning can be useful, but its value depends on someone or something, being able to act before more work proceeds.
This is where the “stop button” becomes more interesting than another cheaper-model announcement.
It need not mean abruptly cancelling every task at an arbitrary threshold. It could mean returning a partial result, requesting approval or switching from autonomous execution to human review.
Those choices also have consequences. Stop too early and useful work remains unfinished. Allow unlimited retries and a modest request can become a costly investigation.
The boundary should reflect the value of the task, not simply the availability of more tokens.
The bill is part of the product
It would be easy to read this week’s reporting as another warning that AI is too expensive. That misses the more useful point.
An expensive task can be worthwhile if it delivers something valuable. A cheap task can be wasteful if it produces an answer nobody can trust.
The harder problem is predictability: understanding what the system is likely to consume, what completion means and what happens when it cannot get there, and, WSJ’s reporting makes that problem concrete.
Businesses are struggling to forecast expenditure, and cheaper models can sometimes consume more money while accomplishing less. Neither finding means AI cannot deliver value. Both challenge the idea that a lower token price settles the economics
For agents, the spending boundary belongs in the product design alongside the instructions and permissions. “Complete this task” needs a companion rule for what happens when completion remains uncertain.
AI may keep getting cheaper to call. The more consequential question is whether we are getting better at deciding when it should stop.





