Insights By Uday Birajdar, CEO & Co-founder, AutomationEdge
When I talk to CIOs, and to the MSP and GSI leaders who serve them, one of the concerns they raise most often is what tokens will cost, and how they can be sure they won’t end up spending more than they have budgeted for AI projects.
I run an automation software company, so I have a position in this. Here it is. If you deliver IT support under a price you have already fixed, the way most AI support tools are being designed right now puts a variable meter on the most predictable work in your queue. That is a margin problem before it is a technology problem, and the fix is architectural rather than commercial.
What this looks like in production
Our field team ran into a version of it at an insurer earlier this year. They were masking IDs in historic policy documents, and because an ID could appear anywhere in a file, the whole document had to go through the model. The prompts had not been tuned. Token consumption came in at roughly ten times the estimate, and the process was pulled out of production within a week. A second workload in underwriting went the same way.
None of that is a service desk workload, of course. It is an insurance back office, two hops away. And yes, the prompts could have been tuned, and the ten times would have come down. That is rather the point. A consumption estimate is a guess until it meets production, and the meter bills you for the guess being wrong.
It happens at scale, too. In April, Uber’s CTO told The Information that the company’s entire 2026 AI budget was gone. Just four months in.
The meter counts attempts
Forrester reports that 30% of enterprise SaaS decision-makers already flag unpredictable usage-based pricing as a problem. CIO Dive, covering the same shift in August, framed the core issue cleanly: with token-metered spend, buyers often don’t know their true costs until the bill arrives.
The important detail is what the meter counts. Not a subscription. Work attempted — invocations, actions, credits, assists, whatever the vendor calls them. A reasoning workflow burns a lot of them where a deterministic workflow burns one or none, and on most contracts a failed attempt is billable and a retry bills again.
Which is why the usual reassurance doesn’t reassure. Model prices fall, but that doesn’t change how many invocations your design spends per ticket.
Why this lands differently on a service provider
If an enterprise IT team blows through its budget in month four, that is a variance. Someone explains it to the CFO, the forecast gets rebuilt, life continues. Uber is doing exactly that.
If you deliver the service, it comes out of gross margin. Beyond that, the exposure looks different depending on which side of the market you sit on.
If you run an MSP. You sold the seat at a fixed price, and on most fixed-fee contracts you can’t reprice until renewal. Because you run one stack across many clients, a design that over-consumes doesn’t over-consume in one place. It does it everywhere at once, in proportion to how well you’re selling. A variable cost underneath a fixed revenue line means growth increases your exposure, which is the same trap as adding L1 headcount to serve new seats, except that this time it never shows up as headcount.
If you run a GSI managed services P&L. The same arithmetic runs over a longer term and a larger base. A five-year fixed TCV locks the revenue line for the entire term, and the productivity commitments written into it were modeled on effort and offshore mix, not on per-invocation AI cost. Unless the contract carries a change-of-technology clause, there is no renewal conversation to correct that until year five. And because the same delivery architecture gets reused across accounts, a consumption assumption that was wrong in one pursuit is wrong in every account that inherited it.
There is a second edge if you sit on the GSI side. AI-led resolution is now a win theme in every large pursuit, which means committing to productivity you then have to deliver against your own effort-based revenue. That is a hard enough trade to make deliberately. It gets harder when the unit cost of the resolution layer moves with every retry.
The shape that fits
The alternative isn’t less AI. It’s AI where the uncertainty actually is.
The front door. Intake is genuinely hard for rules. Users describe symptoms instead of causes and bury three requests in one message. Understanding what someone wants and routing it correctly is exactly what language models are good at.
The core. Known request to known workflow, executed deterministically. Same input, same path, same result, with a log you can hand a client. Little or no meter.
The tail. Ambiguous, compound, or novel tickets, plus incident diagnosis and remediation. Here a reasoning agent earns its cost, because there is real uncertainty to reason about, and the comparison is against human escalation, which costs you a multiple of an L1 ticket.
Deterministic in the core, intelligence at the edges. The design question stops being how agentic you are and becomes which layer a given ticket belongs in.
Cost is the wedge, not the whole argument
If the economics were the only reason, this piece would age badly. The durable reasons don’t depend on anyone’s price list.
Blast radius. You act inside your clients’ own systems, under credentials they gave you. For an MSP, a non-deterministic path through a privileged action carries a different risk when the same automation runs across fifty client environments than when one enterprise runs it in its own. For a GSI, the equivalent is one delivery architecture inherited across accounts and regions, where a design flaw doesn’t stay in the account that created it.
Auditability. When a client asks what happened, “here is the workflow that ran, under this identity, with this result” beats a reconstructed chain of thought. On a regulated account it is the difference between an answer and an incident.
Where this argument is weakest
Deterministic automation has a bad history. It’s what RPA promised, and it broke every time a UI changed, a client’s stack drifted, or the inputs changed. That’s fair, and it’s the honest reason agents got popular. My answer is that the old model failed at intake and at the brittle edges, not in the execution core, which is exactly the seam this design cuts along. But rules rot if nobody maintains them, and no architecture saves you from that.
And most obviously: we price our product, SupportFlo, per resolved ticket at L1, which is a meter too. I’d argue an outcome meter behaves differently from an attempt meter — you pay once per resolution, however many steps it took, and nothing when it fails — but a buyer should interrogate my pricing as hard as anyone else’s.
What we build, since you should know
SupportFlo is our IT support product for service providers, and it is built the way this piece argues. AI at intake, deterministic execution in the core against the systems where the work actually lives, a reasoning agent held back for the tail. We price it per resolved ticket rather than per attempt, and we will model consumption against your own ticket mix before you sign anything.
If you’re working through these numbers for your own desk, I’m glad to compare notes, whether or not there turns out to be anything in it for us.
The question for your next vendor meeting
Ask which layer each capability runs in, and what it consumes per ticket at each layer. Then ask what happens on a retry.
You’ll learn more from how long the answer takes than from the answer itself.
Sources: The Information via Forbes on Uber’s 2026 AI budget overrun; Forrester, “Agentic ERP Won’t Scale Until CIOs Control The Proof, The Price, And The Portability,” August 2026; CIO Dive, “Agentic AI is shifting the pricing models CIOs rely on,” August 2026.
Uday Birajdar,
Co-founder & Chief Executive Officer at AutomationEdge
Uday Birajdar is Co-founder and CEO of AutomationEdge, with 20+ years of experience building and scaling enterprise technology businesses. He leads the company’s product vision around AI and agentic automation, helping enterprises move from task automation to autonomous resolution across IT, HR, Finance, and healthcare operations. Under his leadership, AutomationEdge has evolved into a global enterprise software company focused on agentic AI, automation, and intelligent execution.