Insights By Uday Birajdar, CEO & Co-founder, AutomationEdge
When I talk to CIOs, and to the MSP and GSI leaders who serve them, one of the concerns they raise most often is what tokens will cost, and how they can be sure they won’t end up spending more than they have budgeted for AI projects.
Our field team ran into a version of this at an insurer earlier this year. They were masking IDs in historic policy documents, and because an ID could appear anywhere in a file, the whole document had to go through the model. The prompts had not been tuned. Token consumption came in at roughly ten times the estimate, and the process was pulled out of production within a week. A second workload in underwriting went the same way.
None of that is a service desk workload, of course. It is an insurance back office, two hops away. And yes, the prompts could have been tuned, and the ten times would have come down. That is rather the point. A consumption estimate is a guess until it meets production, and the meter bills you for the guess being wrong.
In April, Uber’s CTO told The Information that the company’s entire 2026 AI budget was gone. Just four months in.
I think you should read these stories as a preview, because the same meter is arriving underneath the service desk, and a service provider is structurally more exposed to it than Uber was.
The shift isn’t a rumour, it’s the roadmap
Forrester puts a number on the discomfort: 30% of enterprise SaaS decision-makers already flag unpredictable usage-based pricing as a problem. CIO Dive, covering the same shift at the end of August, framed the core issue cleanly — with token-metered spend, buyers often don’t know their true costs until the bill arrives.
The important detail is what the meter counts. It isn’t a subscription. Its work attempted: invocations, actions, credits, assists, whatever the vendor calls them. And a reasoning workflow burns a lot of them, where a deterministic workflow burns one or none.
Which is why the usual reassurance doesn’t reassure. Yes, model prices fall every quarter. That doesn’t touch the number of invocations your design spends per ticket, and it certainly doesn’t help when a failed attempt is billable and a retry loop bills again.
Why service providers are exposed differently
If an enterprise IT team blows through its budget in month four, that’s a variance. Someone explains it to the CFO, the forecast gets rebuilt, and life continues. Uber is doing exactly that.
If you do it, it comes out of the gross margin. You sold the seat at a fixed price. On most fixed-fee contracts, you can’t reprice until renewal. And because you run one stack across many clients, a design that over-consumes doesn’t over-consume in one place. It does it everywhere at once, in proportion to how well you’re selling.
Sit with that shape for a second. A variable cost underneath a fixed revenue line means growth increases your exposure. That is the same trap as adding L1 headcount to serve new seats, except that this time it never shows up as headcount.
For a GSI the same arithmetic runs over a longer term and a larger base. A managed services engagement priced as fixed TCV over five years locks the revenue line for the whole term, and the productivity commitments written into it were modelled on effort and offshore mix, not on per-invocation AI cost. Unless the contract carries a change-of-technology clause, there is no renewal conversation to correct that until year five. And because the same delivery architecture gets reused across accounts, a consumption assumption that was wrong in one pursuit is wrong in every account it was reused in.
There is a second edge for services firms. AI-led resolution is now a win theme in every large pursuit, which means committing to productivity you then have to deliver against your own effort-based revenue. That is a hard enough trade to make deliberately. It gets harder when the unit cost of the resolution layer moves with every retry.
The shape that fits
The alternative isn’t less AI. It’s AI where the uncertainty actually is.
The front door. Intake is genuinely hard for rules. Users describe symptoms instead of causes and bury three requests in one message. Understanding what someone wants and routing it correctly is exactly what language models are good at, and it costs one invocation per ticket.
The core. Known request to known workflow, executed deterministically. Same input, same path, same result, with a log you can hand a client. Little or no meter.
The tail. Ambiguous, compound, or novel tickets, plus incident diagnosis where nobody knows the cause yet. Here a reasoning agent earns its cost, because there’s real uncertainty to reason about — and the comparison is against human escalation, which costs you a multiple of an L1 ticket.
Cost is the wedge, not the whole argument
If the economics were the only reason, this piece would age badly. The durable reasons don’t depend on anyone’s price list.
Blast radius. You act inside your clients’ own systems, under credentials they gave you. A non-deterministic path through a privileged action carries a different risk when the same automation runs across fifty client environments than when one enterprise runs it in its own. Every MSP leader I talk to says a version of the same thing: they don’t need AI that’s clever, they need AI that won’t do something stupid in a client’s environment.
Latency. A deterministic workflow completes while the user is still in the chat. A multi-step reasoning loop often doesn’t. On your highest-volume ticket types, that’s the difference between a resolution and a ticket.
Auditability. When a client asks what happened, “here is the workflow that ran” beats a reconstructed chain of thought.
Where this argument is weakest
Two places, and I’d rather name them than let you find them.
Deterministic automation has a bad history. It’s what RPA / ITPA promised, and it broke every time a UI changed, a client’s stack drifted, or the inputs changed. That’s fair, and it’s the honest reason agents got popular. My answer is that the old model failed at intake and at the brittle edges, not in the execution core, which is exactly the seam this design cuts along. But rules rot if nobody maintains them, and no architecture saves you from that.
And most obviously: we price our product, SupportFlo, per resolved ticket, which is a meter too. I’d argue an outcome meter behaves differently from an attempt meter — you pay once per resolution however many steps it took, and nothing when it fails.
The question for your next vendor meeting
Ask which layer each capability runs in, and what it consumes per ticket at each layer. Then ask what happens on a retry.
You’ll learn more from how long the answer takes than from the answer itself.
Sources: The Information via Forbes on Uber’s 2026 AI budget overrun; Forrester, “Agentic ERP Won’t Scale Until CIOs Control The Proof, The Price, And The Portability,” August 2026; CIO Dive, “Agentic AI is shifting the pricing models CIOs rely on,” August 2026.

Uday Birajdar,
Co-founder & Chief Executive Officer at AutomationEdge
Uday Birajdar is Co-founder and CEO of AutomationEdge, with 20+ years of experience building and scaling enterprise technology businesses. He leads the company’s product vision around AI and agentic automation, helping enterprises move from task automation to autonomous resolution across IT, HR, Finance, and healthcare operations. Under his leadership, AutomationEdge has evolved into a global enterprise software company focused on agentic AI, automation, and intelligent execution.