Three levers can cut enterprise AI costs
A coding task can now carry a surprisingly different price tag depending on how the AI system is arranged around the model. In late July, Sarvam AI said its Sarvam Code agent averaged about $2 per task on Terminal-Bench 2.1, compared with $4.1 to $27.8 for Claude Code and OpenAI’s Codex. The difference, Sarvam said, came partly from how the system split work between planner, worker and verifier agents and routed it between models.
That surrounding software is known as an agent harness: it decides what information an agent receives, how a task is divided and how often the model is called. Kausal Malladi, chief technology officer for investments at INDmoney, said a long task can repeatedly resend earlier conversation, instructions and context. Caching — reusing information the model has already processed — can reduce effective cost by roughly 80%, he estimated.
The second lever is model selection. Melento, a document automation company formerly known as SignDesk, routes routine requests to smaller, faster models and sends genuinely complex cases to frontier models. Founder and CEO Krupesh Bhat’s rule is simple: match the price of the model to the difficulty of the task. Enterprises can track the share handled by smaller models, average cost per transaction and task success rate to see whether savings are real rather than merely shifting work elsewhere.
The third lever is more radical: stop treating every step as an AI problem. Arjun Nagulapally, CTO at AIONOS, said a telecom project processing more than a million interactions a month moved roughly 30% of routine steps to ordinary software, including databases and business rules. AIONOS said the cost of resolving each interaction fell by 35%, with the largest saving coming from work that no longer required a model.
So what changes in practice? Teams can use the same AI budget to complete more work by reducing repeated context, reserving expensive models for difficult decisions and removing unnecessary model calls. But higher throughput is not automatically lower spending: the company must also see lower costs for models, people or rework while keeping success rates stable. The reported figures come from the companies themselves.
Comments
Loading the thread…
Sign in to leave a comment. Sign in