Quick Answer: Agentic AI projects in India stall between pilot and payback for ordinary, fixable reasons. KPMG in India found 88% of enterprises surveyed are investing in agentic AI, yet published returns are rare, and Gartner predicts that over 40% of agentic AI projects will be cancelled by the end of 2027. The usual culprits are no measured baseline, a pilot run on clean data, integration debt, nobody owning exceptions, and governance added late. None of them is a model problem. Measure cost per completed task against a baseline, plus work finished without human touch, handoffs, cycle time, error cost, review minutes and compute cost.

An agent demo tends to stop just before the part that decides the return. The agent reads the form and books the refund, and nobody asks what happens when the form is a skewed phone photo or the refund goes to the wrong account. Those cases are where the money is made or lost. The operations head wants to know who picks up the work the agent cannot finish. The board just wants a number.

The published Indian evidence is thin, and it says more about spending than about returns.

What the Indian figures say, and what they leave out

KPMG in India reported in July 2026 that 88% of Indian enterprises surveyed are investing in agentic AI. The figure tells you money is moving, and nothing about whether any of it has come back.

An EY-CII report from November 2025 found that 47% of Indian enterprises have multiple AI use cases live in production. Live in production is a higher bar. It covers AI in general, though, and ‘live’ is not the same as ‘profitable’. Subtracting one survey from the other is tempting and wrong: they asked different questions, and only one of them is about agentic AI.

Then there is the warning. On 25 June 2025, Gartner predicted that over 40% of agentic AI projects will be cancelled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls. It is a general forecast rather than an Indian count. Its three reasons matter more than the percentage, because each one can be tested before signing.

Why agentic AI ROI is harder to prove than chatbot ROI

A chatbot answers questions; its value shows up in deflected calls, and a wrong answer usually costs a follow-up. An agent is given a goal and permission to act. It reads a document, checks a system, updates a record and, in some deployments, moves money. That changes the calculation in five ways.

  • Multi-step work. A task that passes through several steps can fail at any of them. Step-by-step accuracy can look excellent while the share of tasks finished end to end stays lower, because small error rates compound.
  • Exceptions. Real work arrives with missing fields, blurred scans and customers who change their minds mid-call. Agents handle the typical case well. The exceptions decide the economics.
  • Human handoffs. Every case passed to a person carries a cost the business case rarely shows: the reviewer must rebuild context the agent already had. A clumsy handoff can make a case slower than doing it by hand.
  • The cost of a wrong action. A chatbot’s mistake is a bad sentence. An agent’s can be a wrong refund, a duplicate payment instruction or a collections call to the wrong borrower. One such error can wipe out the savings from many correct tasks.
  • Compute cost at volume. Agents call a model several times per task, to plan, check and retry. A cost invisible in a small pilot becomes a real line item in production, growing with task complexity as well as volume.

The chatbot-era instruments, containment rate and satisfaction scores, measure conversations. Agents have to be judged on completed work.

Five places agentic projects stall between pilot and scale

The stall points are predictable. Each one shows up late and is caused early, when the pilot is designed.

A pilot run on clean data

Pilots tend to be fed a curated sample: complete records, recent documents, well-behaved customers. Production gets everything else. When an agent meets real data its completion rate often falls, and that is a data problem before it is a model problem. Our checklist for AI-ready data covers the self-audit to run first.

Integration debt

In a sandbox, an agent reads a copy of the data and writes to nowhere. Production is harder. There it must read from the core system, write back to it, respect batch windows and leave an audit trail, often on a platform older than the team integrating with it. That work frequently outweighs the agent itself and rarely appears in the pilot budget.

Governance arriving late

Risk, compliance and audit teams often first see the agent when it is ready to go live. Their questions are reasonable. What can it do, on whose authority, how is each action logged, and how is a wrong one reversed? If the answers were not designed in, go-live waits while they are retrofitted. Gartner’s third reason, inadequate risk controls, tends to live here.

No owner for exceptions

When the agent cannot finish a case, somebody has to. In the pilot, that is a motivated project team. In production it is an operations desk that was never consulted, has its own targets and treats handed-off cases as extra work. Unless someone owns that queue, exceptions pile up quietly. The first alarm is often a customer complaint.

No measured baseline

This one comes last because it only bites at the payback review. Pilots get approved to prove the agent works, so nobody measures what the current process costs. Months later the agent completes a share of tasks at some cost, and there is nothing honest to compare it with. Finance, reasonably, will not accept a salary-table estimate in its place.

Where agentic AI projects stall between pilot and payback A path from pilot through go-live and scale to the payback review, with five stall points marked where they become visible: a pilot run on clean data, integration debt, governance arriving late, no owner for exceptions, and no measured baseline. Each has a fix that belongs before the pilot starts. Where agentic projects stall on the way to payback Each stall shows up late in the project. Each one is caused when the pilot is designed. PILOT GO-LIVE SCALE PAYBACK REVIEW 1 2 3 4 5 STALL 1 Pilot on clean data STALL 2 Integration debt STALL 3 Governance arrives late STALL 4 No owner for exceptions STALL 5 No measured baseline FIX EARLY Pilot on real, messy records from day one FIX EARLY Budget the core system work into the pilot FIX EARLY Bring risk and audit into the first design FIX EARLY Name the desk that owns the handoff queue FIX EARLY Measure today’s process before the agent starts Each marker shows where a problem becomes visible. Every fix belongs before the pilot starts, when changing the design costs a meeting rather than a rebuild.

What Mahindra Finance and Sarvam disclosed, and what they did not

A useful Indian example arrived at Global Fintech Fest on 10 September 2026, where Mahindra Finance and Sarvam said their voice AI had handled more than 1 crore calls in 12 languages across sales, collections and employee engagement. That is a serious operating number. A system that has run across a crore of calls has met real customers at a volume no pilot reaches.

The announcement did not include conversion, cost or return figures. That is ordinary. Companies seldom publish unit economics for a system that may give them an edge. Nothing in the announcement suggests a problem.

For a buyer, the example shows that voice AI can run at very large volume, in 12 languages, across more than one business function. Whether the same approach would pay back on your own collections desk depends on the figures that were not shared, so you will have to produce them yourself.

Why NPCI’s agent work means governance cannot wait

On 12 September 2026, TechNode, citing Reuters, reported that NPCI is building an AI-agent registry and a protocol for agent-initiated UPI payments, starting with small-value purchases such as groceries. Details may change before anything launches. The direction is still worth planning around.

A registry implies each agent will need an identity. A protocol implies rules on what an agent may do and how that is recorded, and starting small suggests the risk is being contained while those rules settle. Agents are heading into regulated rails, where the controls come first.

Outside payments too, any agent that can change a record or commit money should be able to answer four questions from its own logs: which agent acted, on whose authority, what exactly it did, and how the action can be undone. Designing those answers in costs little, while retrofitting them at go-live is slow and expensive. For lenders and insurers, our piece on AI agents in Indian BFSI goes further.

A seven-metric scorecard for agentic AI ROI

The measures that decide agentic ROI are unglamorous, and none needs anyone else’s benchmark. Each needs a baseline taken before the agent touches the process, over a period that includes your normal peaks, such as month-end or the festival-season rush.

  1. Baseline cost per completed task. What one finished unit of work costs today, with staff time, supervision and rework loaded in. Time-sample the current process and divide by tasks completed, not started. Every other measure is read against this one.
  2. Completion without human touch. The share of tasks the agent finishes end to end with nobody stepping in, taken from its logs and matched against the workflow system so that ‘done’ means done downstream. It tells you whether you bought automation or an expensive assistant.
  3. Exception and handoff rate. How often the agent passes work to a person, and why. Tag every handoff with a reason code from day one. The reasons matter more than the rate: they show whether the fix is better data, a missing integration or a rule nobody wrote down.
  4. Cycle time. Elapsed time from arrival to completion across every case, including those that went to a person. Agents are quick on easy cases, and an average can hide a slow tail in someone’s queue. Report the spread alongside the mean.
  5. Cost of errors. Count wrong actions and price each one from reversal and complaint records. Audit a sample of ‘successful’ tasks too, because some errors surface only when someone checks.
  6. Human review minutes. Time people spend checking agent output, including spot checks on work marked complete, from queue timestamps or a simple time log. Business cases often leave it out.
  7. Compute cost per task. Model and infrastructure spend divided by completed tasks, not by calls or tokens. Tag cloud and model usage by use case so the bill can be split, and track it monthly as volume grows.

Put together, the ROI test is arithmetic. Add compute, platform run costs, review minutes at a loaded rate, exception handling and the priced cost of errors, then divide by tasks completed. Set that against the baseline. If it is lower and error cost is flat or falling, scale; if not, the handoff reason codes tell you what to fix first.

That is the agentic AI ROI India Inc can take to a board: a cost per completed task, before and after, with the mistakes priced in. It will rarely be as dramatic as a vendor slide, and it will survive the audit committee.

Frequently asked questions

What ROI are Indian companies getting from agentic AI?

Very little has been published. KPMG in India found 88% of enterprises surveyed are investing in agentic AI, and EY-CII found 47% have multiple AI use cases live in production, but neither is a return figure. Mahindra Finance and Sarvam say their voice AI has handled more than 1 crore calls, with conversion, cost and return figures not disclosed. The dependable figure is the one you measure yourself.

Why are so many agentic AI projects expected to be cancelled?

Gartner predicted in June 2025 that over 40% of agentic AI projects will be cancelled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls. In practice those show up as compute costs growing with volume, pilots with no measured baseline, and governance added at go-live. Each can be checked before a pilot is approved.

How is measuring an AI agent different from measuring a chatbot?

A chatbot is judged on conversations, such as questions answered and calls deflected. An agent takes actions across several steps and systems, so it has to be judged on completed work. The useful measures are cost per completed task against a baseline, completion without human touch, handoff rate, cycle time, the priced cost of errors, review minutes and compute cost per task.

What should the baseline for an agentic AI pilot include?

The fully loaded cost of completing one unit of the work today, including staff time, supervision and rework, divided by tasks completed rather than started. Measure it before the agent is introduced, over a period that includes normal peaks such as month-end. Without it the pilot has no pass mark, and every review becomes an argument about estimates.

What does NPCI’s reported AI-agent work mean for enterprises?

TechNode, citing Reuters, reported on 12 September 2026 that NPCI is building an AI-agent registry and a protocol for agent-initiated UPI payments, starting with small-value purchases such as groceries. It signals that agents are heading into regulated payment rails, where identity, permissions and audit trails are likely to be required. Enterprises should design those controls in from the start.

Should we wait for Indian ROI benchmarks before investing in agentic AI?

No. Published Indian benchmarks for agentic AI returns are scarce, and imported ones rarely match Indian costs or processes. A better route is a tightly scoped pilot on a process whose current cost you can measure, with cost per task, handoffs, errors and compute captured from day one, and a pass mark agreed in advance. That becomes your benchmark.

WinInfoSoft is ISO 9001:2015 and ISO 27001 certified and assessed at CMMI Level 3, and works with Indian enterprises on AI automation, integration and application development. If you are approving an agentic AI pilot and want the baseline and scorecard designed before the build begins, get in touch.