AI and Transformation
Key takeaways
- Agents pay off now in bounded work: document extraction, reconciliation, ticket triage and exception routing, not open-ended departmental autonomy.
- Most agent failures are process failures, because an agent inherits every ambiguity the process never resolved and cannot improvise the way a person quietly did.
- Keep a named human on the exception path; full autonomy is the fastest route to a cancelled project in any regulated workflow.
- Hyperautomation spend and agent cancellations are rising together, because enterprises are funding integration and orchestration while pulling back on standalone agent bets.
- Buy on demonstrated multi-step reasoning and tool use rather than on the word agent: Gartner counts only around 130 vendors of thousands with real agentic capability.
Where does AI business process automation actually work today?
AI business process automation works today where the task is narrow, the inputs are messy but bounded, and a wrong answer is cheap to catch. Document extraction, invoice and ledger reconciliation, ticket triage, claims pre-screening and exception routing are all running in production somewhere right now. What does not work is the open-ended remit: an agent handed a department and told to run it. The binding constraint is rarely model capability. It is whether the process was ever properly defined before the agent arrived.
The adoption data shows the shape of the problem. Some 88% of organisations report regular AI use in at least one business function, up from 78% a year earlier1. Only 23% are scaling an agentic system somewhere in the enterprise, with a further 39% still experimenting1. Inside any single business function, no more than 10% report scaling agents at all1. A separate count puts 31% of enterprises with at least one agent in production, rising to 47% in banking and insurance2.
Adoption is wide and shallow. Almost every large company has an agent somewhere. Almost none have agents running a function.
Why do most agentic AI pilots never reach production?
A pilot proves the model can do the task. It does not prove the organisation can run the task. MIT’s NANDA study, covering 300 deployments, 153 leader surveys and 52 executive interviews, found that 95% of enterprise generative AI pilots deliver no measurable P&L impact, with only 5% creating significant value4. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls3.
The investment posture underneath those cancellations is more cautious than the headlines imply. In a Gartner poll of 3,412 webinar attendees in early 2025, 19% had made a significant agentic AI investment, 42% a conservative one, 8% none at all, and 31% were still in wait-and-see mode3.
The failure mode is consistent and it is not exotic. Agents fail hardest where the underlying process was never standardised. A clerk handling a non-conforming invoice makes a judgment call, applies an unwritten rule learned from a colleague, and moves on. Nobody records the rule. Hand that process to an agent and the ambiguity does not disappear. It gets executed thousands of times a day, consistently, at speed, and wrongly.
An agent does not resolve an undefined process. It performs the ambiguity at scale.
What separates AI business process automation, RPA and hyperautomation?
The terms get used loosely, and the difference matters as soon as you scope a budget against them.
| Term | What it does | Where it breaks |
|---|---|---|
| RPA | Executes fixed, rule-based steps across systems, exactly as scripted. | Any input that deviates from the script. |
| Intelligent automation | Combines RPA with AI and machine learning so processes with variability or decision points can run without a human at every step. | Decisions needing context outside the training data. |
| AI business process automation | Uses AI models, usually alongside rules, to execute, monitor or improve multi-step workflows involving judgment and unstructured data. | Open-ended remits and unowned exception paths. |
| Hyperautomation | An enterprise strategy: identify and automate as many processes as possible using a coordinated stack of RPA, AI, process mining and orchestration. | Integration and governance debt, not model quality. |
Hyperautomation is not a synonym for agentic AI, and the two are not competitors. Hyperautomation is the operating stack. Agents are one component that plugs into it. Gartner reports that 90% of large enterprises now treat hyperautomation as standard operating practice, with the enabling software market projected at 1.04 trillion dollars by 20265.
That resolves an apparent contradiction. Hyperautomation spend keeps rising while headline agent projects get cancelled, because enterprises are funding the unglamorous plumbing, integration, data pipelines and orchestration, while pulling back on speculative standalone agent bets. The plumbing is what agents need in order to work at all. The same logic governs automation inside SaaS operations, where the tooling is rarely the constraint.
Which processes are safe to hand to an agent today?
Four tests, applied before any tooling decision. A process that fails two of them is not an agent candidate this year.
Is the output verifiable?
A reconciliation either balances or it does not. A summarised customer sentiment does not check itself. Start where correctness is machine-checkable, because that is what lets an agent run at volume without a reviewer attached to every transaction.
Is the exception path owned?
Every real process has a tail of cases the rules do not cover. Before deployment, name the person who receives them, the queue they land in and the service level they carry. Projects that skip this step do not fail loudly. They quietly build a backlog nobody is watching.
Is the decision reversible?
Drafting a response, proposing a journal entry, flagging a claim: all reversible. Posting the entry, paying the claim, closing the account: not reversible without cost. Reversible steps can run with sampled review. Irreversible ones need a checkpoint.
Is the volume worth the governance?
An agent carries fixed overhead: monitoring, evaluation sets, access control, audit logging and model updates. Below a certain transaction volume that overhead exceeds the staff time it displaces, which is much of why costs escalate on work scoped from demo economics.
One further filter applies to vendors rather than processes. Gartner estimates that of the thousands of vendors describing themselves as agentic, only around 130 offer genuine agentic capability, with the remainder amounting to agent washing over existing chatbots and scripted automation3. Test for real multi-step reasoning and tool use against your own data and your own edge cases before signing anything.
How much human oversight does a regulated workflow need?
Enough to own every decision that leaves the building. In finance, healthcare and insurance, the pattern that survives contact with an audit is an agent that does the work and a named human who approves the exceptions, not an agent acting freely with logging bolted on afterwards. Banking and insurance lead adoption at 47% with at least one agent in production2, which is not evidence that regulated firms are relaxed about this. It is evidence that they already have the process documentation, the controls function and the audit habits agents depend on.
Scope is the other half of it. The reported difference between the 5% of pilots that created significant value and the 95% that did not was tight scope and iterative feedback loops, not a larger rollout4. Full autonomy is the fastest route to a cancelled project in a regulated workflow, because the first serious exception becomes a compliance event rather than a ticket.
How do you measure an AI business process automation project beyond the demo?
Four numbers, tracked from the first week in production rather than from the pilot deck:
- Straight-through rate. The share of transactions completed with no human touch. This is the number that turns into money.
- Exception rate and its trend. A flat exception rate after six weeks means the agent is not absorbing your edge cases and somebody else is.
- Fully loaded cost per completed transaction. Inference, licences, review time and engineering maintenance, not model cost alone.
- Rework and reversal rate. How often a completed item is corrected downstream. Automation that moves work later in the chain has not removed it.
Demo accuracy is not on that list. Neither is user satisfaction with the interface. On timing, expect the first honest read when the exception queue stabilises, not at go-live. Internal process automation is a live use case for 48% of organisations deploying agents, against 60% for data analysis and reporting7, and the gap is instructive: reporting work is judged on output quality, process work is judged on what happens to the queue behind it. Agreeing that measurement frame before the build is the part of an AI consulting engagement that pays for itself.
Where should a team start with AI business process automation?
Start with process definition, not procurement. The sequence that works is unromantic:
- Map one process end to end, including the informal workarounds nobody ever documented.
- Write down the rules those workarounds encode, then fix the ones that were never really rules.
- Instrument the baseline: volume, cycle time, exception rate, cost per transaction.
- Automate the deterministic steps with rules. Give the agent only the steps that genuinely need judgment.
- Ship to a narrow slice with a human checkpoint on the exception path, and widen only when the exception rate stabilises.
There is a deadline on this work that most teams have not priced in. Gartner expects 40% of enterprise applications to include task-specific AI agents by the end of 2026, up from under 5% in 20256. Most organisations will therefore acquire agents by software upgrade rather than by deliberate project. Agents arriving inside tools you already run will inherit whatever ambiguity your processes still contain, and they will act on it before anyone has approved a business case.
The work that pays off is the work that has always paid off in AI transformation: define the process, own the exceptions, measure the queue. The agents are the easy part. For a second opinion on which of your processes clear those four tests, talk to our team.
Frequently asked questions
Can AI agents replace RPA and BPM platforms, or do they work alongside them?
They work alongside them, and treating agents as a replacement is a common scoping error. RPA executes fixed rules reliably and cheaply, which an agent does not do better or more predictably. The pattern that holds in production is rules for the deterministic steps and an agent only for the steps that need judgment or unstructured input, with both coordinated by the same orchestration layer.
Which business processes are safe to hand to an AI agent today?
Processes where the output is machine-checkable, the exception path has a named owner, the decision is reversible, and the transaction volume justifies the governance overhead. Document extraction, invoice and ledger reconciliation, ticket triage and claims pre-screening usually clear all four tests. Open-ended work, such as running a function end to end, clears none of them yet.
Why do most agentic AI pilots fail to reach production?
A pilot tests whether the model can do the task, not whether the organisation can run the task. MIT's NANDA study found that 95% of enterprise generative AI pilots deliver no measurable profit and loss impact. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.
What is the difference between AI business process automation, RPA and hyperautomation?
RPA runs fixed, scripted steps exactly as written. Intelligent automation adds AI and machine learning so processes with variability or decision points can run without a human at every step. AI business process automation uses AI models, usually alongside rules, to execute, monitor or improve multi-step workflows involving judgment and unstructured data. Hyperautomation is the enterprise strategy that coordinates all of these plus process mining and orchestration.
How much human oversight does an AI agent need in a regulated workflow?
Enough that a named person approves every exception and every irreversible action. In finance, healthcare and insurance the durable pattern is an agent that does the work with a human checkpoint on the exception path, not an agent acting freely with logging added afterwards. Tight scope with iterative feedback loops consistently outperforms broad autonomous rollouts.
What is a realistic ROI timeline for AI-driven process automation?
Expect the first honest read once the exception queue stabilises in production, not at go-live, because the early weeks understate the manual effort being absorbed elsewhere. Track straight-through rate, exception rate, fully loaded cost per completed transaction and downstream rework from the first week live. Demo accuracy and interface satisfaction are not ROI signals.
Sources
- McKinsey & Company: The State of AI, 2025. mckinsey.com
- S&P Global Market Intelligence: enterprise AI agent adoption data (via Digital Applied), 2026. digitalapplied.com
- Gartner: Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, 2025. gartner.com
- MIT Project NANDA: The GenAI Divide (reported by Legal.io), 2025. legal.io
- Gartner: hyperautomation adoption and enabling software market data (via Zion Market Research), 2026. zionmarketresearch.com
- Gartner: task-specific AI agents in enterprise applications (via Panto), 2026. getpanto.ai
- Azumo: AI agent statistics, aggregated industry data, 2026. azumo.com




