AI and Transformation
Key takeaways
- Sequence beats ambition: govern, prove, industrialise, scale, with a gate between each stage and a named owner at every gate.
- Write the governance rules before the first pilot, not after it; 67% of middle-market firms already apply controls at or before the pilot stage.
- Start in the back office, where the cost being cut already sits in the ledger and attribution is straightforward.
- Data quality and integration, not model choice, are what stop pilots reaching production; budget most of the schedule for them.
- Fund AI as a portfolio with a stop date on every proof of value, because most experiments will close and that is the plan working.
What is an AI transformation roadmap, and how does it differ from a pilot?
An AI transformation roadmap is a sequenced plan that carries a company from governed experiment to production capability, with a defined gate between each stage. A pilot asks whether a tool works. A roadmap asks what has to be true, and in what order, for that tool to move a cost line or a cycle time and keep moving it after the vendor demo ends. The difference is sequence and accountability, not ambition.
Most mid-market firms already hold the tools. 86% of middle-market firms have partially or fully integrated AI into operations, and 97% report satisfaction with their AI investments3. Access to technology is not the constraint. Ordering the work is.
Two definitions hold the rest of this together. AI transformation is the sustained redesign of how work gets done using AI, spanning workflow, data, governance and talent, as distinct from isolated tool adoption. A proof of value is a scoped, time-boxed deployment against a single measurable business outcome, used as the gate before anything scales past its first team.
| Stage | Question it answers | Gate to pass |
|---|---|---|
| 1. Govern | What may we do with which data? | Written data, risk and review rules with a named owner |
| 2. Prove | Does one workflow get measurably better? | One metric moved against a pre-agreed baseline |
| 3. Industrialise | Can that workflow survive real data and real systems? | Integration and data quality fixed, monitoring live |
| 4. Scale | Which initiatives earn more capital, and which stop? | Unit economics finance will sign, plus a retirement list |
Why do most AI pilots never reach the P&L?
Because the failure is organisational, not technical. MIT’s Project NANDA found that 95% of enterprise generative AI pilots deliver no measurable P&L impact, set against $30 to $40 billion of enterprise generative AI spending2. The report puts the divide down to approach rather than model quality or regulation. Firms buy the capability and leave the surrounding workflow untouched, so nothing downstream changes.
The adoption curve and the value curve have come apart. 88% of organisations now use AI in at least one function, up from 78% a year earlier1. About a third have scaled it across the business, and only 6% count as high performers capturing outsized value1. The other 94% use AI without transforming1.
Read that gap as a sequencing problem. Adoption is cheap and fast, needing a licence and a champion. Value needs a workflow redesigned around the model, a data path that holds under load, and a finance owner who agreed the baseline before the work started. Programmes that jump straight to the interesting use case tend to discover the missing pieces at the exact moment they try to scale.
Stage one: set the governance gate before the first pilot
Govern first. In the middle market, 67% of firms apply governance controls before the pilot or production stage rather than after the fact3, and that ordering is the cheapest risk reduction on the table. Gartner predicts that organisations which operationalise AI transparency, trust and security will see a 50% improvement in AI adoption, business goals and user acceptance5.
Governance at this stage should stay thin. Four artefacts are enough for most mid-market firms: a data classification stating which categories may reach a third-party model, a review step for anything customer-facing or regulated, a logging and monitoring standard, and one named accountable owner. Write them before the first pilot. Retrofitting them later costs a rebuild, and it is the rebuild, not the review, that kills momentum.
Governance is not the brake on an AI programme. It is what lets you say yes in a week, because the answer to “can this data go there” is already written down.
Stage two: one workflow, one number, one quarter
Narrow scope wins early. Only 17% of middle-market firms pursue transformational enterprise-wide AI initiatives, while 45% target clear near-term value instead3. The second group is better placed to show a return, because a single workflow has a single owner, a single baseline, and a visible before and after.
Start in the back office. MIT’s analysis points to back-office automation as the highest-return entry point, because it reduces outsourced spend and cuts costs that already sit in the ledger2. Invoice coding, claims triage, contract abstraction, reconciliation and tier-one support deflection all qualify: high volume, structured enough to measure, low brand risk. Customer-facing generative AI is more visible and far harder to attribute. The same logic governs automation inside SaaS operations, where the win that survives scrutiny is a queue that clears faster, not a new interface.
Then time-box it. Agree the metric and its baseline with finance before build starts, hold scope to one team, and set the date on which the proof of value either graduates or stops. A proof of value with no stop date becomes a permanent pilot, and permanent pilots are how the 95% figure gets made.
Stage three: fix data and integration before you scale
The barriers to scaling are known, and they are unglamorous. Middle-market firms name data quality (53%), integration challenges (47%), unclear ROI (33%) and security or compliance (33%) as what stops pilots reaching production3.
“AI-ready” is not a platform purchase. In practice it means four things: every record the model reads has an owner and a defined refresh cadence; systems of record expose data through an interface rather than a nightly export; access follows the classification written in stage one; and outputs are logged so quality can be measured over time rather than argued about. This work is dull, it consumes most of the schedule, and it is the reason a second use case takes a fraction of the time the first one did. It is also where outside AI consulting support tends to pay for itself, because the binding constraint is integration experience rather than model selection.
Stage four: scale what survives, retire what does not
Treat stage four as portfolio management. Deloitte reports that more than two-thirds of enterprises expect 30% or fewer of their AI experiments to be fully scaled within three to six months4. Plan for that ratio instead of against it. Fund several proofs, expect most to close, and move the capital to the ones that cleared their gate.
Agentic systems belong at this stage, not earlier. 62% of organisations are experimenting with AI agents, and 23% report scaling agentic AI somewhere in the enterprise1. Agents multiply the value of a clean process and multiply the damage of a broken one, so they depend on the governance and data work already being finished. A staged AI transformation programme exists to make that dependency explicit up front.
Who should own AI transformation, and how is it funded?
Ownership works best as a thin centre with delivery in the business. The centre owns the rules, the shared platform and the model inventory. The business unit owns the workflow, the metric and the change management, because that is where the process actually lives. IT owns integration, identity and monitoring. When a centre of excellence owns outcomes as well as standards it becomes the bottleneck; when nobody owns standards, every unit renegotiates the same data question from scratch.
Money is not the scarce input. 58% of middle-market firms plan to invest $1 million or more in AI this fiscal year, and 84% expect their AI spending to rise next year3. Across enterprises generally, 78% expect to increase AI spending in the next fiscal year4. With that much capital in motion, the differentiator is a measurement standard finance will accept: a baseline captured before deployment, one primary metric per initiative, cost per unit of work including inference and licence spend, and a quarterly review that can stop an initiative as readily as extend it.
How is AI transformation different for mid-market firms?
The advantage is time to production. Shorter chains of command let a mid-market firm move a concept into production faster than a large enterprise can convene its steering committee. That speed is the asset the roadmap should exploit, which is the argument against importing large-enterprise governance wholesale. Copy the gates, not the committee count.
The matching risk is concentration: fewer engineers, thinner data teams, and more exposure to one vendor’s roadmap and pricing. The mitigations are practical rather than architectural. Keep prompts, evaluation sets and workflow logic in your own repository. Prefer interfaces you can point at a different model. Price the switching cost before signing and revisit it at every renewal, because lock-in usually arrives through data formats rather than through the contract.
Run the four stages in order and the roadmap stops being a slide and starts being a decision record: what was allowed, what was proved, what was fixed, what was funded. If you want a second read on the sequence for your own operation, talk to our team.
Frequently asked questions
How long does an AI transformation take for a mid-market company?
Plan it in quarters per stage rather than as one end date. Deloitte's 2026 enterprise survey found more than two-thirds of enterprises expect 30% or fewer of their AI experiments to be fully scaled within three to six months, so a realistic programme runs a governance stage, then a time-boxed proof of value, then integration work, with scale decisions made at each gate. The variable that moves the timeline most is data readiness, not model selection.
What governance do we need before scaling AI past pilots?
Four artefacts cover most mid-market situations: a data classification stating which categories of data may reach a third-party model, a review step for anything customer-facing or regulated, a logging and monitoring standard so output quality can be measured, and one named accountable owner. This is deliberately lighter than large-enterprise AI governance. Put it in place before the first pilot, because retrofitting controls onto a live deployment usually means rebuilding it.
Why do most generative AI pilots fail to show ROI?
MIT's Project NANDA reported in 2025 that 95% of enterprise generative AI pilots deliver no measurable profit and loss impact, and attributed the gap to approach rather than model quality or regulation. The common pattern is buying a capability without redesigning the workflow around it, so the tool works and the cost line does not move. The second common cause is no agreed baseline, which makes any result arguable after the fact.
Who should own AI transformation: IT, a centre of excellence, or the business units?
Split it. A thin centre owns the rules, the shared platform and the model inventory. The business unit owns the workflow, the metric and the change management, because that is where the process lives. IT owns integration, identity and monitoring. A centre of excellence that owns outcomes as well as standards becomes a queue, and no owner of standards means every unit renegotiates the same data questions.
Which AI use cases should come first?
High-volume back-office processes with an existing cost baseline: invoice coding, claims triage, contract abstraction, reconciliation, tier-one support deflection. MIT's 2025 analysis identified back-office automation as the highest-return entry point because it reduces outsourced spend and cuts costs already visible in the ledger. Customer-facing generative AI is more visible but much harder to attribute, which makes it a poor first proof point.
How do we measure AI ROI so that finance accepts it?
Agree the measurement standard before build starts. Capture a baseline for the target metric before deployment, name one primary metric per initiative, calculate cost per unit of work including inference and licence spend, and hold a quarterly review with the authority to stop an initiative as well as extend it. Anything measured only after launch will be contested, and contested numbers do not survive a budget cycle.
Sources
- McKinsey: The State of AI, 2025. mckinsey.com
- MIT Project NANDA: The GenAI Divide (reported by Legal.io), 2025. legal.io
- RSM US: Middle Market AI Survey, 1,030 respondents, 2026. prnewswire.com
- Deloitte: State of AI in the Enterprise, 2026. deloitte.com
- Gartner: AI Governance and AI TRiSM, 2026. gartner.com




