How should IT delivery adapt to reap agentic AI benefits?

Tech leaders need to rethink the service delivery model
By Xavier Boileau, Charles de Pommerol, Samir Javeri, Zineb El Honsali, Maxime Happy, and Ramiro de la Rosa
Home  // . //  //  How should IT delivery adapt to reap agentic AI benefits?

AI agents can now act across the whole IT delivery chain, from drafting requirements to operating production systems. For technology leaders, the opportunity is no longer simply to accelerate software development. It is to rethink how the entire delivery chain works, where humans remain accountable, and where agentic AI can create the most value.

In The AI Frontier — Where AI Delivery Meets Human Oversight, we examine agentic AI across three stages of the IT delivery chain: intent, covering requirements and clarity (do we know what to build?); build, covering engineering and verification (can we build it reliably?); and run, covering delivery, operations, and control (can agents ship and operate under control?). At each stage, we provide our perspectives on:

  • What should be delegated to an agent
  • What should be kept in human hands
  • How IT can authenticate, after the fact, actions that were taken

Our analysis suggests that the biggest gains are to be had toward the end of the IT delivery chain, particularly in continuous integration and delivery (CI/CD) and run activities.

The IT function in organizations faces a panoply of new challenges with agentic AI. While code generation is becoming cheaper and faster, the scarce capabilities are moving elsewhere: knowing precisely what to build, verifying that it works, and controlling how it operates. For CEOs, CIOs, CTOs, and boards, that changes where to invest and where to keep people firmly accountable.

Intent — As software gets cheaper to build, what to build matters more

Requirements still move at human speed, while code generation moves much faster. Our research shows that investment is already moving upstream. Some 39% of CIOs put data restructuring first among investment priorities, while 33% put requirements quality and governance first; only 3% rank review and test capacity as their top investment priority. More than 70% of stated priorities go to understanding what to build and structuring the data needed to build it.

Exhibit: CIO investment priorities in past 12 months
Source: The Oliver Wyman 2026 Parallel Surveys of CEOs and Tech Leaders on Agentic AI and IT, July 2026.

AI can increasingly do the work of building software. As engineering capacity becomes cheaper and more abundant, the quality of the requirements becomes more valuable. This changes how the teams operate, asking for clearer and more detailed requirements from the start before handing work over to AI for code generation.

At enterprise scale, this is a governance problem. Hundreds of requests may come from different teams and increasingly be drafted with AI. Each can appear reasonable on its own, while unknowingly conflicting with others or leaving gaps that cause agents to build the wrong things at the high speed agentic AI enables.

Technology can identify some of those contradictions, but it needs a shared semantic layer — in other words, agreed definitions for concepts such as customer, active account, and what constitutes a successful outcome.

The harder conflicts are about priorities. Tools cannot resolve them; only clear decision rights can. Each domain needs a named owner, with executive arbitration where priorities collide. The brief becomes a governed, versioned asset, while final sign-off rests with the business.

Four risks affect the front end:

  • Individually credible briefs that conflict with one another
  • AI-generated requirements arriving faster than the business can validate them
  • Definitions drifting as different agents work from different vocabularies
  • Decision queues turning cheap engineering into expensive idle resources

Organizations that know exactly what they want are better positioned to capture the gains from agentic AI. Those that focus only on building faster can create a runaway train of AI-generated outputs that causes confusion and prevents IT departments from focusing on needed work.

Build — As code generation accelerates, verification becomes the bottleneck

The engineering bottleneck is moving from generation to verification. In our 2026 Parallel Surveys of CEOs and Tech Leaders on Agentic AI and IT, we found that 84% of respondents reported productivity gains of up to 20% in their own pipeline after applying agentic AI. Specifically, 27% said that code generation is already outrunning their review and test capacity. A further 29% are still measuring success by commit volume rather than shipped releases.1 The build phase is where agentic AI is most visible and where its productivity gains are easiest to overcount.

Writing code is becoming a commodity. Turning it into reliable, shipped software that delivers the practical value it was intended to in the first place remains complex. That should change the investment equation, prompting organizations to focus more on design and shipping phases rather than code generation itself.

To offset this propensity, organizations should put more capacity into what remains scarce: domain knowledge, review capacity, and enforced quality gates. Give agents work that can be verified automatically, and keep the rest with people. Make verification continuous and tooled rather than relying on reviewers to remain permanently vigilant. A model that is right most of the time can gradually teach its reviewers to stop checking.

Without strong gates, AI productivity gains can quietly become maintenance and security debt. Most CIO respondents estimate that 1% to 15% of AI-assisted defects survive into production, while 29% still measure engineering throughput only at the commit or code-volume level, meaning the true rate may be much higher and may go undetected.

There is a practical warning sign: If code-generation volume rises while review capacity and test gates stay flat, the organization is manufacturing technical debt. Rollout should pause until the controls catch up.

Leaders also must remember to measure productivity on their own delivery pipeline. Published gains describe somebody else's context and somebody else’s success or failure.

Run — Agentic AI can operate in production without controlling it

Agentic capabilities are already emerging in CI/CD on major platforms, with agents moving into operations. The useful design line runs between two planes. The data plane is localized, reversible work: generating a patch, rerunning a test, and restarting a service. The control plane determines the rules of production, including pipeline configuration, deployment policies, and approval gates.

In our work, we have found that a clear distinction between the data plane and control plane is a useful way to define where agentic autonomy should stop.

The data plane’s operations can increasingly be delegated to agents with human oversight. In our CIO survey, 55% of respondents already restrict agents in their CI/CD pipeline to data-plane actions only, and a further 33% allow control-plane actions solely where every decision is human-owned and recorded.

The control plane should remain under governed authority, with human approval, explicit recourse, and a tamper-evident record of who authorized what. Yet only 25% of CIOs and CTOs can reconstruct what an agent changed, when, and who authorized it in a clear audit trail. Seventy-two percent are working with only partial logs.

Organizations also need the infrastructure to support this: a paved-road internal developer platform and agent-ready, observable infrastructure with per-agent identity, telemetry, rollback, and a kill switch. Blast radius must be contained by design because failures can propagate at agentic speed.

The governance test is straightforward. If an agent can change a deployment policy without a human-owned control and a recorded decision, or if the organization cannot reconstruct which agent changed what and under whose authority, control-plane delegation should remain off the table.

Site Reliability Engineering discipline should also extend to the agents themselves. As routine work becomes automated, organizations need to guard against on-call engineers becoming less likely to question the agent.

Agentic AI autonomy requires controls across IT delivery

The market is already choosing bounded autonomy. Almost eight out of 10 CEOs favor preset limits in sensitive processes. More than 90% of CIOs and CTOs place current production autonomy between “human-assisted” and “human-supervised,” and 96% set it workflow by workflow.2 The same pattern holds across requirements, engineering, delivery, and operations.

IT should transfer authority to agents in explicit, bounded steps and back each step with verification the organization can run automatically and evidence it can replay afterward.

Enterprises should measure delivery stability, not adoption, and track patch-to-release latency as a first-class key performance indicator. The controls at each stage should act as a gate for the next stage of autonomy.

The question is no longer whether agents can act across the IT delivery chain. We know they can. The leadership decision now is how far to let them go, step-by-step, across the value chain.

 

Authors