The Work Behind a Larger AI Agent Rollout
Companies expanding agent use need a way to handle exceptions and decide when a workflow is ready for more responsibility.
Forty percent of respondents at companies with more than $1 billion in annual revenue told McKinsey that their organizations were scaling AI agents in at least one function. Among respondents at smaller organizations, 22 percent said their organizations were scaling AI agents in at least one function. McKinsey published the findings on Aug 25, 2026.
McKinsey collected responses from 1,719 participants in 97 countries between May 4, 2026 and Jun 8, 2026. The figures describe reported adoption, and the survey does not establish what caused the gap between company sizes.
A team planning an expansion still has to decide what the agent may do when the usual process breaks down. A successful demonstration can leave that question unanswered because its designers choose the inputs and watch the result. Daily operations put the system in front of cases its designers did not choose.
The operating plan deserves as much attention as the agent itself. For people who do AI work, that plan creates a place for domain knowledge long before anyone needs to train a model.
A defined task gives the team something to test
Anthropic distinguishes workflows, which follow predefined paths, from agents, which choose their next steps and tools. Its engineering guidance recommends starting with the simplest approach that meets the task. A company can adopt an agent for part of a process while keeping other steps fixed.
Consider a hypothetical distributor using an agent to investigate delayed shipments. The agent could read the order record and carrier updates, then prepare a response for the service team. The operations manager would define which records count as evidence and which customer requests fall outside this initial assignment.
That scope leaves useful work for the model without giving it authority over every decision in a service case. A customer asking for a tracking update presents a different problem from a customer demanding a replacement at a new address. The team should write separate rules for those requests before testing either one.
A domain expert can turn that distinction into examples that engineers can implement. One example might show a routine delay with consistent records, while another shows a carrier update that conflicts with the warehouse system. The expert should specify what the agent may conclude from each record and when it must stop making promises.
The team also needs a clear ending for the assignment. Preparing a supported reply is one outcome, while sending the reply is another. A manager who combines them in a single completion measure cannot see whether the agent found the right information but made an unauthorized commitment.
Permissions belong in the connected systems
OWASP’s guidance on excessive agency recommends limiting tool permissions and enforcing authorization in downstream systems. A model’s instruction to stay within its role does not replace those controls. Engineers should give the shipment agent only the access its current assignment requires.
For the distributor’s first rollout, an engineer could permit the agent to retrieve shipment records and save proposed replies. The engineer would leave replacement orders outside that permission set. If the business later adds replacements, the operations manager would first define eligibility and the engineer would enforce it at the order system.
This separation makes a policy change visible. A manager who expands the agent’s authority must change a specific permission and supply evidence for the new task. The team can then test the added action without treating a better conversational answer as proof that the agent should control more of the business.
Someone also needs to own access when the workflow changes. If the distributor retires a carrier connection, its administrator should remove that connection from the agent’s available tools. The evaluation team should confirm that the agent can handle the resulting cases through the approved alternative.
An exception needs a destination
The instruction to escalate a difficult case leaves several operational choices open. In the shipment example, the business must decide who receives a conflict between order records and who can authorize a customer remedy. The agent cannot resolve a missing decision owner by writing a longer explanation.
A service lead could require each handoff to include the conflicting records and the precise decision the agent could not make. The recipient should be able to continue the case without repeating the entire investigation. If the system cannot supply that evidence, the team should count the handoff as incomplete.
The following example assigns responsibility within the hypothetical distributor. Each organization would need to adapt the roles to its own process.
| Case condition | Agent action | Responsible person |
|---|---|---|
| Carrier and order records agree | The agent prepares a reply with the supporting records. | A service reviewer checks the proposed response during the initial rollout. |
| Records conflict | The agent pauses the reply and passes along both records. | A service lead resolves the conflict or assigns an investigation. |
| A customer requests a replacement | The agent records the request without creating an order. | An authorized operations employee decides under the existing replacement policy. |
| The order system times out | The agent records the uncertain result and avoids an automatic repeat. | An operations employee checks whether the action took effect. |
The timeout case deserves its own test. A failed response from a tool does not necessarily tell the team whether the underlying action happened. If a later version of the agent can place replacement orders, the application needs a way to check the order state before it tries again.
Managers should include this exception work in staffing plans. A rollout that sends ambiguous cases to an unstaffed queue has moved the delay to a different screen. The service lead should watch how long customers wait after a handoff and whether recipients have the authority to finish the case.
Testing should establish what actually happened
Anthropic’s guidance on agent evaluations separates the agent’s conversation from the outcome in the surrounding system. Its example contrasts a booking claim with a reservation that actually exists. The same distinction applies when a service agent says it has resolved a shipment problem.
In the distributor example, an evaluator should inspect the saved reply and check whether its delivery estimate follows the available evidence. If the task includes an order change, the evaluator should verify the resulting record. A fluent message cannot establish that the correct customer received the correct action.
The evaluator should keep a set of realistic cases that the team can rerun after a change. That set should include ordinary requests alongside the specific exceptions the business expects. The reviewer can then explain which failures prevent expansion, instead of relying on a single average score.
Some checks permit a definite answer, such as whether the agent accessed the correct order. Others require judgment, such as whether the response gives the customer a misleading impression of certainty. A domain expert should write that judgment rule and supply examples of acceptable wording.
The project manager should also reserve cases that the developers have not used to adjust the agent. Repeated success on familiar examples gives the team less information about a new customer request. A separate test set gives the reviewer a better basis for challenging the proposed expansion.
Expansion needs a measure of the whole service
For this rollout, the operations team could measure the share of eligible cases that reach an acceptable resolution without later repair. It should keep a separate count of cases the agent hands back to people. Otherwise, the team could mistake an increase in easy assignments for an improvement in the service.
The operations team also needs to define and track the eligible case pool. If the team removes difficult requests from the eligible pool, it should record the change and show the excluded volume. Managers can then distinguish a narrower assignment from a more capable system.
Operations leaders should compare waiting time and repair work with the earlier process, using comparable types of cases. A pilot staffed by an unusually experienced service lead may perform differently when other employees take over. The rollout record should identify that staffing difference before anyone attributes the result to the agent.
A manager can expand one part of the assignment at a time, such as adding a carrier while leaving replacement decisions with employees. That choice creates a specific question for the next evaluation. The team can examine whether the new carrier’s records support reliable replies without also changing who may issue replacements.
NIST’s report on monitoring deployed AI systems explains why testing before release cannot capture all the conditions of actual use. The report draws on practitioner workshops and a literature review, and it describes monitoring methods as still developing. Teams therefore need evidence from operation as well as a passing test before launch.
For the distributor, a change in carrier data could justify pausing the affected cases while the rest of the service continues. The operations manager should define who can make that decision and how employees will resume the manual process. A pause should leave customers with a working route to resolution.
Domain expertise becomes part of the operating process
An expert entering AI work can use a small, realistic workflow to demonstrate this kind of judgment. A portfolio example could pair a set of difficult shipment cases with explanations of the correct next action. The author should show the evidence behind each decision and distinguish invented examples from actual client work.
The most revealing example may be a case where a plausible answer would lead to the wrong action. An evaluator who explains that failure gives an engineer something concrete to test. An operations specialist who specifies the right handoff gives the service team a way to recover.
People considering an AI evaluation assignment should ask whether the work includes access to the evidence needed to judge outcomes. They should also find out how the client resolves disagreements about the rubric. Without those arrangements, a reviewer may spend time scoring answers that the client has never defined well enough to assess.
A larger rollout changes who depends on the agent’s decisions. The people responsible for that expansion need evidence that the work finishes correctly, including the cases that return to employees. Domain experts can help supply that evidence by making the unresolved decisions visible and giving the team a workable way to handle them.
Sources
BIJ2 AI reviewed these original sources on October 8, 2026. The notes below identify the source type, survey disclosures, commercial interests, and material limits used in this article.
- McKinsey’s 2026 survey. McKinsey conducted and published the research and does not identify a separate funder. McKinsey also sells AI consulting.
- Anthropic’s workflow guidance. Anthropic published this vendor engineering article on Dec 19, 2024. It supplies the workflow and agent distinction used here.
- OWASP’s Excessive Agency guidance. OWASP’s 2025 edition supplies the recommendations on permissions and authorization. The page presents security guidance rather than a deployment survey.
- Anthropic’s agent evaluation guidance. Anthropic published this vendor engineering article on Jan 9, 2026. It explains the difference between a transcript and the resulting system state.
- NIST’s monitoring report. NIST released AI 800-4 in March 2026 and draws on workshops and a literature review. The report identifies monitoring challenges rather than estimating agent adoption.