Your Containment Rate Can't See the Agent Who Did Nothing

A containment rate can rise for two reasons: customers got what they needed without a person, or they got a fluent reply, a closed ticket, and no help at all. On the dashboard those are the same event, and the second one is cheaper to produce.

That has always been true as a matter of logic. What has changed is which of the two is now more likely to go unnoticed.

In CRM magazine's February look-ahead, Stefan Ostwald of Parloa predicted that AI agents "will increasingly become self-optimizing systems, continuously improving their own performance, reliability, and behaviors without explicit human intervention." He is probably right, and predictions pieces are optimistic by genre, so this is not a complaint. But read all the way down that page and one question never comes up. If the agent is monitoring whether it achieved the goal, what establishes that it did?

Contact center measurement is built almost entirely on the conversation. Containment and deflection count how a conversation ended. Average handle time measures how long it took. Quality monitoring scores what was said inside it,  including tone, disclosures, and adherence. Customer satisfaction asks the minority of customers who answer surveys how the conversation felt.

The customer's actual problem is almost never in the conversation. It is in a different system: the refund posted to billing, the address changed on the account, the order canceled in the order management system. For as long as contact centers have existed, those two things have been held together by a person. A rep who said "I've taken care of that for you" had to go into the other system and take care of it. If the rep didn't, the customer called back, and the next rep opened the record and saw an account that didn't match the promise.

Autonomous agents pull the conversation and the action apart. The conversation is now produced by the same system that reports on it.

The number rewards the failure it can't see.

Picture two agents handling the same contact. The first cannot complete the task, says so, and hands the customer to a person. The second closes the conversation with a confident summary of what it has done, and does not do it.

The second agent wins on containment. It wins on handle time. Quality monitoring, reading a transcript in which it was courteous, on-brand, and fully compliant, has nothing to flag. The customer who was not helped either calls back, arriving weeks later as a repeat contact, often logged under a different intent, or does not call back at all, which is the outcome your instrumentation is least able to see and the one that costs the most.

This is not a measurement bug. It is measurement working exactly as designed, on a system for which it was not designed. Every one of those instruments was built for a process that failed in front of a human being who would have noticed.

The most useful benchmark I know of for service agents is tau-bench, built by the research team at Sierra. Its two domains are retail and airline customer service, and the design decision worth borrowing is what it refuses to score. Not the transcript. The evaluation "compares the database state at the end of a conversation with the annotated goal state" — the world, not the account of it. A customer-service AI company, building an evaluation for customer-service agents, concluded that what the agent says happened is not evidence that it happened.

They also introduced a metric worth asking vendors for by name: pass^k, the probability that an agent completes the same task on all k attempts rather than on at least one. Their scores are stale, and I want to be straight about that—2024 models, and the frontier has moved. The metric has not. Success collapsed as k rose to under 25 percent at pass^8 in retail. Every demo you have ever been shown is pass^1.

As for what an agent does when it cannot finish: researchers at Carnegie Mellon put agents to work inside a simulated company and found that "when the agent is not clear what the next steps should be, it sometimes tries to be clever and create fake 'shortcuts' that omit the hard part of the task." Their example belongs in every CX planning deck. An agent who could not find the person it needed on the company chat platform renamed a different user to that person's name and carried on. Nothing errored. The record showed contact made. Those were 2024 models too; the behavior, not the score, is the finding.

Measure completion, not containment.

The fix is unglamorous and mostly organizational. For each of the handful of intents your agents own, write down the state change that means the customer was actually served—the credit posted, the field updated, the order canceled. Then reconcile a sample of closed conversations against that system on a schedule, using a check the agent does not run and cannot see. Treat "conversation closed, nothing changed" as an exception worth someone's attention rather than a clean containment. And report the two numbers side by side, because the gap between what was contained and what was completed is the number you are actually managing.

None of that requires new technology. It requires deciding that the agent is not the witness to its own work.

Containment counts the conversations that never reached you. It is the one number in the contact center that goes up when nothing happens at all.


Chase W. Hughes is a three-time founder—Pro Business Plans, Equity Up, and ProAI, one of the first commercialized GPT products, which he grew to 300,000 users and seven figures in about 18 months, bootstrapped, before selling it. He has spent 15 years in AI and works with companies building AI products.