Khamas Innovations — Management Consulting, Innovation & Project Delivery

Pages Insights Blog Approach Workshops & Events Wind MENA Careers

What AI agents actually change about the consulting engagement

2026·AI & Adoption·9 min read

What AI agents actually change about the consulting engagement

By Khamas Innovations · Published 7 May 2026 · Updated 27 Aug 2026

The AI agent conversation has reached a familiar inflection point. Twelve months ago everyone was running pilots. Today everyone has at least one production deployment, and a fair number are quietly being walked back. The useful questions have moved on from whether agents work to where they earn their cost, and where they create a category of risk nobody has priced. There is enough deployed practice in the open now to reason from rather than speculate about.

What an agent is in practice

The vendor pitch describes a closed system that perceives, decides and acts on its own. What most enterprises actually run is closer to a managed assistant with permission to call three to five tools, working inside a guardrail that flags anything outside a narrow set of approved patterns. The gap between those two descriptions is where most failure modes live, which is reason enough not to let the brochure set expectations.

A workable definition: an LLM with the authority to invoke tools, take actions in external systems, and choose its own next step within a bounded scope. Almost everything about the risk profile follows from what draws that boundary, whether it is a system prompt, a permission list, or a human checkpoint.

Where the money actually is

The most dependable use case is unglamorous: collapsing the time it takes to answer an internal question that already has a determinate answer buried in unstructured documents. Somebody asks what the standard indemnity position is on a particular contract type, and instead of two hours of searching and asking colleagues, the answer arrives in ninety seconds with citations. The saving per interaction is small. Multiplied across an organisation it is not, and the quality effect matters more than the time: junior staff arrive at the right answer rather than a plausible guess.

Long-tail service triage is the second, and it comes with a cost that rarely appears in the business case. Agents handle the well-formed majority of inbound queries, order status, password resets, document copies, and route the rest to people with a structured handoff. The economics are well documented. What is less discussed is that removing the easy queue starves your team of pattern recognition and raises the average difficulty of everything they do see. That is a real cost, it lands six months later, and it shows up as attrition rather than as a line in the model. It can be managed by rotating staff or investing in training cycles, but it has to be planned for rather than discovered.

Third, and more surprising, are multi-step research workflows that used to need a junior analyst.

Where the economics invert

The clearest rule is about tail risk rather than accuracy. Agents are already drafting contract clauses, issuing internal communications and turning unstructured records into structured notes, and in each case the cost per output looks excellent right up until the first material error. Material errors are inevitable at scale. So average accuracy is the wrong thing to interrogate. What matters is the cost of the worst plausible failure and how far its blast radius extends. Where the honest answer involves a regulatory fine, a lawsuit or harm to a person, the agent is the wrong tool at any accuracy.

The second boundary is judgement. Agents execute well-defined work well and handle ambiguous briefs badly, because responding to an ambiguous brief means asking the right clarifying question and shaping the work, not performing it. Attempts to put agents at the front line of a judgement-heavy advisory practice tend to be retired quietly a few months in. This is the limitation most likely to soften as models improve, so it is worth holding loosely. The other two are structural and will not move.

The third is the quiet one, and it is the most common. The agent completes the task in four minutes. A person then spends twelve minutes verifying it, and verification turns out to be more cognitively expensive than doing the work would have been. When that pattern appears, it is diagnostic: it means the process was not ready to be automated. The correct response is to make the underlying process cheaper to verify, not to go looking for a better agent.

Who owns the agents

Most first agent programmes open with a debate about which model to use. The question that will still matter in three years is who owns the agents and what data they are permitted to touch. In most organisations today the honest answers are whichever team built it, and whatever they could get a service account for. Neither survives a serious data audit.

What works at scale is a split. A central platform team owns the infrastructure, model access and the audit log. Business units own the prompts, the tools and the use cases. Platform holds a veto on data access, the business holds a veto on prompts. It mirrors how competent security functions already operate, and it heads off the two predictable failures: a central team that becomes the bottleneck everyone routes around, and a business team that sends HR records through an inference API in another jurisdiction without anyone noticing.

That governance question is the one that outlasts every model release. Capability arguments have a shelf life measured in months. Who is allowed to touch what does not.

If you’re sketching out an agent strategy and want a sounding board on which use cases are likely to earn their keep, that’s a thirty-minute conversation we’d happily have. Get in touch, or read more about how we run engagements.

← All field notes