Enterprise AI: Separate the Reasoning from the Calculation

Most enterprise AI failures aren’t LLM failures. They’re governance failures, and buying a better LLM won’t fix them.

Estimated Reading Time: 4 minutes

At this year’s Semantic Layer Summit, practitioners from Anthropic, Chevron, WPP, Accenture, NVIDIA, and OpenHands, compared notes on why production AI keeps breaking down. The consensus? The LLM is not the bottleneck. It’s context.

The organizations succeeding in production don’t have smarter AI. They are separating what AI reasons over from what it’s allowed to calculate.

The reasoning is flexible. The calculation is not.

The LLM is not the problem

Launching an AI pilot is a lot like taking a test drive on a closed track. It looks promising because you’re using clean data and a narrow set of questions. Getting AI into production is taking the car out on the road, where you immediately encounter variables like traffic, pedestrians, and weather. 

The failure isn’t always a crash, but context drift can do significant damage to trust and accuracy.

AI might reference a 6-month-old document, applying a metric definition that one team retired and another team kept. It answers confidently using context that was accurate when the pilot ran, but isn’t anymore. 

Context drift is as dangerous as a hallucination because answers may look right but cause real damage, such as failed audits or decisions based on inaccurate numbers.

“Today the model is not the bottleneck,” said André Balleyguier, who leads applied AI teams at Anthropic. “Very often the context is.” 

A recent analysis published in Communications of the ACM confirms that the failure comes from AI not having access to the right information at the right time. Context drift is the gap between what the AI system knows and what is actually true.

When asked what the market most consistently underestimates about getting agents into production, Ted Kwartler from Accenture pointed directly to data quality and retrieval:

“When you have these nonlinear systems, if you provide the wrong context or conflicting information, you can really cause a lot of problems.” 

Ikechi Okoronkwo from WPP echoed his sentiment:

“Data quality and data integration are the main bottlenecks to production scale. It’s not really just about model intelligence.” 

What breaks when you leave the pilot

“The speed side of the equation has been solved. What you need is confidence at scale,” said Ikechi.

There’s no question AI can generate answers quickly, summarizing and surfacing patterns at a pace no human analyst can match. The problem is that enterprises can’t act on an answer they don’t trust. 

MIT Sloan researchers studying the production of agentic AI deployments found that 80% of the implementation work for one AI system focused on data engineering, stakeholder alignment, governance, and workflow integration. The AI was the easy part. The context infrastructure was the hard part.

Here’s how Rajiv Shah from OpenHands explained the metric challenge:

“I have four kids, and if I’m going to plan a vacation, I’m going to have seven different definitions of fun.”

The same fragmentation exists inside every large enterprise. If revenue in the CRM and revenue in the data warehouse aren’t the same number, AI can’t resolve that ambiguity on its own. The model will confidently pick one or blend them, and it will be wrong. The issue is governance.

How Anthropic Wired It (And Why It Works)

André described how Anthropic handles context internally: there is a governed layer managed by the finance team through which financial metric queries are routed. Ask the same question twice, and you get the same answer because the metric definition is fixed and the logic is deterministic. 

“The actual number needs to come from a deterministic system,” he said. “Because this number has been vetted and approved by the compliance teams.”

He explained how agents above that layer can reason probabilistically: planning, orchestrating, and recommending. But when they need a specific metric, they call a tool that returns a deterministic output.

AI handles interpretation and orchestration. The semantic layer handles definitions, calculations, and governance. They work together, each serving a different function.

Once those roles are clear, the work shifts back to the humans who act on the output. Josh Patterson from Nvidia argued: “The output of a business person is not just a number. It’s a strategy.” AI’s job is to help people build strategy, surfacing insights grounded in numbers that are consistent and auditable. 

MCP and the Case for a Governed Semantic Layer

Deterministic metrics require governed infrastructure. A semantic layer that exposes consistent definitions through Model Context Protocol (MCP) gives agents access to institutional context they can’t derive from raw tables or prompt engineering. 

The Futurum Group projects that the semantic layer is on track to be among the fastest-growing segments in data infrastructure, with growth rates expected to double over the next several years as organizations recognize that LLMs cannot operate reliably without governed metrics to ground them. Brad Shimmin, VP at Futurum, explains:

“Without a semantic layer, you don’t have agents; you have hallucination engines.”

The same governed definitions that power your dashboards are the definitions your agents need to operate reliably. AtScale’s semantic layer exposes those metrics and dimensions through MCP, giving agents access to the institutional context they’d otherwise have to guess at.

No matter how powerful an LLM is, if it can’t reliably compute your organization’s definition of “net revenue” quarter over quarter, it’s useless for the enterprise. 

The teams succeeding in production have already figured this out. They’ve separated reasoning from truth and given their agents the tools to reason well, while locking down the business logic they reason over. The teams still struggling are asking their models to do both.


All sessions from the 2026 Semantic Layer Summit are available on demand.

SHARE
How to Evaluate Context Platforms
How to Evaluate Context Platforms - buyer's guide

See AtScale in Action

Schedule a Live Demo Today