My Case for the Semantic Layer, Made Live on Atlan’s Podcast

What I told Austin Kronz when he asked me to define the context layer.

Estimated Reading Time: 7 minutes
Your semantics are your moat

I recently joined Austin Kronz on Atlan’s podcast, WTF is the Context Layer, to discuss one of the biggest sources of confusion in enterprise AI: the distinction between semantic layers, context layers, and ontologies.

We covered a lot of ground, but one theme kept resurfacing: the industry is suffering from a shortage of deterministic business meaning.

Every AI vendor now claims to have a context strategy. Over the last 18 months, I’ve seen context layers, context graphs, ontologies, knowledge graphs, semantic views, AI memory, and “business context” as a category unto itself. Vendors keep adding to the pile because it’s good marketing. But these solutions solve different problems. 

Here’s how I framed it for Austin: the context layer is a stack. An ontology is one part of the stack, serving as a data source that provides entities and their relationships. It sits alongside the physical schema (tables, columns, keys) as raw context. 

The semantic layer is the load-bearing component of the context layer stack. It maps the ontology’s logical entities to the physical schema. An ontology can feed into a semantic model, but it isn’t the semantic layer itself. 

Ultimately, the semantic layer is what makes AI deterministic, meaning you can trust the number instead of watching AI go after your raw data and guess at your metrics. 

You want AI to be creative and probabilistic. But you don’t want it to be “creative” when it’s calculating revenue or gross margin.

Context isn’t the problem. Trust is.

LLMs have access to too much information, and they’re failing because they’re being asked to calculate business metrics themselves.

This problem predates AI. Finance and sales have shown up with two different definitions of revenue for as long as there’s been a revenue number to argue about. What’s changed is velocity. In the BI era, a human asked the question so a human could catch the discrepancy. Now an agent can ask a thousand questions before anyone notices the number is wrong.

Reasoning should stay probabilistic. Business metrics computation should not. That’s the inversion missing from most of today’s AI messaging, which treats “more context” as an unqualified good. More context fed to a model that’s still doing its own math on revenue doesn’t fix anything. It just makes the guess better informed.

You just built a bad semantic layer. In an MD file.

What happens when companies skip this and attempt to build context through prompts instead?

When Databricks introduced Genie, a number of customers tried to work around the problem using Genie instructions (skill files). They started piling up page after page of instructions to get the agent to behave. At some point, you have to call it what it is: you just built a really bad semantic layer in an MD file.

This is where it ends up: every AI platform that treats prompts, instructions, skills, and workflows as a stand-in for governance runs into the same wall. Pile enough MD files under an agent and you’ve built business logic with no owner, no version history, and no audit trail, the exact gaps an auditor would flag first in any other system of record. 

Real governance means three questions have real answers on demand: who defined this metric, when did it last change, and who approved it. A skill file can’t answer any of them. Business logic belongs in infrastructure: versioned, owned, and auditable, not buried in a text file that only works because someone remembered to keep updating it. Once an agent asks the question instead of a person, the human who used to vouch for the number steps out of the loop. Governance is what stands in their place.

The next AI bottleneck is determinism

It’s not enough that an agent can generate SQL to get the right answer. It needs to consistently get the right answer. Run the same question ten times and see what comes back. What we’ve seen is that an agent can get it right on Monday and then do something completely different on Thursday, even when nothing about the underlying data changed.

A single correct answer tells you almost nothing if it isn’t reproducible. I call it the determinism gap: the space between an agent that happens to be right once and a system that’s right on every ask, on every day of the week. That gap is what stands between AI and being trusted with a board-level number.

I watched Databricks’ CEO get on stage and say Genie’s ontology gets it right 84.5% of the time. Run the math: you’re wrong more than 15% of the time. You’d get fired for that in any other part of the business.

Compare that to what governed semantics actually deliver. In a benchmark with a Tier 1 bank, adding a semantic layer took accuracy on real business questions from around 70% (the industry baseline on BIRD-SQL) to 100%. It also cut compute by up to 21,000x and erased roughly $9 million a year in what we call the AI rediscovery tax: the cost of an agent re-deriving the same metric from scratch every time someone asks. 

Your semantics are the only moat you have left

Another topic that came up was AI sovereignty, and Alex Karp’s recent CNBC interview, in which he discussed the risk of handing your competitive knowledge to a frontier model. I’d take it a step further.

Your competitive advantage is your semantics:

  • How you compute margin rate the moment your cost basis shifts mid-quarter
  • How you calculate gross margin return on inventory
  • How you decide whether a markdown was the buyer’s call or the store’s
  • How you define forecast accuracy well enough to actually act on it

Why would you hand that over to a single vendor or platform? It’s yours to own and yours to protect, because it’s what separates you from every competitor running the same models you are.

If the underlying intelligence is commoditized, and increasingly it is, the thing that separates you from a competitor with access to the same model is how you do business. In an agentic world, how you do business gets codified into semantics. That’s exactly why it shouldn’t be locked away in one vendor’s proprietary format. It’s also the argument for open standards like Apache Ossie, because your semantics moving freely between platforms is how you keep that advantage yours.

And it’s not just about protecting what you have now or where you are today. It’s also where you’re going or where you will be. Nobody knows what tool will be trending next year, or which engine ends up cheapest to run on. Lock your semantics to today’s platform and you’re just trading one dependency for another.

AI should ask questions, not invent your metrics

This might be the cleanest way to draw the line between what AI should own and what the semantic layer should own.

The agent should be creative. It should explore, reason, and ask deeper questions. What it shouldn’t do is compute your metrics itself, or decide on its own what revenue means this quarter. That’s the semantic engine’s job: not just describing what a metric means, but actually computing it and returning a deterministic result.

The agent shouldThe semantic layer should
ExploreCompute metrics
ReasonExecute business logic
Generate insightsGovern definitions
Ask follow-up questionsEnforce consistency
Explore many paths to an answerBe deterministic
Ask questions across systemsAccess any database
Work through whatever interface it’s givenRun on open standards

Plenty of vendors today can describe business meaning: here’s what revenue means, here’s the formula, here’s the definition in a catalog somewhere. Far fewer actually execute it, meaning take that definition, apply it to a 750-billion-row fact table, and hand back the same correct number whether it’s asked by a human or an agent, on a Monday or a Thursday.

A computational infrastructure hands you the right number every time, no matter who asks.

Final thought: Stop debating terminology

Austin asked how I’d want people to leave the conversation. My answer was simple: don’t spend months arguing about whether you need an ontology, a semantic layer, or a context layer.

Figure out which metrics actually matter to your business. Then ask yourself honestly whether you trust an AI agent to answer questions about those metrics unsupervised. If you can’t say yes, that’s the gap to close before you do anything else.

Don’t debate it. Benchmark it. Take the one metric your board already argues about and run it two ways: once through an agent guessing at raw tables, once through a governed semantic layer. See the Semantic Layer Economics of AI on the Warehouse case study for how that test played out for the Tier 1 bank (21,000x less compute, accuracy from 70% to 100%), then run the same test against your own warehouse and your own numbers.

Want the full back-and-forth? Watch the full episode of WTF is the Context Layer with Austin Kronz.

SHARE
How to Evaluate Context Platforms
How to Evaluate Context Platforms - buyer's guide

See AtScale in Action

Schedule a Live Demo Today