Est.

Semantic Layers vs Governed Datasets for Agent Context

Governed datasets close gaps that semantic layers alone leave open for agent errors.

Contributing Analyst · · 11 min read
Cover illustration for “Semantic Layers vs Governed Datasets for Agent Context”
The Agentic Data Contract · October 5, 2026 · 11 min read · 2,539 words

An AI agent that queries raw tables will give a confident, wrong answer before it gives a right one, and it will do so without any signal that something went wrong. That is the subject of this piece: why semantic layers, built to solve exactly this problem, still leave agents exposed to it, and what a governed dataset architecture adds that closes the gap.

Why agents querying raw data produce confident wrong answers

Ask three different AI systems to calculate monthly recurring revenue from the same operational database, and three different numbers come back. One assistant sums every invoice. Another excludes trials. Finance, working from its own process, arrives at a third figure. All three queries execute without error. All three draw from the same underlying data. None of them is flagged as wrong, because none of them is wrong in the way a system checks for wrongness: the SQL runs, the aggregation completes, and a number comes back formatted like an answer. The failure sits one level up, in the question of which calculation the business actually means when it says "MRR," and no query engine resolves that question on its own.

This matters because research and data analysis now accounts for 24.4% of agent deployments, the second most common use case after code generation. An agent connected to raw tables has no tribal knowledge to draw on. A human analyst at a company knows which metrics finance trusts, which tables are stale, which column was renamed last quarter and never cleaned up elsewhere. An agent has none of that unless it is made explicitly machine-readable, a point OvalEdge's enterprise guide makes directly: agents do not inherit institutional memory, they infer from column names and schema patterns, and inference is a guess dressed up as a calculation.

The danger compounds with autonomy. A wrong number on a dashboard gets caught eventually, usually when someone reconciling two reports notices they don't match. An agent acting without a human in the loop, calculating a figure and then deciding or executing based on it, has no equivalent checkpoint. The error isn't caught at the moment it's made; it appears later, if at all, during the kind of manual reconciliation the agent was supposed to replace.

What a semantic layer does

A semantic layer exists to end exactly this kind of definitional chaos. It sits between raw data and every consumer of that data, translating physical structures, tables, joins, column names, into a shared business vocabulary that both humans and machines can interpret regardless of their familiarity with the underlying schema. Done well, it standardizes five things across every downstream surface: metric definitions (how "churn rate" is actually calculated), dimensions (customer segments, time periods), relationships (how tables join to each other), business terminology, and access rules, as ATScale's semantic layer glossary lays out.

Two layers of context matter here, and they're often treated as one thing when they aren't. A semantic model describes what data means: which entities exist, how they relate, what synonyms apply, what a given table is for. That tells an agent where to look. Metric definitions are a separate and more exacting layer: the precise business recipe for a specific number, whether revenue excludes refunds, whether it's measured gross or net, what grain it's rolled up at. That second layer is what determines whether the answer an agent produces is one finance would actually defend in a board meeting. A team can build a strong semantic model, get the entities and relationships right, and still watch agents produce wrong numbers, because knowing which table holds revenue isn't the same as knowing which version of revenue the business has agreed to use.

Databricks' architecture guidance frames the absence of this layer as accumulating "decision debt," ambiguity left unresolved that eventually comes due in reconciliation meetings and missed opportunities. The framing is apt because semantic layers, historically, were built for a world where that debt got serviced by a person. Traditional semantic layers assumed a human was looking at a static dashboard, available to question a number that looked off. OvalEdge's guide draws the contrast explicitly: traditional semantic layers support reporting, while governed semantic layers have to support AI reasoning and automated decisions, where certified definitions need to be enforced at query time, automatically, every single time a question is asked. That's a different design target, and most semantic layers in production today were not built toward it.

The three governance gaps a semantic layer alone does not close

Diagram: Three Governance Gaps a Semantic Layer Leaves Open. Visualizes: Show three sequential gaps that a semantic layer alone fails to close, and how a governed dataset closes each one.

A well-built semantic layer still leaves three gaps open, and each becomes serious the moment an agent is making calls on its own.

The first is the problem of competing definitions with no declared winner. Most organizations are running several versions of the same metric simultaneously: one that finance built and trusts, one marketing built for attribution and never fully retired, one left over from a growth experiment nobody deleted, one the data team marked canonical months ago but never enforced downstream. A semantic layer can expose all of these versions faithfully. It does not, on its own, tell an agent which one the business actually stands behind. That resolution requires a layer of governance, ownership, verification, lineage, certification status, sitting on top of the semantics, and resolving it inside the agent's context is expensive: the GROUND research paper puts the token cost at roughly five times a direct schema-only baseline.

The second gap is semantic bypass. Nothing architecturally stops an agent from calling a raw query tool instead of an approved metric interface, if both are available to it. MCP security guidance identifies this as a structural risk rather than a hypothetical one: if raw table access sits exposed alongside governed metric endpoints, an agent can route around the semantic layer entirely, and there is no definition-level check that catches it, because the bypass happens at the tool-selection level, not the query level. The semantic layer governs meaning. It has no say over which tool gets called.

The third gap is timing. A semantic layer translates a question into a query and runs that query at the moment it's asked. Every question an agent poses becomes a live query against the production database or warehouse. It does not control when queries run, how many run simultaneously, what load they put on a system the application also depends on, or whether two agents asking the same question seconds apart get the same answer back. The practical result is what OvalEdge's guide documents as metric drift: marketing defines "active customer" one way, finance another, and an agent facing both options picks whichever field looks statistically closest to what it expects, with no record of having made that choice.

What governed datasets add beyond a semantic layer

Governed datasets close these three gaps by changing when calculation happens, not just how meaning is defined. A governed dataset is pre-modeled, pre-calculated, and access-controlled before any question is asked of it, which separates the moment of calculation from the moment of consumption. Both humans and agents then read from a single answer that was computed, certified, and frozen in advance, rather than assembled fresh from a live system every time someone asks.

That closes the first gap directly. A governed dataset is a calculated artifact tied to one certified definition, not a query that could resolve against any of the competing versions floating around the organization. Ownership, lineage, and certification status travel with the artifact itself as conditions of its existence, not as metadata an agent has to interpret correctly.

It closes the second gap by removing the choice that makes bypass possible. When the governed dataset is the only thing an agent can reach, not raw tables, not an unapproved metric endpoint sitting alongside it, routing around the semantic layer stops being an available option. Dreambase addresses this by pre-modeling business metrics into certified datasets that agents consume rather than query, locking the definition in before the agent ever encounters the raw schema, which eliminates the invisible failure mode where an agent returns the wrong version of a metric with full confidence and no awareness that a choice was made. The interface, whether MCP or an API, surfaces the finished artifact. Production tables are never in scope to begin with.

It closes the third gap because pre-calculation means the answer exists before the question arrives. Agents get fast, cheap results while the production system stays untouched, and the cost of calculation gets paid once, at modeling time. A useful distinction exists between a semantic layer, which defines what data means, and a context layer, pre-materialized operational data an agent can actually consume. A governed dataset is that context layer made concrete: not a translation step but a finished, trusted artifact sitting ready. Atlan Frontier Labs' analysis found that governing this kind of context improved AI-generated SQL accuracy by 38% relative, measured across 174 enterprise-complexity queries, a gain that comes specifically from removing the ambiguity a live, ungoverned query would otherwise have to resolve on the fly.

Governed datasets solve all three gaps simultaneously: they pre-declare which metric definition the business stands behind, eliminating ambiguity over which revenue figure is correct; they exist as immutable artifacts separate from production tables, which prevents semantic bypass by design rather than by policy; and they arrive pre-calculated and access-controlled rather than on demand, so an agent never needs raw query access in the first place. None of this displaces the semantic layer's job. It extends that job past the point where human oversight used to pick up the slack.

How Supabase and Postgres make this architectural choice concrete

Supabase's Series F in June 2026 established that agents are now deploying most of the databases on the platform. The agent-to-data access question is the condition most Supabase teams are already operating under.

Supabase's own guidance draws a clear line around what Postgres is for: excellent for transactional workloads, reading a user profile, inserting an order, but poorly suited to analytics workloads that scan large volumes of data and aggregate across many rows on the same system the live application depends on. The platform recommends separating the two explicitly rather than treating Postgres as an all-purpose engine.

That recommendation has started showing up in the platform's defaults. Starting April 28, with the change effective May 30 for all new projects, Supabase projects can opt out of automatic Data API exposure for public schema tables, meaning explicit Postgres grants are now required before a table becomes reachable through PostgREST or GraphQL. That's a direct architectural response to the reality of AI agents gaining uncontrolled access to production tables. Supabase Pipelines, launched in July 2026 and in public alpha on paid plans, gives teams a way to move production data into systems built for analytics while keeping the application's own workload on Postgres, which is the platform building, in infrastructure, the separation it has been recommending in guidance.

The risk this is responding to is concrete. The Supabase MCP server supports both read and write operations, so a connected agent can query schemas, insert records, and manage data directly. That capability is powerful, and it is dangerous without a governed layer sitting between the agent and the tables it can reach.

How Dreambase implements the governed dataset pattern for Supabase teams

Dreambase functions as that governed layer for Supabase teams specifically. It pre-models business metrics into Parquet datasets, queryable through DuckDB, and exposes all of it through a single MCP server, so agents get fast, accurate, governed answers without ever touching the production database directly.

Pre-calculation closes the problem of agents querying on demand described earlier. Metrics are computed ahead of time and stored as governed Parquet datasets, so the answer exists before an agent asks for it, the calculation cost gets paid once rather than on every call, and the production Postgres instance carries none of the analytics load that would otherwise compete with the application it serves.

A single MCP server closes the bypass gap the same way restricting table exposure does at the platform level. Agents connect to one governed interface, and raw production tables are simply never in scope, mirroring the logic behind Supabase's own shift toward requiring explicit grants before a table becomes reachable. A single pre-modeled definition resolves the ambiguity over which revenue figure is correct: Dreambase treats a single definition as the source of truth across every consumer, dashboards, agents, reports, Slack messages, so metric divergence across teams becomes structurally impossible rather than something a governance policy merely discourages.

None of this requires a Supabase team to build new infrastructure. There's no ETL process to stand up, no schema changes to make, no separate data warehouse or data engineering function required. Dreambase extends a project's existing Postgres setup into full analytics capability, lining up directly with the separation Supabase itself recommends: application workload stays on Postgres, and analytics capability gets added on top of it.

The same governed dataset becomes the shared unit of consumption for people and agents alike. A founder checking a metric in a dashboard and an agent pulling context for a workflow read from the same certified artifact, which removes the divergence that occurs when humans and agents draw from two different surfaces. Metrics get pushed to where people already work, inbox, Slack, rather than waiting to be pulled on demand, extending the governed dataset pattern to human workflows and not only agent ones. None of it assumes a dedicated data team: the pattern is built for founders and small teams who need numbers they can bring to a board meeting without first hiring someone to build the pipeline that produces them.

Checklist for a governed data layer before production

Before connecting any agent to a data layer, a team should confirm five architectural properties hold, and treat the absence of any one of them as a reason to stop, not a gap to patch after launch.

Production database isolation has to be real and unconditional: agents should never reach live tables, even under read-only permissions, because the query itself carries risk independent of what it returns. Confirm the layer in question exposes only governed artifacts and nothing resembling raw schema access.

Pre-calculation has to replace on-demand querying as the default. A layer that computes metrics at request time turns every agent call into a database hit. A layer built on governed datasets computed in advance is cheaper, faster, and consistent across simultaneous calls, properties that matter more as agent volume grows.

A single certified definition has to exist for each metric that matters. A layer that exposes multiple versions of "revenue" and leaves the choice to the agent has not resolved which figure is correct. It has delegated it to the one party least equipped to resolve it. One definition, one owner, one published artifact is the standard a layer should be held to.

Lineage and auditability need to live at the metric level, not just the query level. When an agent's answer gets challenged, and eventually one will be, the team needs to trace that answer back to the exact dataset version and calculation that produced it, rather than attempt to reconstruct the query after the fact and hope the production data hasn't shifted underneath it in the meantime.

Sources

  1. GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions
  2. Dreambase — The Data Factory for your Agents

More in The Agentic Data Contract