Semantic drift: the invisible risk in multi-agent systems
When several agents share the same words but not quite the same business definitions, every component can work correctly while the system as a whole quietly starts to drift.
Moving from a single agent to a system made of dozens of agents changes the nature of the architecture problem. With a single agent, a business definition can still be encapsulated in a prompt, a knowledge base, a function or a few application rules.
With several agents exchanging information, querying different sources and executing actions on the information system, a new property becomes critical: do they actually share the same understanding of the business? Not just the same data — the same meaning. This is the problem of semantic drift.
Semantic drift is not hallucination
Consider an apparently simple notion: ActiveCustomer. For the sales agent, an active customer may be one who placed an order in the last twelve months; for the finance agent, an open account with no payment incident; for the support agent, a customer holding a valid maintenance contract. All three agents can be right.
The problem appears when they start exchanging that information in a simplified form:
Customer A is active.
No hallucination necessarily occurred. The source data, the retrieval, the local reasoning and the tool call can all be correct. And yet the agent receiving the information can make a bad decision, because it assigns a different semantics to the word active.
Semantic drift, as used here, is the divergence in meaning of a concept shared across several components of an agentic system. It should not be confused with concept drift in machine learning models, where the data distribution shifts over time: here the data does not necessarily change — it is the interpretation of the same word that diverges between the components exchanging it. So the question is no longer only whether an answer is true: you also need to know in which model of the business it is true.
How drift appears
The problem already exists in traditional information systems: two applications can hold different definitions of the customer, the order, the revenue or the risk. Agentic architectures, however, increase the exposure surface. The same business rule can now be encoded simultaneously in:
- a system prompt;
- a knowledge base;
- a SQL query;
- the code of a tool;
- a business API;
- an agent memory;
- a semantic model;
- the instructions passed from one agent to another.
Each representation can evolve independently. A seemingly harmless change is then enough to introduce a divergence.
Agent Sales
|
| "Customer is active"
v
Agent Finance
|
| active = account_open && payment_status == OK
v
DecisionThe problem is not the message. The problem is the absence of a contract about its meaning.
RAG solves access to knowledge, not unity of meaning
RAG architectures have greatly improved the ability of models to access enterprise information. But finding the right document and sharing a business model are two different problems.
RAG primarily answers the question: which information should be provided to the model?
Semantic drift raises another one: which definition should be used to interpret that information?
An agent can perfectly retrieve the latest sales procedure while applying a customer definition that comes from another source. Adding more documents to the context does not necessarily solve this problem — it can even multiply the available definitions. Retrieval quality remains indispensable; it is simply not sufficient.
The harness does not guarantee semantic consistency either
The same reasoning applies to the harness. A robust agentic runtime can handle planning, tool calling, errors, retries, permissions, traces and execution policies. It can guarantee that the agent correctly executes:
plan -> act -> observe -> update -> respondBut a perfectly executed loop running on an incorrect business definition is still an incorrect loop — technically correct and semantically wrong. The harness controls execution; it cannot, by itself, guarantee that two agents give the same meaning to the concept they manipulate.
The semantic layer changes role
This is probably one of the most interesting architectural evolutions of the moment. Historically, the semantic layer was mainly developed for analytics: it defines dimensions, measures, relations and indicators so that different reports use the same interpretation of the data. The arrival of agents considerably extends its role, and Microsoft Fabric IQ illustrates this evolution well, with a progression from unified data to semantic models, ontologies and agents.
The ontology no longer merely describes a data structure. It can represent:
Customer
|
+-- places --> Order
|
+-- owns ----> Contract
|
+-- has -----> RiskProfiletogether with the associated properties, relations and rules.
The semantic layer then stops being only a way to produce consistent analyses: it starts providing an operational representation of the business, usable by software systems and agents. This is an important shift — once agents start acting, semantics progressively becomes a runtime dependency. Analysts are converging on the same observation: Gartner predicts that 60% of agentic analytics projects relying solely on tool-access protocols such as MCP will fail by 2028, for want of a consistent semantic layer underneath.
From semantic model to semantic contract
This evolution suggests a change of perspective: the critical concepts of the enterprise should be treatable as contracts. The idea is familiar to data teams — it is the natural extension of data contracts, which make explicit the schema and guarantees of a data flow between a producer and its consumers. The semantic contract is to agents what the data contract is to pipelines: it makes explicit not the structure of the data, but the business definition used to interpret it.
A semantic contract could, for instance, spell out:
concept: ActiveCustomer
version: 3
definition:
customer_status: OPEN
commercial_activity_within: 365d
source:
system: CRM
entity: Customer
owner:
domain: Sales
valid_from: 2026-09-01This format is not a standard. It simply illustrates an important property: the definition becomes explicit, identifiable and versionable.
One point deserves emphasis: the contract does not impose a single definition on the whole enterprise. Sales, Finance and Support can keep their respective definitions — Sales:ActiveCustomer:v3 and Finance:ActiveCustomer:v1 can perfectly coexist, each with its own owner. The problem is not the plurality of definitions, which is often legitimate; the problem is exchanging them under the same word without qualifying which one is being used. The contract makes the divergence explicit and addressable, where it used to be implicit and invisible.
An agent should then no longer merely produce:
{
"customer": "A123",
"active": true
}but be able to attach that decision to the concept and version it used.
{
"customer": "A123",
"active": true,
"semantic_contract": "Sales:ActiveCustomer:v3"
}We then obtain something that agentic systems still often lack: the semantic provenance of the decision.
The problem becomes critical at scale
For three agents built by the same team, maintaining this consistency manually remains possible. For fifty agents spread across several domains, teams and platforms, the situation changes quickly. Consider just five concepts:
Customer
Order
Revenue
Eligible
RiskFive concepts shared across fifty agents already means two hundred and fifty possible local interpretations. And the consistency to verify is not linear: it applies to every pair of agents exchanging those concepts — potentially more than a thousand combinations for a single ambiguous word.
If each team encodes its own interpretation of these concepts in its prompts and tools, the divergences become combinatorial, and the organization progressively recreates the semantic silos already present in its information system.
With one major difference: these new silos can make decisions and trigger actions.
This is probably where semantic drift becomes an architectural risk rather than a mere answer-quality problem.
Towards semantic observability
AI platforms are starting to provide increasingly fine-grained observability:
prompt
model
tokens
latency
tool calls
traces
errors
evaluationsFor multi-agent systems, this may no longer be enough. When an agent makes a decision, several questions become useful:
Which business concept was used?
Which definition?
Which version?
Which source is authoritative?
Which agents use this definition?
Has the definition changed since the last evaluation?We can call this capability semantic observability. The goal is no longer only to detect: this agent produces bad answers. But also: these two agents no longer reason from the same model of the business. This distinction matters — the first problem can be detected by a local evaluation; the second sometimes only appears at the moment several agents start collaborating.
Two concrete mechanisms are already emerging on the data-platform side and can be transposed to agents: definition consistency checks integrated into CI/CD, which surface a conflict at build time rather than at runtime, and semantic lineage — knowing who changed which definition, when, and with what downstream impact.
Testing semantics like code
If semantics takes part in execution, it must progressively adopt some properties of software engineering. A definition change should not be treated as a mere documentation update: it can change the behavior of a whole set of agents. An evaluation chain could therefore include semantic consistency tests.
def test_active_customer_semantics():
sales = sales_agent.classify(customer)
finance = finance_agent.classify(customer)
# Both agents declare their contract; the pair must be
# either identical or explicitly mapped.
assert (sales.semantic_contract == finance.semantic_contract
or is_mapped(sales.semantic_contract, finance.semantic_contract))The test is deliberately simplified, but the principle matters: the consistency of shared concepts becomes a testable property of the system — including when two distinct definitions legitimately coexist, provided their mapping is explicit. These divergences can now be treated as a measurable property of the system.
We can then imagine quality gates covering:
- the version of the semantic model;
- the compatibility between versions;
- the provenance of definitions;
- the inter-agent consistency;
- the impact of a change;
- the concepts used during a decision.
Semantic drift thus becomes observable before it becomes a business incident.
Do not build a universal ontology
The answer to semantic drift is not necessarily to model the whole enterprise before building a single agent: that approach would quickly lead to an architecture too heavy to move. Not all concepts carry the same risk; priority must go to those that cross agent boundaries or trigger important decisions.
A good starting point is to identify the concepts that have at least one of these characteristics:
- they are used by several agents;
- they trigger an action;
- they represent a regulatory or financial decision;
- they have several historical definitions in the information system;
- their misinterpretation can produce a significant business effect.
The semantic architecture can then grow with the agentic system.
One more engineering discipline
Production agentic architectures progressively accumulate several layers:
Models
Context
Memory
Harness
Tools
Identity
Policy
Evaluation
ObservabilityWe should probably start considering another one explicitly:
SemanticsNot because models would be unable to understand business language, but precisely because they are able to. An LLM can interpret several plausible definitions of the same concept and keep producing a perfectly convincing answer. At the scale of one agent, this flexibility is a strength; at the scale of a distributed system of agents, it can become a source of divergence.
The question is therefore no longer only: how do we give our agents enough context? It becomes: how do we guarantee they interpret that context from the same model of the business?
Semantic drift may be one of the first genuinely distributed problems of agentic architecture. And mastering it will probably rely less on a better prompt than on practices architects already know well: contracts, versioning, governance, testing and observability.
References
The reflection presented here notably builds on the evolution of Microsoft architectures around Fabric IQ and Microsoft IQ: semantic models, ontologies, knowledge bases and shared context for agents. It nevertheless develops a general architecture problem, independent of any particular platform. The phenomenon is also documented on the data-platform side — notably under the terms semantic drift and context drift — confirming that it extends beyond agentic architectures alone.
The prediction quoted about agentic analytics projects was presented at the Gartner Data & Analytics Summit 2026; see also Gartner, "Lack of Semantics Causes Inaccurate AI Agents and Wasted Spending", May 2026.