A customer asked my agent a question about a number, and it answered with a paragraph of policy. Well written, entirely on topic, the wrong shape. The customer wanted a figure; what came back was an explanation of how such figures are generally arrived at.
My first instinct was to blame the model. That instinct was wrong, and unlearning it is the most useful thing I can hand you today. The model had invented nothing. It had been handed a policy article and asked to answer from it, and it did exactly that. The mistake was mine, one layer down, in the part of the system that decided what to hand it.
That layer is called grounding, and it is not magic. It is retrieval.
Retrieval is a step, and it happens before the writing
Here is the sequence, plainly. A question arrives. Before the model writes a single word, something searches your data and pulls back material that looks relevant. That material goes into the prompt alongside the question. Only then does the model write — and what it writes is a rephrasing of what it was given.
Two consequences follow, and they are the whole post.
The first is the good news: the model never gets a vote on what the facts are. If the retrieved row says one thing, the answer says that thing. This is why grounding stops an agent inventing plausible fiction.
The second surprises people: the answer can only ever be as good as the retrieval. If the search returns the wrong material, the model will present it fluently and confidently, because it has no way of noticing that it was handed the wrong thing.
Grounding does not make an agent right. It makes an agent repeat whatever the retrieval step found — accurately.
Two shapes of question
Once you see grounding as retrieval, a distinction appears that changes the design.
Some questions are explanation-shaped. “Why was I charged for an estimated reading?” “How does the tariff work?” The right answer is prose. It exists as written policy somewhere, it applies to many customers, and it changes when the business decides it changes.
Other questions are record-shaped. “What was my reading last month?” “Has my order shipped?” The right answer is a value belonging to one specific customer, and it changes whenever the data changes.
These need different retrieval mechanisms, because they live in different places and are found in different ways: prose by meaning, a record by identifier.
Split the grounding layer on purpose
In the build I keep coming back to on this blog — HanseWatt, a fictional DACH energy retailer in a demo org, where the company is invented but the agent and its grounding are real — I split the grounding layer along exactly that line.
Knowledge articles carry policy and explanation. When a customer asks why something works the way it does, the agent is grounded in the published article that says so. The benefit is not only accuracy: the agent’s explanation and the company’s documented position are the same sentence, and they change together.
Data 360 retrievers carry the record-shaped facts. Data 360 is the platform many of us still call Data Cloud; a retriever is the configured search that fetches the relevant records for a question. When the answer is a number about one customer’s meter, it comes from there — queried from the source of truth, and phrased, not decided, by the model.
Two sources, one agent, routed by the shape of the question rather than by convenience.
What mixing them produces
Now the failure I opened with. If you pour everything into one undifferentiated pile and hope relevance sorts it out, you get an agent that quotes a policy where a number was needed.
Look at why, because it is not a bug in the search. A policy article about billing is genuinely relevant to a question about a bill. Relevance scoring is doing its job. It simply has no concept of “this question needs a value, not a paragraph”. That distinction is yours to encode, and if you do not, retrieval will make a reasonable guess and be wrong in a way that reads perfectly.
The same failure runs the other way, and it is uglier: a general “how does this work” question answered with one customer’s record. Now you have a wrong answer that is also somebody else’s data.
The quieter failure: identity resolution
There is a second failure mode in the record-shaped half, and it is the one to remember longest.
Data 360 performs identity resolution: it decides which records from different systems represent the same person. A billing account, a web signup, a support contact, a meter installation — matching rules link them, and the result is one unified profile that grounding then queries.
When those rules are too generous, records get merged that should have stayed apart. Two people sharing a surname and a postcode become one person. And now an agent can speak about the wrong customer.
Here is what makes this worse than the first failure: every individual query still looks correct. Ask “what is this profile’s latest reading?” and you get a real reading, formatted properly, from a real record. Nothing is malformed, nothing errors. You can read the retrieval log line by line and find no fault, because the fault is not in any query — it is in the assumption underneath all of them, that this profile is one person.
You cannot debug this by inspecting answers. You have to inspect the identity graph itself: take a handful of unified profiles, list the source records folded into each, and check by hand whether they belong together. It is slow, unglamorous work with no substitute. Where the matching and reconciliation settings live varies, so verify in your own org.
Test the retrieval, not just the answer
So: when an agent gives a poor answer, do not start by rewriting the instructions. Start by asking what was retrieved.
Run the retrieval on its own and read the rows. Data 360’s SQL is inspectable — you can see exactly what was asked and what came back, before any model touches it. Nine times out of ten the answer to “why did the agent say that?” is sitting in those rows, and no amount of prompt tuning would have fixed it.
Then build the habit forwards. Write down a handful of real questions of each shape, run the retrieval for each, and check what comes back before you evaluate a sentence of prose. A retrieval layer you have tested is a grounding layer. An untested one is a hope.
Your next step
Take ten questions your users actually ask — real ones, in their words — and sort them into two columns: explanation-shaped and record-shaped. That sort alone usually reveals a source you have not configured, or two you have quietly merged.
Then pick three unified profiles at random and look at what was folded into each. If you find a merge you cannot defend, you have found the failure that hides best — and found it before a customer did.