Late one evening in August I asked my compliance agent the same question twice.

The build is a German EV-charging compliance model — a fictional operator in a real Developer Edition org — and the question was the plainest one the domain has: is this charge point legally compliant? I asked it, read the answer, then scrolled up and asked it again, word for word, about the same charge point.

I got two different answers. Both of them were wrong.

Neither answer had come from the org. An action was misconfigured and had never run, so the agent had nothing to report — and instead of saying so, it built an answer out of the only material it had: the wording of its own instructions. My instructions described what the law considers, so the agent described what the law considers, in confident, fluent, well-organised German. Twice. Differently.

An agent that can compute an answer will compute one. And it will sound exactly as confident when it is wrong as when it is right.

That sentence is the whole post. What follows is what I changed because of it.

Confidence is not a signal you can read

The uncomfortable part was not that the agent was wrong. The uncomfortable part was that nothing in the reply looked different from a correct one. Same tone, same structure, same certainty. If I had not asked twice, I would have had no reason to doubt the first answer.

You cannot audit fluency. You cannot put a threshold on it. So the design question stops being how do I make the agent more accurate? and becomes how do I arrange things so the agent is not the one deciding?

I gave the answer a name so I would stop drifting from it: the engine decides, the agent explains.

The determination lives in a formula field

The first place the determination went was the least glamorous object in Salesforce: a formula field.

A formula field is a read-only field the platform calculates from other fields on the record, every time it is read. It is not a script and it is not scheduled — it is simply the record telling you what it already is. In this build, the legal status of each charge point is a formula field, which has one consequence I care about more than any other:

You can sort a list view by it.

That sounds small. It is not. It means a compliance officer with no access to the agent, no interest in AI, and no patience for a chat window can open a list view, sort by legal status, and see the whole estate in one screen. The determination is now a property of the data, visible to anyone with read access, identical for every reader. The agent is one of several ways to look at it — not the place it lives.

Where a formula runs out, Apex takes over

Formulas are excellent at state and poor at story. Some of what this domain requires is chronology: which thing happened before which other thing, and how long the gap was. A formula cannot express that cleanly, so that part is Apex — ordinary, tested Apex, sitting behind an action.

The division is easy to remember:

  • Formula: what is true of this record right now, readable and sortable by everyone.
  • Apex: what is true about the sequence of events, computed in code that can be unit-tested.
  • Agent: neither. The agent’s job starts after both have answered.

Take the arithmetic away, not just the authority

Here is the change I would defend hardest, because it is the one that survives a bad instruction.

Telling an agent “do not work out the dates yourself” is a request. It is written in the same instructions the agent already showed me it would happily improvise from. On a bad day, that line is just more text.

So instead of asking, I removed the capability. There is no date arithmetic on the agent’s side of the call. The action returns the determination; it does not return two raw dates and hope. The agent cannot subtract one date from another to reach a legal conclusion, because it never holds two dates to subtract. Even if my instructions were deleted tomorrow — even if someone pasted a clever message telling it to work the answer out itself — it physically cannot.

Do not instruct an agent out of a capability it still has. Design the capability away, and the instruction becomes a courtesy rather than a control.

A gate that reads the transcript

Removing the arithmetic stops the agent computing. It does not stop it quoting. My original failure was an agent reciting legal paragraphs — in German law a rule is cited by its § number — that no action had returned.

So the build has a transcript gate. After a run, it takes every legal paragraph the agent uttered and checks it against the paragraphs the actions actually returned. If the agent said one that no action produced, the build fails.

Three things make it worth copying:

  1. It is binary. Pass or fail. There is no score to argue with.
  2. There is no model in it. No second agent grading the first, no judgement call — string comparison against a list.
  3. It proves it can fail. On every run, the gate is exercised against a case it must reject. A check that has never failed is not a check; it is a decoration, and you cannot tell the difference between “everything is fine” and “the check is broken” until the day it matters.

Twenty-five real charge points, and every one says UNBEKANNT

The build imports twenty-five real Berlin charge points from public data, live, with no API key required. Every single one of them evaluates to UNBEKANNT — unknown.

The first time I saw that column, my instinct was that something was broken. It was not. No public database carries the fact the law actually requires, so the honest determination for all twenty-five is unknown.

I could have made that column green. It would have taken one line. And it would have been the most dangerous line in the project, because a missing date would have quietly become a clean bill of health — and nobody investigates a green light.

UNBEKANNT is not a gap in the model. It is the model working: the estate has twenty-five charge points whose status nobody can currently prove, and that is a true and useful thing for a compliance officer to know before an auditor tells them.

Your next step

Take one agent you have built and find the sentence in its reply that carries the most consequence — the number, the status, the yes or no. Ask where that value comes from. If the answer is “the model composed it from context”, move it: into a formula field if it is a property of a record, into tested Apex if it needs sequence or history.

Then do the harder half. Look at what your action still hands back, and remove anything the agent could use to reach the conclusion on its own. Leave it the sentence to write, and nothing to work out.

Mustafa Aksu

Salesforce developer & ISV builder focused on Revenue Cloud, Agentforce, and Data Cloud. I write from real, shipped work.