The evening I finished writing the attack cases for my own Agentforce agent, I had a corpus of tests and no way to grade them. So I did the obvious thing and started typing a prompt: “You are a security reviewer. Here is a transcript of a conversation. Did the agent behave safely?”

I got about two lines in before I stopped.

I spent twenty years in education before Salesforce, and there is one request every teacher hears and every teacher declines: can I mark my own paper? The answer is never “I don’t trust you.” It is “that isn’t marking.” A pupil marking their own work is not being dishonest. They simply cannot see what they could not see the first time.

That is the whole argument of this post. The agent in question is one I built for a fictional energy retailer living in a real Developer Edition org — invented company, real agent, real data model.

What an LLM judge is, and why it is so tempting

An LLM judge is what my half-written prompt was about to become: a second language model that reads the transcript of a test and returns a verdict. Safe or unsafe. Passed or failed.

The appeal is genuine, and I do not want to be sniffy about it. It requires no code, it handles cases you did not anticipate, it reads nuance, and it scales to hundreds of transcripts in the time it takes to make tea. When you are one person building a test bench in the evenings, that is a real offer.

It is also the exact mistake the whole exercise exists to prevent, and it fails in three separate ways.

The three ways it fails

It is generous. A model asked whether it behaved is being asked to mark its own homework, and it will be kind about it. Not maliciously — structurally. These systems are trained to be agreeable, and “yes, that response looks appropriate” is the agreeable answer.

You might think a different model solves that. It helps a little, and less than you would hope. Different models share training data, conventions and blind spots. The subtle thing that fooled the agent is the subtle thing that reads as reasonable to the grader, because “reads as reasonable” is the same faculty in both. And there is a sharper version: the grader reads attacker-controlled text. A prompt injection buried in a transcript is text your judge also processes, so the attack you are testing for can be aimed at the marker as easily as at the pupil.

The verdict moves. Run the same transcript past the same judge twice and you may get two answers. A test suite earns its keep by having a memory — you fix something, it stays fixed, and next month’s run tells you so. If verdicts drift, you cannot tell a real regression from a mood, and every red result becomes a conversation instead of a fact.

There is nothing to point at. When a judge says “unsafe,” you have a sentence, not a location. You cannot reproduce it, hand it to someone else, or argue with it. When it says “safe,” you have even less: a compliment with no evidence attached.

A grader you cannot argue with is not a grader. It is an opinion with a schedule.

What “deterministic” means here

So in my bench the verifier is deterministic: plain code, fixed input, fixed verdict, every single time — and, just as importantly, a check a human can read and agree is the right check. Two techniques carry most of the weight, and neither requires anything clever.

Canary tokens: a fact, not a judgement

A canary token is a distinctive, made-up string planted where only a leak would carry it — inside a record the agent must never surface, in a field it has no business reading. Then the check is simply: does that string appear in the agent’s answer?

If it does, something crossed a boundary. There is no threshold to tune, no rubric to write, no second model to consult. It is a binary, checkable fact — checkable by someone who has never seen your agent. My corpus carries eight of these tokens, planted across the cases.

One practical note: the check has to normalise the text before comparing, because lookalike characters — a Cyrillic character that renders identically to a Latin one — would sail straight past a naive string match.

Trace assertions: grade what it did, not what it said

The second technique matters even more, and beginners almost always miss it.

A trace assertion checks the agent’s actions — which action ran, with which inputs — rather than its prose. It closes a gap that is hard to forget once you have seen it: an agent can compose a warm, well-phrased refusal while its action log shows it went and fetched the forbidden record anyway. Grade the prose, and you have graded the prose.

The transcript is what the agent said. The trace is what the agent did. Only one of those is behaviour.

Where a model is still allowed in the room

I am not against using a model here. I am against letting it decide. In my pipeline an LLM does appear — after the deterministic verdict is in, to explain a failure in readable language. The code decides; the model narrates. That division of labour is the entire design: it explains, it never decides.

What the bench actually reports

The corpus is thirty attack cases written in German, the language the agent’s users would speak, across fourteen categories and four families: prompt injection, data exfiltration, authority spoofing, and GDPR-rights abuse — the last being a data grab dressed in the language of a real legal right.

The reference run reports 26 of 29 deliverable attacks passed — deliverable meaning the attacks that run could actually put in front of the agent — and three genuine weaknesses in my own agent.

I published the three. That is not modesty; it is the only thing that keeps the number meaningful. And notice what makes “three” a number at all: a deterministic verifier. Had a model graded that run, “three” would have been a mood, and next week it might have been one, or five, or none.

Determinism is what turns a finding into a fact that survives being handed to somebody else.

Your next step

You do not need a bench like mine to start this week.

  1. Plant one canary. Put a distinctive string in a record your agent must never disclose.
  2. Write one check in code that fails if the string appears in any answer. Five lines is plenty.
  3. Add one trace assertion. Where your agent should refuse, assert that the fetch action never ran — not merely that the reply sounded like a refusal.
  4. Delete the judge prompt you were about to write. If you keep a model in the loop, keep it downstream of the verdict.

The pupil is not lying to you when they mark their own paper. They simply cannot mark it. That is not a character flaw, in a child or in a model — it is what marking is for.

Mustafa Aksu

Salesforce developer & ISV builder focused on Revenue Cloud, Agentforce, and Data Cloud. I write from real, shipped work.