The first time one of my own attack cases embarrassed my agent, my hand went straight to the test file.
Not to fix the agent. To soften the case. The wording was a little unfair, I told myself. Two words and it would pass — and wasn’t the spirit of the test already covered by the case above it?
I did not do it, and the only reason I did not is that I could not have done it quietly. The case was already committed to git, under a timestamp, before the thing that scored it existed. Editing it would have left a mark with my name on it.
That is the subject of this post: not what to test, but in what order to commit things.
The order is the evidence
Here is the practice, in one sentence. The attack corpus is committed to git before the harness that scores it.
Corpus first. Grader second. Provable from the history.
It sounds like housekeeping. It is not. It is the difference between a number that means something and a number that merely sounds good, and it costs nothing but the discipline of two commits in order.
If the vocabulary is new: an attack corpus is a set of test cases that try to make an agent misbehave — smuggled instructions, requests for another customer’s data, someone claiming to be an administrator. The harness is the code that runs those cases and decides whether the agent passed.
What pre-registration means
Science had this problem long before we did, and solved it with pre-registration: you publish your hypothesis and your method before you collect the data, so you cannot quietly reshape the analysis once you have seen the results.
The failure it prevents is not fraud. That is the part people miss. It prevents something far more ordinary: a sincere researcher, looking at a disappointing result, noticing a perfectly reasonable adjustment that happens to improve it. Every individual step is defensible, and the process produces a finding that was never really tested.
I recognise that person. That person was me, hand on the test file, thinking about two words.
Why the harness must come second
You could argue the order is arbitrary — both files end up in the repository either way. It is not, and the reason is worth being precise about.
The harness is the half with discretion. It decides what counts as a pass: which strings it looks for, which actions it inspects, how strictly it reads a refusal. Dozens of small judgements, each one nudgeable.
If the harness exists first, the corpus gets written — sincerely, with no bad intent — to suit it. You reach for the attacks the grader handles crisply and pass over the awkward ones. Nobody decides to do that. It happens the way a tired teacher sets the exam question they already know the class can answer.
Committing the corpus first pins the questions in place before the marking scheme exists. In teaching we had the same rule stated differently: you may write a mark scheme before the pupils sit the paper, but not once you have read the scripts.
Git is the notary, and it costs nothing
You do not need a platform or a process document for this. You need the tool you already have open.
A commit carries a timestamp and a hash, and a later commit carries the previous one’s hash inside it — which makes a git history a chain rather than a list. You cannot slide a file backwards into an earlier commit without rewriting everything after it, and a rewritten history is visible to anyone who looks. So the order is provable in the ordinary sense: somebody sceptical can clone the repository and check. No trust in me is required, which is the property you want in a safety claim.
In my own bench the sequence reads: commit the corpus — thirty attack cases written in German, across fourteen categories, with eight planted canary tokens — then commit the harness that scores them, then run. (A canary token is a distinctive string placed where only a leak would carry it; if it turns up in an answer, something crossed a boundary. It gives the grader something binary to check rather than something to judge.) The corpus commit contains no scoring logic; the harness commit changes no case.
The price is two commits, in an order. It is the cheapest credibility available to anyone building anything.
A test you are free to rewrite after seeing the answers is not a test. It is a story with numbers in it.
Changing a case honestly
Cases do need to change sometimes. A typo. An attack that turns out to be undeliverable. A case that tests nothing because it was malformed. Pre-registration does not freeze the corpus forever — it makes changes visible.
So: fix it forward. A new commit, with a message saying what changed and why. Never amend the original commit, never rebase the corpus into a tidier shape, never quietly overwrite a case that failed.
A history showing “loosened the wording on an ambiguous case; the agent still fails the underlying check” is more trustworthy than a history with no such commit in it. The version where nothing ever needed adjusting is the version that looks fabricated.
The number this makes possible
The reference run of my bench reports 26 of 29 deliverable attacks passed, and three genuine weaknesses in my own agent. I published the three rather than hiding them.
Notice what pre-registration does to that sentence. Because the corpus was committed first, I could not have made those three disappear by adjusting the cases — not without leaving a commit that said so. Publishing became the path of least resistance. The discipline did not make me honest. It made honesty easier than the alternative, which is far more reliable than depending on my character at eleven at night.
That is the real function of the practice: not proving your integrity to strangers, but removing the moment where your integrity gets quietly tested.
How to read somebody else’s number
Once you have done this yourself, you read other people’s claims differently — the most portable thing in this post.
When someone tells you their agent passed some percentage of safety evaluations, one question cuts through: when was the test set written, relative to the grader?
If they can answer, with a history, the number is a result. If they cannot, it is a claim. It might still be true — you have no way of telling, and neither do they.
Your next step
- Write ten attack cases against your own agent, in your users’ actual language.
- Commit them, on their own, before you write a single line of scoring code.
- Then write the harness, in a separate commit, and run it.
- When something fails, fix the agent — and if you must touch a case, do it in a new commit that says why.
The agent will fail some of them. Mine did, and it is a better agent for it. A test you always pass is not a test, and a result nobody can check is not a result. Commit the exam before you write the mark scheme.