The number that made me trust this setup was 88.

I had been building out a hospital demo org — a fictional clinic in a real Developer Edition org — with a loop that deploys metadata, deploys Apex, runs the tests, and reads what comes back. On one pass, the runner reported a trigger handler at 88% coverage. Under target.

Nothing dramatic happened. The loop did not shrug, and it did not announce that the deployment had gone well and everything looked fine. It read which lines were uncovered, wrote a test for that path, and re-ran until the handler was green.

A small moment — and the whole argument for building this way, so let me unpack it slowly.

What “Apex as an MCP tool” actually means

MCP — the Model Context Protocol — is the open standard for letting a model call tools that live somewhere else. Salesforce Hosted MCP Servers publishes tools from inside your org, and an External Client App authorises an outside client to connect. Here, that client is Claude.

The tools themselves are the least exotic part. Each is an Apex method carrying the @InvocableMethod annotation — ordinary Apex, with ordinary unit tests, deployed the ordinary way. It does not become special because a model calls it; it simply gets a new kind of caller.

I have written before that testing an MCP tool is the same Apex with a sharper contract. This is the other half of that idea.

Six tools, all about deployment

There are six custom tools in this org, and they share a theme: they are deployment tools. Not “look up a patient”, not “summarise a record” — tools for the mechanical work of getting metadata and code into an org and learning whether it survived.

That choice was deliberate. Deployment work has a quality most agent tasks do not: every step has an objective verdict. The deploy succeeded or it returned errors. The tests passed or they failed. Coverage is a number. There is very little to interpret, and therefore very little to talk your way past.

The repo is public, tools and tests included: github.com/aksumustafa1625/hospital-org-mcp.

The loop

Written out, it is unglamorous, and that is a compliment:

  1. Deploy the metadata, in dependency order — objects before the fields on them, fields before the layouts that reference them, and so on.
  2. Deploy the Apex.
  3. Run the tests.
  4. Read the actual failures.
  5. Fix them.
  6. Go back to step one.

The ordering in step one is not a detail. Metadata deployment fails in confusing, cascading ways when a dependency is missing, and half the errors you get back are downstream noise from the first genuine problem. Dependency order removes that whole category — which means every error the loop does see is far more likely to be real.

The stop condition must be one the model cannot argue with

Here is the design decision I would keep in any loop I build again.

The loop stops when the test runner says all green at or above 90% coverage. Not when the work looks complete. Not when the model believes it has finished. The runner decides.

This matters because a language model is very good at producing a satisfying summary. Left to judge its own work, it will write a confident paragraph about what was accomplished — and that paragraph reads just as well whether or not the org ended up in the state you wanted. If the stop condition is “the model thinks it is done”, you have built a loop that terminates on optimism.

Coverage and pass/fail are not opinions. They come from outside the conversation, and nothing said inside it changes them.

A loop needs a stop condition its own author cannot negotiate with. If the model decides when the work is finished, the work is finished exactly when the model gets tired.

That is why 88 was such a useful number. It was under target, there was no way to round it up, and the next action was decided by the runner rather than by a judgement call.

Read the actual failure, not the failure you expected

The step people skip is number four, and it separates a loop from a slot machine.

When a test fails, the message tells you which assertion broke, on which line, with which values — genuinely more informative than anything you could guess from outside. Yet the tempting move, for a person as much as for a model, is to note that something failed, form a hypothesis about why, and fix the hypothesis.

Sometimes the hypothesis is right. When it is wrong, you have changed working code for no reason and the original failure is still sitting there, now harder to see.

With the 88% handler, the runner did not merely say “coverage too low”. It said which lines had never executed — a shopping list. The fix was not a guess about what might be untested; it was a test written for the path the report named. One pass, then green.

When the error is opaque, enumerate — do not retry

The other moment worth reporting is uglier. A deploy came back with an opaque script exception: the kind of platform error that tells you something went wrong and almost nothing about what.

The obvious move is to try again. Most of the time you burn a minute, get the identical message back, and try once more anyway — because retrying feels like progress.

What the loop did instead was enumerate candidate root causes: the plausible reasons a deploy of this shape produces an error of this shape, listed out so they could be checked one at a time. Turning one unhelpful message into a short list of testable explanations is not a clever trick. It is the same discipline as reading a failure properly, applied to a case where the failure refuses to say much.

Repeating a failing call is not a diagnosis. If you cannot say what changed between two attempts, you have not made two attempts — you have made a wish, twice.

Where this pattern stops working

The pattern is easy to over-sell, so here is the boundary. It works because deployment produces an objective verdict at every step. Point the same loop at work where “done” is a matter of taste — is this data model elegant — and the confident-summary problem walks straight back in, because there is no runner to overrule anyone.

The loop is only ever as trustworthy as its verdict. Where an external checker already exists, the pattern is remarkably solid. Where none exists, build one first, or keep a person in the seat.

Your next step

Take a job you already do by hand — a deploy-and-test cycle in a scratch or Developer Edition org is the perfect candidate — and write down its stop condition in one sentence before you automate anything.

If that sentence contains the words “looks”, “seems”, or “should be”, you do not have a stop condition yet. You have a hope. Replace it with something a runner can print: all tests pass, coverage at or above a number you chose in advance, zero deployment errors. Then let the loop run until the runner agrees.

Mustafa Aksu

Salesforce developer & ISV builder focused on Revenue Cloud, Agentforce, and Data Cloud. I write from real, shipped work.