I was reading back over one of my own MCP tools recently — an Apex method I had exposed for a model to call — and I found a parameter I had been quietly proud of when I wrote it. A mode string. Pass "read" and it fetched. Pass "update" and it wrote.

It felt tidy at the time. One tool instead of two. Less metadata, less to maintain.

Then I asked the question I now ask about everything a model can call: what happens if the caller gets that argument wrong? The answer was that a single mis-chosen word turns a lookup into a write. Not a rejected call. Not an error I could catch. A different action entirely, executed exactly as designed.

I split the tool in two that evening. This article is why.

First, the vocabulary

Briefly, so nobody is lost. MCP — the Model Context Protocol — is the open standard that lets an AI model call tools living somewhere else, such as inside your Salesforce org. On Salesforce Hosted MCP Servers, a custom tool is an Apex method carrying the @InvocableMethod annotation, published through an McpServerDefinition.

The part beginners underestimate is who the caller is now. Not a button. Not a Flow you wrote. A language model, choosing on its own whether to call your method at all, and what to pass it — from your parameter names and descriptions alone. There is no meeting where you explain your intent. Everything below follows from that one fact.

A tool a model cannot misuse

Here is the design rule I have settled on, and it fits in five words: one tool, one verb.

A tool that does one thing, named after the thing it does, is a tool a model cannot misuse. It can call it at the wrong moment — a different problem, solved with instructions and permissions — but it cannot make it do something other than what its name says. The set of behaviours the tool has is one.

A tool that takes a mode flag and branches has a different property: its behaviour depends on the model getting an argument right. You have moved a decision you could have made at design time, in metadata you control, into an inference the model makes in the moment. That is a strange trade to make on purpose, and I made it without noticing.

A branching tool does not fail when the model guesses badly. It succeeds — at the wrong thing.

Compare the two shapes:

// Before: one tool, two behaviours, decided by an argument
// mode = 'read' | 'update'
@InvocableMethod(label='Manage Reseller')
public static List<Result> manageReseller(List<Request> reqs) { ... }
// After: two tools, each with a single verb in its name
@InvocableMethod(label='Get Reseller Status')
public static List<Status> getResellerStatus(List<Lookup> reqs) { ... }

@InvocableMethod(label='Update Reseller Status')
public static List<Result> updateResellerStatus(List<Change> reqs) { ... }

The second version is more metadata and more lines. It also has something the first could never have: a read tool structurally incapable of writing.

The audit trail argument, which I find even more persuasive

Now the part I did not expect to care about as much as I do.

When you review what an agent did — because something went wrong, or because someone asked you to prove nothing did — you are reading a list of tool calls. With narrow tools, that list is a sentence in English:

getResellerStatus → getResellerStatus → updateResellerStatus

You can read that from across the room. Something was looked up twice and then changed once. Nobody has to open the arguments to understand what happened.

With the branching version, the same activity reads:

manageReseller → manageReseller → manageReseller

Three identical lines. To learn whether anything was written, you must open each call and inspect an argument. Multiply that by a busy day and you have an audit trail that technically contains the truth but does not communicate it — which in practice means nobody reads it.

Narrow tools make the audit trail readable, because the tool name tells you what happened. That alone would justify the split even if the safety argument did not exist.

Where to draw the line around “one thing”

The rule is easy to over-apply, so let me be concrete about what “one thing” means.

One thing is one verb against one subject. Get a reseller’s status. Update a reseller’s status. Create an onboarding case. Three tools, and each name is a complete description of the effect.

One thing is not “one field.” Splitting an update tool into six tools, one per field, gives the model six chances to pick the wrong one and gives you an audit trail nobody can read for the opposite reason: it is too long. The point is not maximal fragmentation. It is that each tool has exactly one effect a reader could name.

The clearest test I know: write the tool’s name without using “and” or “or”. If you cannot — “search and update”, “get or create” — you are looking at two tools wearing one coat.

Two more places the line lands, from my own splitting:

  • Read and write never share a tool. I would enforce this one even when it is inconvenient. The difference between fetching and changing is the difference a reviewer cares about most, and it belongs in the tool’s name, not in an argument’s value.
  • A flag that only changes formatting is usually fine. A parameter that decides what the tool does is a second tool. A parameter that decides how much detail comes back is just a parameter.

The objection: does this not give the model too many tools?

It is a fair worry, and the answer is discipline about names rather than a return to flags. Since the model chooses from names and descriptions alone, the risk is not the count — it is two tools whose descriptions could each plausibly answer the same request. So when you split, make the names diverge sharply.

And a small honesty from my own build: splitting a tool made its tests better, not merely more numerous. A tool with one behaviour has one success path, one empty-input path, one error path and a bulk path. A tool with a mode flag has all of those twice, plus a category I had never tested at all — what happens when the mode value is neither word I expected.

Your next step

Open your list of MCP tools and read only the names, ignoring the code entirely.

For each one, ask: can I say what this tool does in one verb, with no “and” and no “or”? Where the answer is no, look at the parameters and find the one that is really a hidden branch. Split it, give each half a name that says its verb out loud, and check whether each description now writes itself in one sentence. If it does, you have made a tool that is easier for a model to choose correctly and easier for a human to audit afterwards — for the same reason.

If MCP is new to you, begin at What is MCP?, then your first custom Apex tool. To prove a tool behaves, testing your MCP tools is the companion piece. The series lives in MCP.

Mustafa Aksu

Salesforce developer & ISV builder focused on Revenue Cloud, Agentforce, and Data Cloud. I write from real, shipped work.