An admin friend asked me a fair question about a hosted MCP setup I had running: “Can you show me what it did last week?”
I opened the org expecting to walk him through it in five minutes. I could answer him comfortably twice. The third time, I found myself opening an Apex class to work out the answer — and that pause is what this post is about, because the pause was the finding.
If MCP is new to you: the Model Context Protocol is an open standard that lets an AI client call tools your Salesforce org exposes. Salesforce’s hosted MCP servers make those tools available over an authenticated connection, so a model can query records or run your own Apex on request.
The three questions an audit has to answer
Strip away the vocabulary and an audit of an MCP rollout is three questions, in this order:
- Which tool was called?
- With which arguments?
- Under whose permissions did it run?
Answer all three and you can reconstruct any event. Answer only the first two and you have a log that describes activity without establishing accountability — which was the shape of my own rollout at the moment my friend asked. The difficulty is not evenly spread, so take them one at a time.
Question one: which tool was called
This one is easy, for a structural reason worth understanding.
An MCP server does not expose your whole org. It exposes a tool list — a named, finite, curated set of things a model may ask for, and that list is configuration you chose. So “which tool ran” is a small answer from a closed set, not an investigation.
Which is the first argument for keeping tool lists short and purpose-shaped. A server offering a handful of tools to one team produces an audit trail a human can read. A server offering everything to everyone produces a log that is technically complete and practically useless.
Question two: with which arguments
Also answerable, and more revealing than people expect. The tool tells you the kind of action; the arguments tell you the subject — which record, which account, which date range, which search string. “Ran the order lookup” is activity. “Ran the order lookup for this one customer, over and over, inside a few minutes” is a story, and the story is what you need when somebody asks what happened.
The habit that makes this useful is unglamorous: read the trail during a quiet week, not during an incident. Logs you only open in a crisis are logs you cannot interpret, because you never learned what normal looks like. Twenty years in schools taught me the same about attendance registers — the pattern only means something if you have been reading it all term.
Question three: under whose permissions — and the gap
Here is where my five-minute walkthrough stalled.
The first half of the answer is genuinely good news. Every hosted MCP call runs as the authenticated user. Each person connects through their own login, so there is no anonymous service account sitting between the model and your data. The trail carries a real name, not “the AI did it.”
But “runs as the user” is a statement about identity, not automatically about enforcement. Those are two different things, and the difference is where rollouts get into trouble.
Two fences, and only one of them is on the diagram
Think of it as two fences.
The connection fence: scopes. The External Client App that brokers the OAuth connection is granted a set of scopes — named slices of capability. If a capability was never granted, no tool call can use it, however the model phrases the request.
The data fence: the Apex. For a custom tool, the code behind it decides what data is actually touched, and Apex can run in two very different ways. In user mode, the platform enforces the running user’s object and field permissions on every query and write. In system mode — Apex’s default in many contexts — it does not. The code sees everything, because nobody asked it to check.
And now the sentence I would put on the wall of any MCP rollout:
A tool whose Apex runs in system mode is not constrained by a narrow scope, because the Apex was never asked to enforce anything.
Read that again if you are new to this, because it is counter-intuitive. You can grant a beautifully minimal scope, connect a named user with a modest profile, feel appropriately careful — and still have a tool that reads a field that user cannot see, because inside the class the enforcement was never requested. The scope bounded the door. It said nothing about the room.
The fix is small and specific: write your SOQL and DML WITH USER_MODE inside the Apex behind any tool. Then the user’s own permissions bind the data, and “runs as the user” becomes a fact about enforcement rather than about the login screen.
One related warning, because it catches experienced admins: hiding a field on a page layout is discoverability, not authorisation. It changes what a person happens to see on a screen, not what a tool acting as them can read.
Measuring the gap instead of guessing at it
The uncomfortable part is that you cannot see this gap in the tool list, the scopes, or the audit trail. It lives inside the code, and it looks like nothing.
That is the sort of thing a machine should read for you. I built a static analyser for it — static meaning it reads the code as text without running anything, so it costs no credits and touches no data. It measures how far an agent’s code can reach beyond its running user, and it is MIT-licensed and open at github.com/aksumustafa1625/agent-blast-radius.
You do not need a tool to begin, mind you. Open every Apex class behind an MCP tool and search it for USER_MODE. The classes where you find nothing are your list.
What good looks like in practice
If I were writing the runbook for a team, it would be four lines.
- Keep each server’s tool list short enough that a human can read a week of logs in one sitting.
- Connect every person through their own login. No shared user, ever, however convenient it looks on a Friday.
- Put
WITH USER_MODEin every custom tool’s SOQL and DML, and treat a class without it as unfinished. - Read the trail on day three and again on day seven of the pilot, while everything is still fine.
Two honest limits. Admin labels and settings here are still moving between releases, so treat my wording as the shape of the thing and check the current documentation in your own org. And no amount of scoping replaces judgement: for anything consequential, keep a human confirmation step in the client.
Your next step
Take one tool you have already exposed and answer the three questions out loud, in order. Which tool. With which arguments. Under whose permissions.
If the third answer requires you to open a class and read it — as mine did — that is not a failure. That is the audit working, a week before anyone official asks.