A Runtime Where an Agent Cannot Call a Tool Around the Controls
"Agents propose. Cradle decides what may execute" sits on the product's front page. As long as a single code path could call a tool directly, that sentence was a recommendation, not a rule. Here is how the runtime is built and how I closed the side doors mechanically rather than by convention.
Three separations
Cradle separates probabilistic reasoning from deterministic control of execution. The whole construction rests on three statements:
A proposal is not a decision. A tool's output is not verified truth. Successful execution is not verification of the result.
The risk assessment I wrote about earlier governs replies to people. The runtime generalises it to governing actions on systems.
Four entities instead of one "agent action"
In a typical agent loop a "tool call" is one entry in the transcript. In the runtime it is four separate tables: proposals, decisions, observations, verifications, plus an append-only audit_events.
| Entity | What it records | Who writes it |
|---|---|---|
| Proposal | what the agent wants to do: action type, target, arguments, side-effect class, pre- and postconditions | the agent |
| Decision | the verdict allow / deny / repair / require_approval and the authorised identity | policy |
| Observation | what the tool actually returned | the runtime, after the call |
| Verification | whether the postconditions hold, read from an independent source | a verifier |
The separation is not there to make the schema pretty. It lets the invariants be enforced with cheap checks over records, with no graph and no reasoner:
- a denied proposal cannot produce an observation (
EXECUTION_NOT_AUTHORIZED); - the executor must match the authorised identity (
IDENTITY_MISMATCH); - one idempotency key, one proposal; one proposal, one decision;
- the verification status is derived from the checks, not set by the caller: an empty list of checks yields
inconclusive, because "we did not look" is not "all is well".
The last point is the whole thesis in one line of code:
const verification = recordVerification(db, {
observationId: observation.id,
checks: [{ name: 'refund.recorded', status: 'fail', source: 'ledger-api' }]
})
verification.status // 'failed' — even though the tool returned 200
A failed verification comes back as a failure even when the tool reported success. The source field on every check records that an independent source was read, not the same call that did the writing.
Audit as a by-product
The log is not layered on top of the runtime. Every proposal, decision, observation and verification record produces an audit event carrying payload_hash, prev_hash and hash. verifyAuditChain() catches a modified payload, modified metadata and a deleted row (a gap in seq).
This is what I mean by assurance by construction: you cannot execute an action without leaving a trace, because the trace is not a separate logger call but part of the same write path. One pitfall along the way: canonicalJson() sorts keys at every level. Without it the chain broke on a round-trip through SQLite: JSON.stringify preserves insertion order, so a freshly built object serialised differently from one read back from the database.
ToolGateway: the single door
On top of the existing MCP client sits ToolGateway.execute(): propose → decide → invoke → observe. A policy denial is a result, not an exception: the caller gets the decision with the full list of reasons, and the transport is never touched.
import { ToolGateway, registerTool, mcpInvoker } from './core/runtime'
registerTool(db, {
serverId: 'mcp_billing',
toolName: 'refund.issue',
sideEffectClass: 'compensatable',
requiredScopes: ['billing:write'],
allowedEnvironments: ['staging', 'production'],
dryRunToolName: 'refund.validate'
})
const result = await new ToolGateway(mcpInvoker).execute(db, {
taskId, agentId,
identity: { id: 'user:operator_7', scopes: ['billing:write'] },
serverId: 'mcp_billing',
action: {
actionType: 'refund.issue',
target: { kind: 'order', id: 'A-119', env: 'production' },
arguments: { amount: 4200 },
sideEffectClass: 'compensatable',
preconditions: ['order.exists', 'order.paid'],
postconditions: ['refund.recorded']
},
idempotencyKey: 'refund:A-119'
})
result.decision.verdict // allow | deny | repair | require_approval
result.observation // null unless allow — the tool was never invoked
result.replayed // true when the idempotency key was already used
Two decisions inside the gateway deserve to be named.
Discovery is not authorization. An MCP server reports a tool's name and, at best, a destructiveHint. It does not say what the tool can touch or who may call it. So an unregistered tool is never invoked, and when registration is automatic a new tool is classed irreversible unless the server explicitly says otherwise. A remote server's optimism about its own safety is no reason to lower the bar locally.
A human approval is an object, not a string. The gateway used to accept approvalId as a string and trust it; any caller satisfied the rule by inventing a value. Now it is a record with its own lifecycle: it points at a specific proposal or scope, it has an expiry, it is spent once, and an agent cannot approve itself.
The guard in CI
The gateway on its own guarantees nothing while a direct getMcpManager().callTool() exists next to it. The docstring of gateway.ts says so outright: without a single chokepoint the invariants are advisory.
Checking the code turned up three bypasses. One was not MCP at all but eleven in-process tools against the product's own database, four of them destructive. Another, which I had classified as read-only from line 80, wrote through the same helper: it updated notes and marked them completed. Shallow classification was a lesson in itself.
Closing the doors took three steps. Types: getMcpManager() returns a connection pool that has no callTool. Code: each bypass is either registered in the registry with a side-effect class and scopes, or explicitly taken out from under the gateway with the reason written in the code. And a mechanical guard: the script scripts/check-gateway.mjs, also known as pnpm check:gateway.
Why a script and not a grep. The rule is about where a symbol is used, not whether a string appears: callTool legitimately lives inside the runtime and inside the MCP client. The guard checks three things: a call outside src/core/runtime/, an import of any symbol that can reach a tool, and the set of exports from src/core/mcp/, which must stay exactly two named exemptions. The one exemption (fetching a note's screenshot for the operator UI) must keep its tool name as a string literal; widen its signature to take an arbitrary name and the guard fails.
The guard was tested on a deliberate violation; otherwise it is a wish, not a guard:
✗ src/core/__probe/violation.ts:2 — calls .callTool() outside src/core/runtime/
→ ПРОВАЛ: 1 нарушений в 234 файлах вне src/core/runtime/
On clean code: zero violations across 233 files, and the runtime tests are green.
Honest status
What does not exist, and must not be described as working:
- Ready verifiers. The engine exists; the registry is empty. Until something is registered, any action with postconditions ends in
inconclusive→aborted. That is the right default, and it is of no use yet. - Rollback in the sense of restoring a snapshot. There is only compensation through a tool the server itself declared.
- Domain packs as a deliverable. The manifest exists; a distribution format for on-premise installations does not.
- Tool use for the channel agent. Triage replies with text; tools live in the sandbox and in domain packs. Whether the agent attached to a channel needs tool calls through the gateway is a product question, and it is open.
The fallback path, if you do not want an agent loop of your own at all: an external agent such as Claude Code, started inside the perimeter with cradle launch claude against the Anthropic-compatible /v1/messages on cradle-server. Then Cradle does not build the agent; it only stands between the agent and the tools. That was the original brief.