← Back to blog

A Runtime Where an Agent Cannot Call a Tool Around the Controls

"Agents propose. Cradle decides what may execute" sits on the product's front page. As long as a single code path could call a tool directly, that sentence was a recommendation, not a rule. Here is how the runtime is built and how I closed the side doors mechanically rather than by convention.

Three separations

Cradle separates probabilistic reasoning from deterministic control of execution. The whole construction rests on three statements:

A proposal is not a decision. A tool's output is not verified truth. Successful execution is not verification of the result.

The risk assessment I wrote about earlier governs replies to people. The runtime generalises it to governing actions on systems.

Four entities instead of one "agent action"

In a typical agent loop a "tool call" is one entry in the transcript. In the runtime it is four separate tables: proposals, decisions, observations, verifications, plus an append-only audit_events.

EntityWhat it recordsWho writes it
Proposalwhat the agent wants to do: action type, target, arguments, side-effect class, pre- and postconditionsthe agent
Decisionthe verdict allow / deny / repair / require_approval and the authorised identitypolicy
Observationwhat the tool actually returnedthe runtime, after the call
Verificationwhether the postconditions hold, read from an independent sourcea verifier

The separation is not there to make the schema pretty. It lets the invariants be enforced with cheap checks over records, with no graph and no reasoner:

  • a denied proposal cannot produce an observation (EXECUTION_NOT_AUTHORIZED);
  • the executor must match the authorised identity (IDENTITY_MISMATCH);
  • one idempotency key, one proposal; one proposal, one decision;
  • the verification status is derived from the checks, not set by the caller: an empty list of checks yields inconclusive, because "we did not look" is not "all is well".

The last point is the whole thesis in one line of code:

const verification = recordVerification(db, {
  observationId: observation.id,
  checks: [{ name: 'refund.recorded', status: 'fail', source: 'ledger-api' }]
})

verification.status // 'failed' — even though the tool returned 200

A failed verification comes back as a failure even when the tool reported success. The source field on every check records that an independent source was read, not the same call that did the writing.

Audit as a by-product

The log is not layered on top of the runtime. Every proposal, decision, observation and verification record produces an audit event carrying payload_hash, prev_hash and hash. verifyAuditChain() catches a modified payload, modified metadata and a deleted row (a gap in seq).

This is what I mean by assurance by construction: you cannot execute an action without leaving a trace, because the trace is not a separate logger call but part of the same write path. One pitfall along the way: canonicalJson() sorts keys at every level. Without it the chain broke on a round-trip through SQLite: JSON.stringify preserves insertion order, so a freshly built object serialised differently from one read back from the database.

ToolGateway: the single door

On top of the existing MCP client sits ToolGateway.execute(): propose → decide → invoke → observe. A policy denial is a result, not an exception: the caller gets the decision with the full list of reasons, and the transport is never touched.

import { ToolGateway, registerTool, mcpInvoker } from './core/runtime'

registerTool(db, {
  serverId: 'mcp_billing',
  toolName: 'refund.issue',
  sideEffectClass: 'compensatable',
  requiredScopes: ['billing:write'],
  allowedEnvironments: ['staging', 'production'],
  dryRunToolName: 'refund.validate'
})

const result = await new ToolGateway(mcpInvoker).execute(db, {
  taskId, agentId,
  identity: { id: 'user:operator_7', scopes: ['billing:write'] },
  serverId: 'mcp_billing',
  action: {
    actionType: 'refund.issue',
    target: { kind: 'order', id: 'A-119', env: 'production' },
    arguments: { amount: 4200 },
    sideEffectClass: 'compensatable',
    preconditions: ['order.exists', 'order.paid'],
    postconditions: ['refund.recorded']
  },
  idempotencyKey: 'refund:A-119'
})

result.decision.verdict  // allow | deny | repair | require_approval
result.observation       // null unless allow — the tool was never invoked
result.replayed          // true when the idempotency key was already used

Two decisions inside the gateway deserve to be named.

Discovery is not authorization. An MCP server reports a tool's name and, at best, a destructiveHint. It does not say what the tool can touch or who may call it. So an unregistered tool is never invoked, and when registration is automatic a new tool is classed irreversible unless the server explicitly says otherwise. A remote server's optimism about its own safety is no reason to lower the bar locally.

A human approval is an object, not a string. The gateway used to accept approvalId as a string and trust it; any caller satisfied the rule by inventing a value. Now it is a record with its own lifecycle: it points at a specific proposal or scope, it has an expiry, it is spent once, and an agent cannot approve itself.

The guard in CI

The gateway on its own guarantees nothing while a direct getMcpManager().callTool() exists next to it. The docstring of gateway.ts says so outright: without a single chokepoint the invariants are advisory.

Checking the code turned up three bypasses. One was not MCP at all but eleven in-process tools against the product's own database, four of them destructive. Another, which I had classified as read-only from line 80, wrote through the same helper: it updated notes and marked them completed. Shallow classification was a lesson in itself.

Closing the doors took three steps. Types: getMcpManager() returns a connection pool that has no callTool. Code: each bypass is either registered in the registry with a side-effect class and scopes, or explicitly taken out from under the gateway with the reason written in the code. And a mechanical guard: the script scripts/check-gateway.mjs, also known as pnpm check:gateway.

Why a script and not a grep. The rule is about where a symbol is used, not whether a string appears: callTool legitimately lives inside the runtime and inside the MCP client. The guard checks three things: a call outside src/core/runtime/, an import of any symbol that can reach a tool, and the set of exports from src/core/mcp/, which must stay exactly two named exemptions. The one exemption (fetching a note's screenshot for the operator UI) must keep its tool name as a string literal; widen its signature to take an arbitrary name and the guard fails.

The guard was tested on a deliberate violation; otherwise it is a wish, not a guard:

✗ src/core/__probe/violation.ts:2 — calls .callTool() outside src/core/runtime/
→ ПРОВАЛ: 1 нарушений в 234 файлах вне src/core/runtime/

On clean code: zero violations across 233 files, and the runtime tests are green.

Honest status

What does not exist, and must not be described as working:

  • Ready verifiers. The engine exists; the registry is empty. Until something is registered, any action with postconditions ends in inconclusive → aborted. That is the right default, and it is of no use yet.
  • Rollback in the sense of restoring a snapshot. There is only compensation through a tool the server itself declared.
  • Domain packs as a deliverable. The manifest exists; a distribution format for on-premise installations does not.
  • Tool use for the channel agent. Triage replies with text; tools live in the sandbox and in domain packs. Whether the agent attached to a channel needs tool calls through the gateway is a product question, and it is open.

The fallback path, if you do not want an agent loop of your own at all: an external agent such as Claude Code, started inside the perimeter with cradle launch claude against the Anthropic-compatible /v1/messages on cradle-server. Then Cradle does not build the agent; it only stands between the agent and the tools. That was the original brief.

This article was created in hybrid human + AI format. I set the direction and theses, AI helped with the text, I edited and verified. Responsibility for the content is mine.

← Back to blog