Sari la conținut

Guardrails — one policy for every mutating action

Acest conținut nu este încă disponibil în limba selectată.

Agents in Control Center can delete files, push branches, publish pull requests and call external APIs. Each of those is a different kind of risk, and for a long time only one of them — shell commands — had a policy. The guardrails exist to answer one question uniformly, for every mutating action an agent can take, regardless of which runtime or tool proposes it: is this allowed, does it need an approval, or is it denied?

There is exactly one policy engine. It is consulted at every point where an action can leave the system: the built-in harness’s tool registry, the server-side MCP dispatcher gate and repo-op mutations. A rule written once therefore applies to the built-in runtime and to external agent CLIs alike — there is no second, quieter path around it. The older bash-only command policy still works: a rule can match a command prefix instead of an action class and the two forms coexist in the same store.

The engine is not, however, a substitute for the OS sandbox and the sandbox is not a substitute for it. They answer different questions: the guardrails decide whether a call Control Center can see is allowed to happen at all, while the sandbox bounds where the resulting process may reach. cc_server applies the sandbox where the host offers a backend — but Windows offers none and Linux needs bwrap and socat. On those hosts the guardrails are the outermost enforcement layer rather than a policy sitting on top of a floor, which is worth knowing when you decide how permissive to be.

ActionClass: a deliberately closed taxonomy

Section titled “ActionClass: a deliberately closed taxonomy”

Every mutating tool declares which ActionClasses it belongs to, from a closed set of thirteen effect classes. Five of them prompt out of the box; the rest allow and a rule is what changes that.

Class Built-in default What it covers
fileDelete prompt Deleting a file (a delete-shaped write, rm)
gitPush prompt Pushing to a git remote
prCreate prompt Creating a pull request
prPublish prompt Publishing a review, merging a PR
vendorSyncWrite prompt Writing to an external ticket vendor (Linear, Jira, GitHub sync)
fileWriteOutsideWorktree allow Writing a file outside the isolated worktree
gitCommit allow Creating a git commit
networkEgress allow Network egress (web fetch, arbitrary HTTP)
secretAccess allow Reading a secret or credential
packageInstall allow Installing a package (npm, pip, pub, brew)
processSpawn allow Spawning a process (shell and anything shelling out)
workspaceMutation allow Mutating workspace structure (repos, spaces, agents)
enclosureControl allow Driving an enclosure (rig): booting a VM, sending it input

The eight allow-by-default classes were chosen on the assumption that the sandbox floor applies beneath them. It does where the host offers a backend, and it does not on Windows or on a Linux box missing bwrap and socat — so if you want fileWriteOutsideWorktree or packageInstall gated there, that is a rule you have to write.

The set is closed on purpose. A new mutating tool that declares no class fails the ratchet test, so the taxonomy can only grow by an explicit, argued decision — “taxonomy sprawl is the death of this feature.” Thirteen classes is enough to express the policies people actually write (never push without asking, never delete outside a worktree, always allow reads) and few enough that the settings UI stays comprehensible.

A rule maps (scope, ActionClass | command prefix) → allow | prompt | deny. Scopes nest from most to least specific:

space > agent > workspace > mode preset > built-in default

Resolution is specificity first: the first scope with a matching rule decides and resolution stops there. Most-restrictive is only a tie-break — between two equally specific rules within one scope and when combining the several classes a single action declares. Within a scope, a command rule matches on the longest command prefix. Allow does not beat deny anywhere.

Two decisions carry hard guarantees:

  • Fail-closed prompts. A prompt with no approver connected resolves to denied. An unanswered question is never a yes.
  • Most-restrictive combination. An action spanning several classes combines them most-restrictively; when several classes prompt, one confirmation lists them all.

There is deliberately no per-turn latch. Every tool call re-resolves the policy from scratch, so repeated prompting for the same action is bounded by the operator’s own “remember this decision” rather than by anything turn-scoped.

A rule can constrain the ARGUMENTS of an action, not just its verb. “May push” and “may push to feature/*” are different claims, and only the second is a control — a capability gate that authorizes the verb alone is the shape behind most published agent incidents.

A constraint may name:

  • paths — glob patterns (repos/**, !**/.env)
  • refs — git refs, with ! negation (['**', '!main'] reads as “any branch except main”)
  • hosts — network destinations, with subdomain wildcards (*.internal)
  • commands — command prefixes, matched on a word boundary
  • ceilings — maxCount / maxCents, so “delete up to 50 files” is expressible

Three properties fall out of keeping the grammar closed and loop-free, and all three are the point: every rule renders as a sentence you can read back, every denial can name the constraint that matched, and the whole policy can be enumerated — you can see what an agent may do before it runs.

Three rules that matter:

  • A constrained rule is more specific than an unconstrained one at the same scope, so “deny push to main” beats “allow push” without anyone thinking about ordering.
  • A restrictive rule (deny / prompt) applies when the request touches any value it names; a permissive one (allow) applies only when the request is entirely inside it. Reading both the same way would let an agent launder a forbidden path by batching an innocent one beside it.
  • A rule that names a facet the request says nothing about is neither applied nor skipped: it escalates to an approval. Silently skipping would make every protected-branch rule inert against a tool that does not report its ref; silently denying would refuse every push the extractor cannot describe. With no approver connected, the existing fail-closed rule turns that escalation into a denial.

A commandPrefix rule matches on a word boundary: git push covers git push origin main and not git pushx. It does not normalise the shell. git push (two spaces), cd sub && git push and GIT_DIR=… git push are all different strings and none of them matches a rule about git push.

That is a real limit and it is stated rather than papered over: command rules are an operator convenience for shaping ordinary agent behaviour, not a boundary against an adversarial one. The boundaries are the effect classes (which a tool declares and cannot re-spell), the mode command net beneath the policy store, and the OS sandbox beneath both. A command rule that resolves to allow never loosens either of those — the mode net runs afterwards and its deny is final.

Answering a prompt with “remember this” writes a real policy rule: scoped (space / agent / workspace), argument-constrained (approving a push to feature/login grants feature/**, not main) and self-revoking on an explicit TTL. Saying yes for the next few hours is not the same as quietly rewriting your permanent policy, and the rules it writes land in the same store with provenance remembered and an expiry, so they can be revoked early.

Each rule carries one:

  • advisory — allow, but record that the rule matched. This is the adoption path: roll a strict rule out in advisory mode, read the audit trail for a week to see exactly what it would have blocked, then promote it. Without this tier, strict policies never get turned on.
  • soft — deny, overridable by someone holding the override permission, with a recorded justification. The override is itself an audited event.
  • hard (the default) — deny. Nothing overrides it.

An install operator can pin policy that no workspace admin can loosen. Managed rules live server-side and are merged most-restrictive with each workspace’s own chain — they are not a scope at the head of it, because a head-of-chain managed allow would let the install override a workspace’s deny, the opposite of what a clamp is for. A managed rule can therefore only ever tighten.

Setting CC_SERVER_MANAGED_POLICY to a JSON file outranks the stored rules entirely, which is what lets an operator pin a posture that no admin UI can flip — the answer to “can I stop my developers from disabling the safety controls?”.

The autonomy dial is a profile, not a parallel system

Section titled “The autonomy dial is a profile, not a parallel system”

The per-space autonomy dial is a named profile over this same policy store, not a second mechanism. Its three settings are stored and sent as proposeOnly, actWithApproval and actFreely (the hyphenated forms are UI prose only and autonomy.setForSpace rejects anything else). Leaving it unset behaves as actWithApproval.

What each one does to a resolved decision:

  • proposeOnly denies every gated tool outright. The agent can still reason and reply; it just cannot act and its message says why.
  • actWithApproval (and unset) is the fail-closed approval gate described above.
  • actFreely pre-approves anything that did not resolve to a hard deny.

That last one deserves to be stated plainly rather than softened: under actFreely a prompt decision is not escalated to an approver — it is allowed and no one is asked. Only an explicit deny rule survives the dial. So actFreely is a deliberate grant of autonomy, not a convenience setting, and the fail-closed rule protects exactly the case where it matters (actWithApproval, the default).

Delegation between agents is guarded at the delegate_task chokepoint by all four guards: a depth cap (default 3), cycle detection, an autonomy ceiling (a delegate can never act with more autonomy than its delegator — privilege cannot be laundered by handing work to a freer agent) and a budget envelope (delegation bills the delegator’s remaining budget and cannot mint more). Each is refused loudly, with the guard’s reason returned to the agent verbatim.

Not every runtime lets Control Center intercept every action. External agent CLIs run some tools natively, in-process, where no gate can sit. Rather than pretend otherwise, each adapter carries an honesty matrix: per adapter, whether Control Center filters the tool surface, intercepts tool calls, observes the completion contract, can see the runner’s native tools and whether in-process tools are covered by a sandbox profile.

Two entries are worth reading before you trust a configuration. The built-in harness declares that its in-process file tools are not sandboxed — the tool surface and this policy engine are the only filesystem boundary it has. Claude Code declares that its own tool calls are not interceptable, because it is launched with --dangerously-skip-permissions; Control Center sees only its mcp__* calls and its own read, write, edit and shell tools run unseen.

The settings screen shows this matrix, because the one unforgivable failure for a guardrail system is claiming coverage that does not exist.

The guardrails live at Settings → Workspace → Agent permissions (/workspaces/<id>/settings/workspace/permissions). The screen has the policy matrix where rules are written, policy templates, a what-if probe, the authorization audit trail and the adapter honesty matrix.

The probe takes an action class (or a command) and a scope and shows which rule decided and why. Because resolution is pure and deterministic — same action, same scope, same decision — the probe’s answer is the answer the agent will get.