Guardrails — one policy for every mutating action
Nội dung này hiện chưa có sẵn bằng ngôn ngữ của bạn.
Agents in Control Center can delete files, push branches, publish pull requests and call external APIs. Each of those is a different kind of risk, and for a long time only one of them — shell commands — had a policy. The guardrails exist to answer one question uniformly, for every mutating action an agent can take, regardless of which runtime or tool proposes it: is this allowed, does it need an approval, or is it denied?
One engine, every chokepoint
Section titled “One engine, every chokepoint”There is exactly one policy engine. It is consulted at every point where an action can leave the system: the built-in harness’s tool registry, the server-side MCP dispatcher gate and repo-op mutations. A rule written once therefore applies to the built-in runtime and to external agent CLIs alike — there is no second, quieter path around it. The older bash-only command policy still works: a rule can match a command prefix instead of an action class and the two forms coexist in the same store.
The engine is not, however, a substitute for the OS sandbox and the sandbox is
not a substitute for it. They answer different questions: the guardrails decide
whether a call Control Center can see is allowed to happen at all, while the
sandbox bounds where the resulting process may reach. cc_server applies the
sandbox where the host offers a backend — but Windows offers none and Linux
needs bwrap and socat. On those hosts
the guardrails are the outermost enforcement layer rather than a policy sitting
on top of a floor, which is worth knowing when you decide how permissive to be.
ActionClass: a deliberately closed taxonomy
Section titled “ActionClass: a deliberately closed taxonomy”Every mutating tool declares which ActionClasses it belongs to, from a closed set of thirteen effect classes. Five of them prompt out of the box; the rest allow and a rule is what changes that.
| Class | Built-in default | What it covers |
|---|---|---|
fileDelete |
prompt | Deleting a file (a delete-shaped write, rm) |
gitPush |
prompt | Pushing to a git remote |
prCreate |
prompt | Creating a pull request |
prPublish |
prompt | Publishing a review, merging a PR |
vendorSyncWrite |
prompt | Writing to an external ticket vendor (Linear, Jira, GitHub sync) |
fileWriteOutsideWorktree |
allow | Writing a file outside the isolated worktree |
gitCommit |
allow | Creating a git commit |
networkEgress |
allow | Network egress (web fetch, arbitrary HTTP) |
secretAccess |
allow | Reading a secret or credential |
packageInstall |
allow | Installing a package (npm, pip, pub, brew) |
processSpawn |
allow | Spawning a process (shell and anything shelling out) |
workspaceMutation |
allow | Mutating workspace structure (repos, spaces, agents) |
enclosureControl |
allow | Driving an enclosure (rig): booting a VM, sending it input |
The eight allow-by-default classes were chosen on the assumption that the
sandbox floor applies beneath them. It does where the host offers a backend,
and it does not on Windows or on a Linux box missing bwrap and socat — so
if you want fileWriteOutsideWorktree or packageInstall gated there, that is
a rule you have to write.
The set is closed on purpose. A new mutating tool that declares no class fails the ratchet test, so the taxonomy can only grow by an explicit, argued decision — “taxonomy sprawl is the death of this feature.” Thirteen classes is enough to express the policies people actually write (never push without asking, never delete outside a worktree, always allow reads) and few enough that the settings UI stays comprehensible.
Rules and scope resolution
Section titled “Rules and scope resolution”A rule maps (scope, ActionClass | command prefix) → allow | prompt | deny.
Scopes nest from most to least specific:
space > agent > workspace > mode preset > built-in default
Resolution is specificity first: the first scope with a matching rule decides and resolution stops there. Most-restrictive is only a tie-break — between two equally specific rules within one scope and when combining the several classes a single action declares. Within a scope, a command rule matches on the longest command prefix. Allow does not beat deny anywhere.
Two decisions carry hard guarantees:
- Fail-closed prompts. A
promptwith no approver connected resolves to denied. An unanswered question is never a yes. - Most-restrictive combination. An action spanning several classes combines them most-restrictively; when several classes prompt, one confirmation lists them all.
There is deliberately no per-turn latch. Every tool call re-resolves the policy from scratch, so repeated prompting for the same action is bounded by the operator’s own “remember this decision” rather than by anything turn-scoped.
Argument-level rules
Section titled “Argument-level rules”A rule can constrain the ARGUMENTS of an action, not just its verb. “May
push” and “may push to feature/*” are different claims, and only the second
is a control — a capability gate that authorizes the verb alone is the shape
behind most published agent incidents.
A constraint may name:
- paths — glob patterns (
repos/**,!**/.env) - refs — git refs, with
!negation (['**', '!main']reads as “any branch except main”) - hosts — network destinations, with subdomain wildcards (
*.internal) - commands — command prefixes, matched on a word boundary
- ceilings —
maxCount/maxCents, so “delete up to 50 files” is expressible
Three properties fall out of keeping the grammar closed and loop-free, and all three are the point: every rule renders as a sentence you can read back, every denial can name the constraint that matched, and the whole policy can be enumerated — you can see what an agent may do before it runs.
Three rules that matter:
- A constrained rule is more specific than an unconstrained one at the
same scope, so “deny push to
main” beats “allow push” without anyone thinking about ordering. - A restrictive rule (deny / prompt) applies when the request touches any value it names; a permissive one (allow) applies only when the request is entirely inside it. Reading both the same way would let an agent launder a forbidden path by batching an innocent one beside it.
- A rule that names a facet the request says nothing about is neither applied nor skipped: it escalates to an approval. Silently skipping would make every protected-branch rule inert against a tool that does not report its ref; silently denying would refuse every push the extractor cannot describe. With no approver connected, the existing fail-closed rule turns that escalation into a denial.
What a command rule can and cannot catch
Section titled “What a command rule can and cannot catch”A commandPrefix rule matches on a word boundary: git push covers
git push origin main and not git pushx. It does not normalise the
shell. git push (two spaces), cd sub && git push and
GIT_DIR=… git push are all different strings and none of them matches a
rule about git push.
That is a real limit and it is stated rather than papered over: command rules
are an operator convenience for shaping ordinary agent behaviour, not a
boundary against an adversarial one. The boundaries are the effect classes
(which a tool declares and cannot re-spell), the mode command net beneath the
policy store, and the OS sandbox beneath both. A command rule that resolves to
allow never loosens either of those — the mode net runs afterwards and its
deny is final.
Standing approvals
Section titled “Standing approvals”Answering a prompt with “remember this” writes a real policy rule: scoped
(space / agent / workspace), argument-constrained (approving a push to
feature/login grants feature/**, not main) and self-revoking on an
explicit TTL. Saying yes for the next few hours is not the same as quietly
rewriting your permanent policy, and the rules it writes land in the same
store with provenance remembered and an expiry, so they can be revoked early.
Enforcement levels
Section titled “Enforcement levels”Each rule carries one:
- advisory — allow, but record that the rule matched. This is the adoption path: roll a strict rule out in advisory mode, read the audit trail for a week to see exactly what it would have blocked, then promote it. Without this tier, strict policies never get turned on.
- soft — deny, overridable by someone holding the override permission, with a recorded justification. The override is itself an audited event.
- hard (the default) — deny. Nothing overrides it.
The managed tier
Section titled “The managed tier”An install operator can pin policy that no workspace admin can loosen. Managed
rules live server-side and are merged most-restrictive with each
workspace’s own chain — they are not a scope at the head of it, because a
head-of-chain managed allow would let the install override a workspace’s
deny, the opposite of what a clamp is for. A managed rule can therefore only
ever tighten.
Setting CC_SERVER_MANAGED_POLICY to a JSON file outranks the stored rules
entirely, which is what lets an operator pin a posture that no admin UI can
flip — the answer to “can I stop my developers from disabling the safety
controls?”.
The autonomy dial is a profile, not a parallel system
Section titled “The autonomy dial is a profile, not a parallel system”The per-space autonomy dial is a named profile over this same policy store,
not a second mechanism. Its three settings are stored and sent as
proposeOnly, actWithApproval and actFreely (the hyphenated forms are UI
prose only and autonomy.setForSpace rejects anything else). Leaving it
unset behaves as actWithApproval.
What each one does to a resolved decision:
proposeOnlydenies every gated tool outright. The agent can still reason and reply; it just cannot act and its message says why.actWithApproval(and unset) is the fail-closed approval gate described above.actFreelypre-approves anything that did not resolve to a harddeny.
That last one deserves to be stated plainly rather than softened: under
actFreely a prompt decision is not escalated to an approver — it is
allowed and no one is asked. Only an explicit deny rule survives the dial.
So actFreely is a deliberate grant of autonomy, not a convenience setting,
and the fail-closed rule protects exactly the case where it matters
(actWithApproval, the default).
Delegation between agents is guarded at the delegate_task chokepoint by all
four guards: a depth cap (default 3), cycle detection, an autonomy ceiling
(a delegate can never act with more autonomy than its delegator — privilege
cannot be laundered by handing work to a freer agent) and a budget
envelope (delegation bills the delegator’s remaining budget and cannot mint
more). Each is refused loudly, with the guard’s reason returned to the agent
verbatim.
The adapter honesty matrix
Section titled “The adapter honesty matrix”Not every runtime lets Control Center intercept every action. External agent CLIs run some tools natively, in-process, where no gate can sit. Rather than pretend otherwise, each adapter carries an honesty matrix: per adapter, whether Control Center filters the tool surface, intercepts tool calls, observes the completion contract, can see the runner’s native tools and whether in-process tools are covered by a sandbox profile.
Two entries are worth reading before you trust a configuration. The built-in
harness declares that its in-process file tools are not sandboxed — the
tool surface and this policy engine are the only filesystem boundary it has.
Claude Code declares that its own tool calls are not interceptable, because
it is launched with --dangerously-skip-permissions; Control Center sees only
its mcp__* calls and its own read, write, edit and shell tools run unseen.
The settings screen shows this matrix, because the one unforgivable failure for a guardrail system is claiming coverage that does not exist.
Probing policy before it bites
Section titled “Probing policy before it bites”The guardrails live at Settings → Workspace → Agent permissions
(/workspaces/<id>/settings/workspace/permissions). The screen has the
policy matrix where rules are written, policy templates, a what-if probe, the
authorization audit trail and the adapter honesty matrix.
The probe takes an action class (or a command) and a scope and shows which rule decided and why. Because resolution is pure and deterministic — same action, same scope, same decision — the probe’s answer is the answer the agent will get.
Related concepts
Section titled “Related concepts”- Sandbox and security: what actually constrains an agent run and how far the OS floor beneath this engine reaches
- Multiplayer — identity, membership and presence: spaces, principals and where the autonomy dial lives
- Configure guardrails: writing and testing rules in practice