C=USStop my agent →

The kill switch: a control layer outside the sandbox

A sandbox decides what an agent's program can touch on one machine. The C=US kill switch decides whether the agent may act at all, anywhere C=US is asked, and it puts that decision in a person's hands. It does not replace the sandbox; it is built on top of it. This page explains how it works, and why it rests on an old idea from cryptography: keep the key secret, publish everything else, and put control where it can be proved.

How it works

The C=US control layer outside the sandbox On the left, a sandbox holds the agent's program, its model and its tools, with no private key inside. Every request leaves the sandbox through an identity step that holds the agent's key, and reaches the C=US control layer over mutual TLS. The control layer has the agent gateway, the directory, the stops in force and the signed evidence log. A person with a passkey can stop or restart the agent, and another system can send a signed kill signal. The gateway's answer reaches relying parties such as the trust registry and the sugarbushagent.com services. Sandbox e.g. NVIDIA OpenShell: what the code can touch on this machine Agent programits reasoning and logic Modellocal or remote Tools and fileslimited by the sandbox No private key inside files, network and system calls contained Identity step holds the agent's key (hardware where possible) outside the sandbox every request C=US control layer cryptographic, outside the sandbox Stops in forceby the person, or by a kill signal Agent gatewaydecides every request, fails closed Directory c=US (X.500 / LDAP)agent, grants, accountable person, passkeys Signed evidence logevery decision, stop and restart mTLS Another systemsigned kill signal (CAEP) Accountable person passkey (e.g. YubiKey): stop, restart Relying partiestrust registry, edges, credential checks "not authorized"
The sandbox contains the program. Everything that decides whether the agent may act, and the people and systems who can stop it, sits outside the sandbox, where the agent's code cannot reach.
  1. The agent runs in a sandbox. The sandbox limits the files, network and system calls its program can use. Its private key is not in there: the step that proves the agent's identity runs outside, as tested with NVIDIA OpenShell (identity versus runtime).
  2. Every request is decided outside the sandbox. The agent gateway checks the certificate, the directory entry, the grant and any stop in force, then signs its decision into the log.
  3. A person can pull the switch. The person accountable for the agent stops it with their own passkey on the stop-my-agent page, and only they can restart it. Drilled on 2026-10-10: the person enrolled a passkey (Windows Hello), an administrator added it to the directory, and the person stopped and restarted the research intern, each step confirmed by the passkey and signed into the log.
  4. Other systems can pull it too, but only one way. A guardrail or security monitor that C=US trusts can send a signed kill signal. A signal can take authority away; it can never grant it.
  5. The stop reaches everyone who asks. The gateway, the trust registry and the services that rely on it refuse the agent on its next request. In the 2026-10-10 drill this took 0.35 seconds.

An addition to sandboxing, built on it

Sandboxing is a fundamental component: without it, an agent's program could read its own key, or anything else on the machine. But a sandbox answers only one question, about one machine. The kill switch answers the others.

SandboxC=US kill switch
Question it answersWhat can this program touch on this machine?May this agent act at all, anywhere, right now?
Where it worksOn the machine running the agentWherever C=US is asked: the gateway, the registry, the services relying on it
If the program escapesThe containment is lostAn escape gives no authority: the key and the grants are outside, so requests are still refused
Who decidesWhoever configured the sandboxThe accountable person, an administrator, or a trusted system that can only stop
EvidenceLocal logs, if anyEvery decision, stop and restart signed in a hash-chained log anyone can check
What it cannot doStop the agent acting through another machineEnd the program, or reach services that never ask C=US

Used together, each covers the other's blind spot: the sandbox keeps the program away from its key, and the control layer keeps authority away from the program.

The story of public keys: keeping one secret

For most of history, two people who wanted to communicate secretly first had to share a secret key, delivered in advance by a courier or a meeting. The key had to travel, and anyone who copied it on the way could read everything.

The lesson for privacy is that only the private key has to stay secret. Public keys and certificates can be published, copied and looked up by anyone. A person or an agent can prove who they are without ever revealing the secret that proves it, and anyone can send them something only they can read, without arranging anything in advance. Less that must be hidden means less that can leak.

X.500 and LDAP: a directory that can be private

X.500 (ITU-T, first published in 1988) describes a world-wide directory: entries named by distinguished names such as c=US, held by many servers, each organisation and country holding its own part. LDAP, first defined in 1993 (RFC 1487), is the lightweight protocol that lets ordinary software read and search such a directory over the internet. X.500 supplies the model, the names and the certificates; LDAP is how applications reach it. Together they help privacy in C=US in four ways:

The only thing that must stay secret is the secret

In 1883 Auguste Kerckhoffs set out principles for military ciphers in La cryptographie militaire. The most lasting is that a system must stay secure even if the enemy knows everything about it except the key. Claude Shannon restated it in 1949 as a maxim: assume the enemy knows the system.

A frontier AI model is close to that enemy, or ally, in practice. Trained on a large share of the public written record (standards, textbooks, source code, research papers), it can explain how X.509, TLS and LDAP work, and infer much that was never written down. Any security that depends on an AI not knowing the method is already lost. So C=US follows Kerckhoffs:

What a neural network knows is, in this sense, like the method: wide, shared and impossible to take back. What it must not have is the key. Control rests on the one thing it cannot learn by reading.

Shannon and Turing: what can be kept secret, and what can be known

Alan Turing (1912–1954)Claude Shannon (1916–2001)
Founding workOn Computable Numbers (1936): what any machine can compute, and that some questions no machine can answerA Mathematical Theory of Communication (1948): information measured in bits
In cryptographyAt Bletchley Park in the Second World War, with others, designed the method and machines (the bombe) used to break German Enigma trafficCommunication Theory of Secrecy Systems (1949, from a classified 1945 report): proved what perfect secrecy needs (a one-time pad) and put the key at the centre
On thinking machinesComputing Machinery and Intelligence (1950): the imitation gameBuilt early learning machines, such as a maze-solving mouse; met Turing at Bell Labs in 1943, where they discussed machines that might think
What it means for agentsThere is no general way to decide, by inspecting a program, what it will do (made general by Rice's theorem, 1953). An agent's behaviour cannot be fully verified from the insideSecrecy can be made exact: it rests on the key, measured and protected, not on hoping the method stays unknown

Put together, they say where control belongs. Turing shows that no inspection can prove what an arbitrary agent will do, so control cannot depend on reading its mind or its code. Shannon shows that a key can be kept secret even when everything else is known. So the control point is a cryptographic check of keys and signed records, which can be proved, outside the program, which cannot be.

The point of all this

In C=US, cryptography creates a control layer outside the sandbox, and that control is delegated to a person. It is not delegated to agents in a collective pursuing ill-defined tasks of their own choosing.

Grants are dated, scoped and issued by people. An agent cannot register another agent or give it authority. The accountable person stops and restarts with their own passkey. Other systems may only take authority away. Every decision is signed, so anyone can check who decided what.

Sources

Drafted by Claude Code (Claude Opus 5.5) at the project lead's request, 2026-10-10, from the project's own pages and the sources above; the history is summarized from them and from general reference, and should be checked before it is cited. The link between Kerckhoffs's principle and what frontier models know is the project lead's idea.