The kill switch: a control layer outside the sandbox
A sandbox decides what an agent's program can touch on one machine. The C=US kill switch decides whether the agent may act at all, anywhere C=US is asked, and it puts that decision in a person's hands. It does not replace the sandbox; it is built on top of it. This page explains how it works, and why it rests on an old idea from cryptography: keep the key secret, publish everything else, and put control where it can be proved.
How it works
- The agent runs in a sandbox. The sandbox limits the files, network and system calls its program can use. Its private key is not in there: the step that proves the agent's identity runs outside, as tested with NVIDIA OpenShell (identity versus runtime).
- Every request is decided outside the sandbox. The agent gateway checks the certificate, the directory entry, the grant and any stop in force, then signs its decision into the log.
- A person can pull the switch. The person accountable for the agent stops it with their own passkey on the stop-my-agent page, and only they can restart it. Drilled on 2026-10-10: the person enrolled a passkey (Windows Hello), an administrator added it to the directory, and the person stopped and restarted the research intern, each step confirmed by the passkey and signed into the log.
- Other systems can pull it too, but only one way. A guardrail or security monitor that C=US trusts can send a signed kill signal. A signal can take authority away; it can never grant it.
- The stop reaches everyone who asks. The gateway, the trust registry and the services that rely on it refuse the agent on its next request. In the 2026-10-10 drill this took 0.35 seconds.
An addition to sandboxing, built on it
Sandboxing is a fundamental component: without it, an agent's program could read its own key, or anything else on the machine. But a sandbox answers only one question, about one machine. The kill switch answers the others.
| Sandbox | C=US kill switch | |
|---|---|---|
| Question it answers | What can this program touch on this machine? | May this agent act at all, anywhere, right now? |
| Where it works | On the machine running the agent | Wherever C=US is asked: the gateway, the registry, the services relying on it |
| If the program escapes | The containment is lost | An escape gives no authority: the key and the grants are outside, so requests are still refused |
| Who decides | Whoever configured the sandbox | The accountable person, an administrator, or a trusted system that can only stop |
| Evidence | Local logs, if any | Every decision, stop and restart signed in a hash-chained log anyone can check |
| What it cannot do | Stop the agent acting through another machine | End the program, or reach services that never ask C=US |
Used together, each covers the other's blind spot: the sandbox keeps the program away from its key, and the control layer keeps authority away from the program.
The story of public keys: keeping one secret
For most of history, two people who wanted to communicate secretly first had to share a secret key, delivered in advance by a courier or a meeting. The key had to travel, and anyone who copied it on the way could read everything.
- 1976: Whitfield Diffie and Martin Hellman published New Directions in Cryptography. Each person could have two keys: a private one they never share, and a public one anyone may see. Two strangers could agree on a secret over an open line.
- 1977: Ron Rivest, Adi Shamir and Leonard Adleman described RSA, which made public-key encryption and digital signatures practical.
- 1970–1974, revealed in 1997: James Ellis, Clifford Cocks and Malcolm Williamson had found the same ideas at the UK's GCHQ, kept secret until the British government declassified them.
- 1988: X.509, published as part of the X.500 directory standards, defined the certificate: a public key bound to a name and signed by a certificate authority. This is public key infrastructure (PKI), and it is how every website, and every C=US agent, proves who it is.
The lesson for privacy is that only the private key has to stay secret. Public keys and certificates can be published, copied and looked up by anyone. A person or an agent can prove who they are without ever revealing the secret that proves it, and anyone can send them something only they can read, without arranging anything in advance. Less that must be hidden means less that can leak.
X.500 and LDAP: a directory that can be private
X.500 (ITU-T, first published in 1988) describes a world-wide directory: entries named by distinguished names such as c=US, held by many servers, each organisation and country holding its own part. LDAP, first defined in 1993 (RFC 1487), is the lightweight protocol that lets ordinary software read and search such a directory over the internet. X.500 supplies the model, the names and the certificates; LDAP is how applications reach it. Together they help privacy in C=US in four ways:
- Access control on every attribute. Nobody can read the C=US directory anonymously. Signed-in services may read it. The server's root login, connecting locally, may change only certificate fields; anything else needs the directory administrator's password.
- Publish proofs, not identities. The directory holds public keys, certificate fingerprints and, for a person's identity proofing, a digest rather than the identifier itself.
- Each jurisdiction keeps its own data. Quebec's entries live on the c=CA server; c=US reaches them through a read-only chain and stores no copy.
- Identity is separate from activity. Who an agent is, and who answers for it, is in the directory. What it did (grades, readings, stops) is in a separate database with its own access rules.
The only thing that must stay secret is the secret
In 1883 Auguste Kerckhoffs set out principles for military ciphers in La cryptographie militaire. The most lasting is that a system must stay secure even if the enemy knows everything about it except the key. Claude Shannon restated it in 1949 as a maxim: assume the enemy knows the system.
A frontier AI model is close to that enemy, or ally, in practice. Trained on a large share of the public written record (standards, textbooks, source code, research papers), it can explain how X.509, TLS and LDAP work, and infer much that was never written down. Any security that depends on an AI not knowing the method is already lost. So C=US follows Kerckhoffs:
- The method is described openly: the schema, the gateway's rules and how each check works are explained on this website.
- The secrets are keys, kept in hardware where possible: a YubiKey or a TPM signs, but will not hand its key to anyone, including an AI that knows exactly how the system works.
What a neural network knows is, in this sense, like the method: wide, shared and impossible to take back. What it must not have is the key. Control rests on the one thing it cannot learn by reading.
Shannon and Turing: what can be kept secret, and what can be known
| Alan Turing (1912–1954) | Claude Shannon (1916–2001) | |
|---|---|---|
| Founding work | On Computable Numbers (1936): what any machine can compute, and that some questions no machine can answer | A Mathematical Theory of Communication (1948): information measured in bits |
| In cryptography | At Bletchley Park in the Second World War, with others, designed the method and machines (the bombe) used to break German Enigma traffic | Communication Theory of Secrecy Systems (1949, from a classified 1945 report): proved what perfect secrecy needs (a one-time pad) and put the key at the centre |
| On thinking machines | Computing Machinery and Intelligence (1950): the imitation game | Built early learning machines, such as a maze-solving mouse; met Turing at Bell Labs in 1943, where they discussed machines that might think |
| What it means for agents | There is no general way to decide, by inspecting a program, what it will do (made general by Rice's theorem, 1953). An agent's behaviour cannot be fully verified from the inside | Secrecy can be made exact: it rests on the key, measured and protected, not on hoping the method stays unknown |
Put together, they say where control belongs. Turing shows that no inspection can prove what an arbitrary agent will do, so control cannot depend on reading its mind or its code. Shannon shows that a key can be kept secret even when everything else is known. So the control point is a cryptographic check of keys and signed records, which can be proved, outside the program, which cannot be.
In C=US, cryptography creates a control layer outside the sandbox, and that control is delegated to a person. It is not delegated to agents in a collective pursuing ill-defined tasks of their own choosing.
Grants are dated, scoped and issued by people. An agent cannot register another agent or give it authority. The accountable person stops and restarts with their own passkey. Other systems may only take authority away. Every decision is signed, so anyone can check who decided what.
Sources
- Diffie, W. and Hellman, M. (1976). New Directions in Cryptography. IEEE Transactions on Information Theory 22(6).
- Kerckhoffs, A. (1883). La cryptographie militaire. Journal des sciences militaires (with an English summary of the six principles).
- Shannon, C. E. (1948). A Mathematical Theory of Communication; (1949) Communication Theory of Secrecy Systems. Bell System Technical Journal.
- Turing, A. M. (1936). On Computable Numbers, with an Application to the Entscheidungsproblem; (1950) Computing Machinery and Intelligence.
- ITU-T X.500 and X.509; IETF RFC 1487 (LDAP, 1993) and RFC 4510 (LDAP today).
- This project: the kill switch on the Zero Trust page, identity versus runtime, stop my agent.
Drafted by Claude Code (Claude Opus 5.5) at the project lead's request, 2026-10-10, from the project's own pages and the sources above; the history is summarized from them and from general reference, and should be checked before it is cited. The link between Kerckhoffs's principle and what frontier models know is the project lead's idea.