p(doom): a feedback loop for agent control
Why build a control architecture for AI agents at all? One common answer is p(doom). This page explains the term, shows how published estimates have changed, records how that feedback shaped C=US, and is plain about which risks the architecture does and does not address. Then it asks for your own estimate.
What p(doom) is
p(doom) is the probability a person assigns to artificial intelligence causing human extinction, or a similarly severe and permanent catastrophe. It is a personal judgment, not a measurement. People define “doom” differently, use different time horizons (some say “in 30 years”, some give no horizon) and revise their numbers as AI changes. Comparing two numbers therefore says more about how worried two people are than about the world.
It is still useful here. A high estimate argues for strong control. A low estimate asks a fair question: does the architecture pay for itself even if catastrophe is unlikely?
Published estimates, and how some have changed
Values as compiled in Wikipedia's P(doom) article, which links each person's original statement. Ranges are drawn as bars.
Scale: 0% to 100%. Cyan bars are published estimates; your estimate appears in green after you take the self-assessment below.
| Person | Role | Estimate | When | Change over time |
|---|
How this feedback changes the architecture
Each estimate comes with a reason. The reasons, more than the numbers, are what an architecture can respond to.
Geoffrey Hinton: profit motive alone will not keep us safe
Hinton raised his estimate from 10% to “10% to 20%” over the next three decades, said development was faster than he expected, and called for government regulation, warning that “the invisible hand is not going to keep us safe” (The Guardian, December 2024).
Architectural response: authority over agents comes from accountable, state-endorsed identity rather than from vendors alone. A government stop sits at the top (the proposed AI Kill Switch Act, pending), and the duty of loyalty is enforced at every use of a person's data.
Yoshua Bengio: risk comes in three kinds
Bengio chaired the International AI Safety Report, which groups risks into malicious use, malfunctions (including loss of control) and systemic risks such as market concentration, single points of failure and environmental cost.
Architectural response: the coverage table below is organised by those categories, so gaps are visible rather than implied.
Yann LeCun: catastrophe is very unlikely
At under 0.01%, LeCun's estimate is the low anchor.
Architectural response: the design must earn its keep on everyday failures, not only on doom. Agent identity, consent and audit pay off against fraud, impersonation and an agent serving the wrong interest, whatever one's p(doom).
Dan Hendrycks and Eliezer Yudkowsky: very high estimates
Hendrycks's estimate rose from about 20% to over 80%; Yudkowsky's is above 95%. Both concern systems far more capable than today's agents.
Architectural response: none that would be honest to claim. A registry can limit what registered agents are authorised to do and record who is accountable. It cannot control a system that ignores its authorisation, and it does not address that scenario.
A reviewer: the danger is systemic, such as the power grid
In a conversation recorded on sugarbushagent.com (Reviewer 3), a reviewer thought AI itself unlikely to kill us, but that systemic effects such as data-center demand on the power grid might.
Architectural response: favour small, task-specific models on local, renewable power for well-defined business tasks, and record each agent's shared dependencies in the systemic dependency graph.
What the architecture does and does not address
Agentic risks use the OWASP Top 10 for Agentic Applications (2026), an expert taxonomy. Misuse, malfunction and systemic risks follow the International AI Safety Report, plus the energy and water concerns raised by reviewers who are not AI specialists.
| Risk | Coverage | C=US control, or why not |
|---|
“Cybersecurity” marks risks assigned to an organisation's security team rather than to the registry; the registry supplies the identity, revocation and audit evidence they work from.
Changes made, and planned, in response
Reading the estimates for their reasons led to two changes in the dependency graph, each tested on synthetic data:
- Grant limits (done). A grant can cap operations, spend and data volume. Usage counts against every grant in the delegation chain, so delegating cannot multiply a budget. This limits how far a hijacked or faulty agent can go before it is stopped.
- Circuit breakers (done). Suspending a shared dependency stops every agent that would fail without it, while agents with a working alternative carry on.
- Sandbox versions as shared dependencies (done). If agents run in NVIDIA OpenShell sandboxes, each OpenShell version is recorded as a runtime every agent depends on. A vulnerable version can be suspended, stopping exactly the agents still on it; agents moved to the fixed version carry on.
- Identity path through the sandbox (tested). In a local OpenShell sandbox, an agent authenticated to a mutual-TLS service through an identity proxy outside the sandbox that holds its key; OpenShell allowed only the one permitted request and refused everything else. The same test found the sandbox could read a private key placed inside it, so keys stay outside the sandbox.
Planned, not yet built:
- Approval from the originating person for actions above a threshold, so injected instructions cannot produce it (goal hijack).
- Model and runtime recorded at enrollment, covered by the device’s hardware attestation, with unknown versions refused (supply chain).
- “Who is behind this agent?” A lookup of the accountable person and organisation for anyone an agent deals with (human–agent trust).
- Energy and water attributes, compute budgets and concentration alerts in the dependency graph (grid and water).
- A defined handover to cybersecurity: a revocation feed and the signed audit trail, exported to the security team’s monitoring tools.
Your p(doom)
Take the self-assessment. Your answers are processed in this page only; nothing is stored or sent unless you press “Send as feedback”, which opens a Google Form with your answers filled in. You can change them, add a comment, and choose whether to submit.
Your result
Sent feedback may be quoted anonymously on this page as part of the feedback loop.
The feedback loop
- Listen. Published estimates, expert taxonomies, reviewer conversations and your self-assessment.
- Find the reason behind the number. Regulation, misuse, loss of control, energy, water, concentration.
- Map it to a control, or record that there is none. The coverage table is the record.
- Change the architecture where it can help, recording the change the same way corrections are recorded in the break/fix cycle.
- Ask again. Estimates change; the table is revised when they do.
The architecture does not lower anyone's p(doom) by itself. It makes authority, accountability and dependencies visible for agents that are registered, which addresses many everyday and systemic risks, and it says so plainly where it does not.
Sources
- P(doom), Wikipedia: compiled estimates with links to each original statement
- The Guardian (2024): Hinton shortens odds of AI wiping out humanity over the next 30 years
- Bengio et al. (2025): International AI Safety Report
- OWASP GenAI Security Project: Top 10 for Agentic Applications (2026)
- IEA (2025): Energy and AI
- Li et al. (2023): Making AI Less “Thirsty”: the water footprint of AI models
- Belcak et al. (2025): Small Language Models are the Future of Agentic AI, NVIDIA Research
Provenance: drafted in Claude Code (Claude Opus 5.5) at the project lead's request, from the sources above and this project's own pages. The coverage judgments are the project's own assessment, not those of the people quoted.