C=US Liquid Accord Scorecard← Control-plane architectures vs. compliance standards

Accord scorecard · agentic risk & compliance follow-up

Scoring the White House AI accord against the risks this project already tracks.

On 2026-09-29 the President and a group of AI company leaders announced a voluntary accord on AI safety. This page asks a narrow question: how well does that accord alone — with no other law, standard, or technical control — address the ten agentic-AI risks and four compliance standards already documented elsewhere in this project, and does it add anything once real technical controls (OpenShell sandboxing plus c=US identity/authorization) are already in place?

This is not an attack on the accord's sincerity or an assertion about any signing company's intent. It only tests whether the accord's own stated text, taken at face value, constitutes a technical or independently-verifiable control for these specific risks and standards — not whether AI safety work is happening at these companies through other means not described in the accord itself.

Method

Two different jobs on this page, kept visibly separate

Per this project's own convention (see AGENTS.md), every AI-generated hypothesis is attributed to the tool that produced it. Every number on this page came from Jev; every fact, source, and framing choice came from Claude Code.

WHAT CLAUDE CODE DID
  • Fetched the actual White House executive order and, separately, news coverage of the accord (Forbes, Axios, CNN, Washington Times) to write a factual, non-editorial description of what the accord commits signing companies to — and, just as importantly, what it does not contain (no enforcement, no shared technical standard, no penalty, no disclosure requirement).
  • Pulled the ten risk descriptions straight from two pages this project had already published — site/agentic-risks.html and research/web-connected-agent-security.md — rather than inventing new risk language for this analysis.
  • Reused the same four compliance standards and the same OpenShell/c=US architecture description already scored in research/jev-compliance-mapping-analysis.json and research/jev-followup-gaps-and-control-analysis.json, so this run's numbers are directly comparable to those earlier ones.
  • Wrote every question's instructions and every scoring rubric's levels (what "not addressed" vs. "substantially addressed" means, etc.) — Jev only ever sees the rubric Claude Code writes.
  • Ran the calls, saved the raw request/response JSON to the repository, and wrote the interpretation and page copy below.
WHAT JEV DID
  • Given that fixed state and those fixed questions, returned every numeric judgment shown on this page: the 0–3 risk-coverage scores and their full probability distributions, the 0–1 compliance-relevance scores for the four standards, and the comparison score for whether adding the accord to an existing technical combination helps.
  • Jev did not choose the risks, write the rubric levels, fetch the accord's text, or decide which standards to test against — it only judged the specific, pre-written question against the specific, pre-written state each time.
  • Model version (jev-1.13.0), full state/question text, and token usage for every call are preserved in the linked JSON files in Sources, exactly as returned by the API — nothing here is a paraphrase of Jev's output.

Result 1 of 3

Ten risks, scored 0 (not addressed) to 3 (substantially addressed)

Jev scored how well the accord alone would address each risk. No risk received any probability of landing on level 3, and none scored above 1.60 — Claude Code wrote the risk descriptions and the four rubric levels; Jev returned every score and probability below.

Insecure, unvalidated model output
1.60
Goal drift beyond task boundary
1.59
Unapproved egress / SSRF
1.59
Indirect prompt injection
1.58
Privilege chaining / excessive agency
1.56
Human over-reliance, no approval gate
1.48
Detection lag / cross-agent blindness
1.33
Emergent multi-agent coordination
1.07
Unaccountable agent identity
0.87
Third-party attribution blindness
0.75

Scale: 0 not addressed · 1 nominally addressed (touches the theme, no verifiable mechanism) · 2 partially addressed (relevant but voluntary and unenforceable) · 3 substantially addressed (specific, independently verifiable mechanism). The risks c=US's own architecture targets most directly — identity, cross-agent coordination, third-party attribution — score lowest, because the accord has no identity or cross-company concept at all.

Result 2 of 3

The accord alone, against four real compliance standards

Same four standards, same question wording, and the same c=US comparison numbers used in the earlier control-plane compliance analysis — Claude Code defined the standards and reused that exact wording; Jev returned every cell.

NIST SP 800-63
(identity)
NIST AI RMF /
AI 600-1
OWASP Agentic
Threat Modeling
ISO 27001 /
CIS Controls
White House accord, alone0.120.540.240.38
c=US (X.500/LDAP + mTLS), alone0.760.510.720.65

Each cell: probability that the row, on its own, would meaningfully help satisfy or provide relevant technical evidence toward that standard. The accord ties c=US only on NIST AI RMF — itself a voluntary, process-oriented governance framework rather than a technical control standard — and trails badly everywhere a technical control is actually expected.

Result 3 of 3

Does the accord add anything once real controls exist?

The strongest architecture this project has tested is OpenShell (kernel-level agent sandboxing) combined with c=US's own identity/authorization layer. Claude Code asked Jev to score that combination with the accord layered on top, and to judge directly whether the accord makes it meaningfully stronger.

StandardOpenShell + c=US+ White House accord
NIST SP 800-630.820.84
NIST AI RMF0.770.80
OWASP Agentic0.830.78
ISO 27001 / CIS0.840.82

Flat within the noise of a single run — up slightly on two standards, down slightly on the other two. Asked directly whether adding the accord to the existing technical combination is a meaningful compliance gain, Jev returned 0.84 out of 2 (confidence 0.66): 78% probability on “small, mostly cosmetic benefit — not a material compliance gain,” 19% on “no meaningful difference,” and only 3% on “meaningful, additive benefit.”

What this checked, and what it found

The accord is a process commitment, not a technical control

Its four steps — internal controls, an internal team, a company-chosen external assessor, a board committee — are all single-company, self-selected, and unenforceable. That shape is why it caps out around 1.6/3 on every risk tested and never reaches "substantially addressed."

It scores best where the bar is already low

Its highest compliance score (0.54, NIST AI RMF) is on the one standard in this set that is itself a voluntary governance framework rather than a technical standard. Against NIST SP 800-63 (identity) it scores 0.12 — the lowest score any candidate has received on any standard in this project's Jev comparisons to date.

It adds essentially nothing once real controls exist

Layered on top of the already-strongest tested architecture (OpenShell sandboxing plus c=US identity/authorization), the accord moves every standard by two points or less and Jev puts only 3% probability on it being a meaningful additive benefit.

Sources and underlying data

What this page draws on

  • The accord's own commitments, as summarized by Claude Code from White House and press coverage: Forbes ("White House Releases 'Accord' Between Billionaire AI Execs"), Axios, CNN, and the Washington Times, all dated 2026-09-29.
  • Risk descriptions: site/agentic-risks.html and research/web-connected-agent-security.md, both already published elsewhere in this project.
  • Raw Jev request/response data, including full question text, rubric levels, model version, and token usage, for every number on this page: research/jev-accord-risk-score.json, research/jev-accord-compliance-mapping.json, and research/jev-accord-triple-combo-analysis.json in this repository.
  • Comparison baselines: research/jev-compliance-mapping-analysis.json and research/jev-followup-gaps-and-control-analysis.json, and their write-up on the control-plane compliance analysis page.
  • Follow-up: is the accord compatible with Chinese AI regulation, and could c=US support shared US-China control?