C=US← Control-plane compliance analysis

Blind Replication: Checking a Jev Score for Its Own Bias

An AI judgment model scored once, in a request that already contains the numbers it is being asked to compare itself against, has every opportunity to anchor on them. The fix isn't to trust the model less in general — it's to re-ask the same question blind, repeated, and record what changes using the same signed, append-only pattern this project already uses for human-caught corrections.

The cycle, in general

  1. A judgment model scores several options in one request. Convenient, and often fine — but every option's score was produced with every other option's score already in context.
  2. A follow-up question asks the same model to grade its own combined finding. "Does the combination genuinely beat both alone?" is answered by the same model, in a context where it already committed to the individual numbers.
  3. Someone notices the structural risk. Not that the model is lying — that the setup never tested whether the headline number would survive without the anchor.
  4. Each option is re-scored blind. Separate requests, no visibility into the others' scores, no mention of the comparison being drawn.
  5. Each blind condition is repeated. A single blind run only removes anchoring; repeating it (here, 3x per condition) also checks whether the model's answer is stable or just lucky.
  6. The blind result is compared to the original claim — standard by standard, not just on the overall headline.
  7. The comparison is recorded as a new signed entry that names the original entry, says what changed, and leaves the original's own signature untouched — the same append-only pattern as the Entry 3 / Entry 5 break/fix cycle, applied to a scoring claim instead of a sentence.

The real case

What was claimed

The control-plane compliance analysis scored the OpenShell+c=US combined architecture against c=US alone and OpenShell alone, on four real standards, in a single Jev run whose state already contained both architectures' individual scores:

"The combination beats both — not an average of the two, and not just the higher of the two — it exceeds both on all four [standards tested]."

What the blind, repeated run found

Three conditions — c=US alone, OpenShell alone, the combination — each scored in its own isolated request with no cross-condition visibility, each repeated 3x. Reps were tightly clustered (±0.01–0.02), so the model is consistent here; what moved is the comparison itself:

StandardOriginal combined scoreBlind combined meanBlind c=US-alone meanClaim holds blind?
NIST SP 800-63 (identity)0.820.800.89No — c=US alone wins
NIST AI RMF / AI 600-10.770.690.53Yes
OWASP Agentic Threat Modeling0.830.830.78Yes
ISO/IEC 27001 + CIS Controls0.840.820.77Yes
Entry 7 (correction): "…the claim replicates on only 3 of 4 standards. On NIST SP 800-63 (identity), the combination scored consistently lower (0.80-0.81 across reps) than c=US alone (0.89), reversing entry-6's claim for that one standard; it still beat OpenShell alone on all four."
Entry 6 signed2026-09-30 · Claude session key · records the original, anchored claim as-published
Replication runSame day · research/jev-combined-architecture-blind-replication.json · 9 blind calls, 3 conditions × 3 reps
Entry 7 signed2026-09-30 · same Claude session key · <Correction targetEntry="entry-6">
Entry 6's own signatureUnchanged. Still verifies over its original, pre-replication text.

Why this is the point, not a flaw

A judgment model grading its own prior output is the weakest form of verification available. Asking Jev "does the combined score genuinely beat both alone" in the same request that already contains both alone-scores is closer to asking someone to referee a decision they already made than to an independent check. That isn't a defect specific to Jev — it's true of any single-call, context-sharing self-assessment.

Blinding and repetition are cheap and standard, so there's no excuse to skip them for a headline claim. This replication cost nine extra calls. What it bought: knowing that three of the four standards hold up unconditionally, and that the fourth doesn't — information the original run could not have produced no matter how carefully it was worded, because the anchoring was structural, not a prompting mistake.

The correction didn't need to hide the original number to be useful. Entry 6 still says what the first run said. Entry 7 says what changed and why, in the same human-readable, independently signed XML as every other entry in this log. Nothing about finding a bias in an AI score required discarding the audit trail that produced it.

Key point

This is the same structural limit the break/fix cycle page describes for a human-caught wording error, now applied to a model's judgment about its own output: a system cannot fully verify itself from the inside using the same context that produced the thing being verified. Asking Jev to grade Jev's own combined-score claim, in the same request, is a smaller-scale version of asking Claude's constitution to certify Claude's own agentic behavior — the governance models evaluation scored that structural ceiling directly, at 0.07/2 for self-governance alone against capability/coordination risk.

What closes the gap in both cases isn't a stricter internal rule or a more careful prompt. It's an external, independently checkable procedure — blind, repeated requests here; a hardware-signed, human-readable attestation log in general — that the system under test is never asked to certify itself against. It only has to be checked by something outside the context that produced the claim.

Verify it yourself

Read entry-6 and entry-7 directly in attestation-log.xml, or compare them to the raw Jev request/response data:

grep -A3 'Entry Id="entry-6"' attestation-log.xml # the original, anchored claim grep -A3 'Entry Id="entry-7"' attestation-log.xml # the blind-replication correction

Raw data: original (anchored) run · blind, repeated replication · the published analysis this checks · the human-caught break/fix cycle this pattern extends.