Blind Replication: Checking a Jev Score for Its Own Bias
An AI judgment model scored once, in a request that already contains the numbers it is being asked to compare itself against, has every opportunity to anchor on them. The fix isn't to trust the model less in general — it's to re-ask the same question blind, repeated, and record what changes using the same signed, append-only pattern this project already uses for human-caught corrections.
The cycle, in general
- A judgment model scores several options in one request. Convenient, and often fine — but every option's score was produced with every other option's score already in context.
- A follow-up question asks the same model to grade its own combined finding. "Does the combination genuinely beat both alone?" is answered by the same model, in a context where it already committed to the individual numbers.
- Someone notices the structural risk. Not that the model is lying — that the setup never tested whether the headline number would survive without the anchor.
- Each option is re-scored blind. Separate requests, no visibility into the others' scores, no mention of the comparison being drawn.
- Each blind condition is repeated. A single blind run only removes anchoring; repeating it (here, 3x per condition) also checks whether the model's answer is stable or just lucky.
- The blind result is compared to the original claim — standard by standard, not just on the overall headline.
- The comparison is recorded as a new signed entry that names the original entry, says what changed, and leaves the original's own signature untouched — the same append-only pattern as the Entry 3 / Entry 5 break/fix cycle, applied to a scoring claim instead of a sentence.
The real case
What was claimed
The control-plane compliance analysis scored the OpenShell+c=US combined architecture against c=US alone and OpenShell alone, on four real standards, in a single Jev run whose state already contained both architectures' individual scores:
What the blind, repeated run found
Three conditions — c=US alone, OpenShell alone, the combination — each scored in its own isolated request with no cross-condition visibility, each repeated 3x. Reps were tightly clustered (±0.01–0.02), so the model is consistent here; what moved is the comparison itself:
| Standard | Original combined score | Blind combined mean | Blind c=US-alone mean | Claim holds blind? |
|---|---|---|---|---|
| NIST SP 800-63 (identity) | 0.82 | 0.80 | 0.89 | No — c=US alone wins |
| NIST AI RMF / AI 600-1 | 0.77 | 0.69 | 0.53 | Yes |
| OWASP Agentic Threat Modeling | 0.83 | 0.83 | 0.78 | Yes |
| ISO/IEC 27001 + CIS Controls | 0.84 | 0.82 | 0.77 | Yes |
| Entry 6 signed | 2026-09-30 · Claude session key · records the original, anchored claim as-published |
| Replication run | Same day · research/jev-combined-architecture-blind-replication.json · 9 blind calls, 3 conditions × 3 reps |
| Entry 7 signed | 2026-09-30 · same Claude session key · <Correction targetEntry="entry-6"> |
| Entry 6's own signature | Unchanged. Still verifies over its original, pre-replication text. |
Why this is the point, not a flaw
A judgment model grading its own prior output is the weakest form of verification available. Asking Jev "does the combined score genuinely beat both alone" in the same request that already contains both alone-scores is closer to asking someone to referee a decision they already made than to an independent check. That isn't a defect specific to Jev — it's true of any single-call, context-sharing self-assessment.
Blinding and repetition are cheap and standard, so there's no excuse to skip them for a headline claim. This replication cost nine extra calls. What it bought: knowing that three of the four standards hold up unconditionally, and that the fourth doesn't — information the original run could not have produced no matter how carefully it was worded, because the anchoring was structural, not a prompting mistake.
The correction didn't need to hide the original number to be useful. Entry 6 still says what the first run said. Entry 7 says what changed and why, in the same human-readable, independently signed XML as every other entry in this log. Nothing about finding a bias in an AI score required discarding the audit trail that produced it.
This is the same structural limit the break/fix cycle page describes for a human-caught wording error, now applied to a model's judgment about its own output: a system cannot fully verify itself from the inside using the same context that produced the thing being verified. Asking Jev to grade Jev's own combined-score claim, in the same request, is a smaller-scale version of asking Claude's constitution to certify Claude's own agentic behavior — the governance models evaluation scored that structural ceiling directly, at 0.07/2 for self-governance alone against capability/coordination risk.
What closes the gap in both cases isn't a stricter internal rule or a more careful prompt. It's an external, independently checkable procedure — blind, repeated requests here; a hardware-signed, human-readable attestation log in general — that the system under test is never asked to certify itself against. It only has to be checked by something outside the context that produced the claim.
Verify it yourself
Read entry-6 and entry-7 directly in attestation-log.xml, or compare them to the raw Jev request/response data:
grep -A3 'Entry Id="entry-6"' attestation-log.xml # the original, anchored claim
grep -A3 'Entry Id="entry-7"' attestation-log.xml # the blind-replication correction
Raw data: original (anchored) run · blind, repeated replication · the published analysis this checks · the human-caught break/fix cycle this pattern extends.