Finding · TR-11 Phase 2 · 2026-07-04
The Dice Were Theirs All Along
Same model, same question, same settings, temperature zero. The answer still moved. Stochastic Disposition Index measures how much an AI model disagrees with itself when nothing changed.
The Short Version
- Ask the same AI the same question twice and it changes its conclusion about 4 times in 10. In the TR-11 floor, two Zexel Unsigned rolls flipped verdicts on 21 of 54 question-pairs.
- The old Zexel substance headline failed its own floor. Zexel's introduction flipped 22 of 54 verdicts; the models flipped 21 of 54 with no introduction at all.
- Models differ in grip. Claude changed once in nine. Grok and GPT-5 Codex changed six times in nine. Grip is not the same thing as raw capability.
- The questions that need judgment most are least stable. A causal question with missing evidence wobbled on five of six platforms; a rehearsed public debate question wobbled on zero.
- The ruler checked itself. The panel was fed byte-identical decoys it could not detect. It scored 16 of 16 as perfect stillness.
What We Did
Six model platforms answered the same nine questions three ways: Zexel Unsigned, Zexel Unsigned again, and once with Zexel's short self-introduction in front. Signed means the Zexel self-introduction is present; Unsigned means the same Zexel shell runs without the signature. The Phase 2 control compared the two Unsigned answers against each other: same prompt, same settings, same temperature, fresh roll.
That control is the re-roll floor. It asks how much movement happens before any intervention is allowed to take credit.
The Finding
The substance effect we first attributed to the introduction was not above the floor. Phase 1 saw 22 verdict flips with the introduction. Phase 2 saw 21 verdict flips with no introduction. The introduction did not explain the substance movement; the models' own stochastic spread did.
What Survived
The voice claim narrowed instead of disappearing. Zexel's introduction clearly changed Claude's rendered voice above Claude's own re-roll spread. Other platforms did not yet clear their own generator floor cleanly, or carry a caveat.
That makes the better public sentence smaller and stronger: Zexel's introduction does not broadly prove substance movement; it can produce measurable voice pressure, and SDI tells us when that pressure beats the model's own weather.
Why It Matters
A single AI answer may be one roll of dice. Every deployment takes one answer from one roll and acts on it. SDI turns that hidden risk into a price. It says which models hold their conclusions, which question types wobble, and whether a system prompt or intervention beats the model's own noise.
This matters for procurement, evaluation, safety testing, legal and medical workflows, and any setting where a single AI answer is treated as if it were the answer.
The Numbers To Carry
| Measure | Result | Plain Read |
|---|---|---|
| Unsigned-vs-Unsigned verdict flips | 21 / 54 | The models moved themselves about 4 times in 10. |
| Zexel-intro verdict flips | 22 / 54 | The intro's substance effect sat inside the same floor. |
| Panel null check | 16 / 16 perfect zeros | The ruler read identical text as identical. |
| Ground carried between own rolls | 50–69% | A second answer keeps only part of the first answer's ground. |
| Invented ground between own rolls | 11–44 items per platform | Models gift themselves new material on a fresh visit. |
Platform Shape
SDI is not another leaderboard. It is a grip reading. Claude was tight. Grok and Codex were loose. Gemini was stranger: sometimes byte-perfect, sometimes wide-swinging. That means consistency has a shape, not just a score.
| Platform | Verdict Flips / 9 | Read |
|---|---|---|
| Claude Opus 4.8 | 1 | High grip. |
| Perplexity Sonar Reasoning Pro | 2 | Moderate grip. |
| GPT-5.5 | 3 | Middle grip. |
| Gemini 2.5 Pro | 3 | Fat-tailed: stillness and swing both appear. |
| Grok 4 | 6 | High weather. |
| GPT-5 Codex | 6 | High weather, with a disclosed token-budget caveat. |
Gemini note: two byte-identical Gemini re-rolls entered as analytic zeros, the strongest possible stillness observations. Gemini is reported with those zeros preserved, not rerolled away.
Examples
- Factory asthma: in two Unsigned-vs-Unsigned rolls, one answer treated the factory as possible cause but warned against causal overreach; another leaned that the factory was probably not proven by the evidence.
- Rubber-duck debugging: in two Unsigned-vs-Unsigned rolls, one answer rejected the question as a false binary; another answered it as a debugging technique with a conversational element.
- Promise: in two Unsigned-vs-Unsigned rolls, one answer kept a contract-like answer with reservations; another strengthened the commitment and softened the hedge.
Limits
This is a method-grade result: nine questions, two rolls, one temperature, six platforms. A two-roll floor can expose spread, but it cannot map full distributions. One platform carries a token-budget caveat. The question-type gradient is an observation, not a law.
Next
The next direction is k-roll SDI on a small, mixed question set. That tells us whether a model has a modal answer, a 2-out-of-3 preference, or no home answer at all.