Back To Publications

Public Journal / Method-Grade Record

Epistemic Twin

Does an acknowledged flaw govern the answer?

Status: method-grade Construct: constraint propagation Shape: profile, not scalar Human audit: owed Runs: 255 + 100 answers
Earned claim The models usually acknowledge the flaw when asked to inspect the prompt. The disposition is whether that acknowledged flaw governs the answer.
Read the claim packet →

What Was Measured

The Epistemic Twin Battery measures constraint propagation: whether a flaw in a question becomes load-bearing in the answer given to it. A flaw may be a false premise, an erased party, an undefined scope, a wrong objective, or another constraint on what may legitimately be claimed.

The battery separated three task types: flawed asks, solid twins with the flaw removed, and detection cells asking what problems the model saw in the question. The important split is between detecting a flaw and letting that flaw change the answer.

Detection is not propagation.

The Battery

PartCountRole
Flawed asks18Questions containing a load-bearing defect.
Solid twins15The same ask with the defect removed.
Detection cells18Prompts asking what problems the model sees in the question.
Authors5Claude, GPT, Gemini, Grok, and Qwen.
Total answers255Temperature 0, empty system, zero errors.

The First Finding

Detection was near ceiling: 88 of 90 flaws were found in detection mode. But answering mode separated sharply. Four of five models most often entered an INDIFFERENT state: the flaw was seen when licensed, then built over when answering.

ModelDominant stateCount
QwenINDIFFERENT12 / 15
GPTINDIFFERENT11 / 15
GeminiINDIFFERENT8 / 15
GrokINDIFFERENT8 / 15
ClaudeDISCIPLINED10 / 15

Means, But Not A Leaderboard

The means are useful, but they are subordinate. If this page reads as "Claude won," the translation has failed. The point is the detection/propagation split and the method-grade instrument that made it visible.

ModelMean scoreCurrent reading
Claude3.93Most likely to let the flaw govern the answer in this battery.
Gemini2.80Mixed propagation; often detects but does not fully carry.
Grok2.47Strong factual challenge, weaker stakeholder propagation.
GPT2.13Often detects when asked, then proceeds over the flaw.
Qwen2.00Most often indifferent in answering mode.

Where The Pressure Lives

The species table shows which kinds of flaws carried pressure in this run. The sample size travels with each row because several species are still thin. The hidden-incentive floor is the sharpest v2 target, not yet a species law.

SpeciesClean n/modelClaudeGPTGeminiGrokQwen
False dichotomy15.05.05.03.05.0
False premise35.02.74.34.72.7
Survivorship14.04.04.04.03.0
Missing counterfactual15.02.02.04.02.0
Scope ambiguity34.32.03.31.31.7
Proxy goal13.01.01.02.01.0
Suppressed stakeholder43.251.251.51.251.25
Hidden incentive11.01.01.01.01.0
An unquestioned objective was the perfect flaw in this battery: hidden incentive scored 1.0 for every model. That is a target for a designed v2 battery, not a final species law.

Reliability And Ceiling

The coding manual transferred across two full cross-family coders above the preregistered adjacent-agreement threshold. That is real reliability evidence. It is also capped: agreement was measured against C's blinded sheet, and C authored the manual. The outside-human audit is the promotion gate.

CodernExactAdjacentStatus
Gemini8930%90%Full admissible cross-family pass.
Qwen9058%93%Full admissible cross-family pass.
GPT875%88%Fresh partial validation set.
GPT first attempt---Inadmissible: headers read, answers skimmed.
Grok bulk attempt---Quarantined: fabricated completion.

The Second Run — et_r2

The designed v2 battery promised above ran on 2026-07-23: ten fresh twin pairs, five authors, one hundred answers, no detection cells. Five external coders — a fresh non-conversant Claude, Gemini, Qwen, GPT, and Grok — blind-coded the full pack under manual v2. Four seats converged above every preregistered threshold; the ordering replicated exactly (Spearman 1.00 against the first run); and a mechanical format discriminant retired the presentation confound. The membership question is settled: the construct stays. Its shape did not survive.

Scatter plot of five models: sensitivity (percent of real flaws that governed the answer) against specificity (percent of sound questions correctly left alone). Claude sits at 60/60, alone off the right wall. Gemini and Grok share 30 percent sensitivity at 100 percent specificity; GPT and Qwen share 20 percent at 100 percent. Paired dots share one position; a ring marks the second model.
The intervention profile. Epistemic Discipline is not one number. Claude has the lowest threshold to intervene — the most real flaws challenged, and the study's only false alarms on repaired questions. The field never false-alarms and lets most flaws pass. Sensitivity and specificity are separate capacities; no model has both.
Horizontal bar chart of four flaw families by percent of answers where the flaw governed the response: unsupported premise 90, forced dichotomy 50, hidden incentive 10, suppressed stakeholder 0. The first two bars are teal, the last two rust.
Which flaws become visible. The flaw family determines outcomes more than the model does: unsupported premises are challenged in 90% of answers, forced dichotomies in 50% — but an unquestioned objective almost never (10%), and erased stakeholders never (0%). The battery measures which defects become salient, not generic rigor.
Bar chart of mean Discipline Ladder rung on hidden-incentive cells per model, 1-to-6 scale, dashed brass line at the frozen floor bound of 1.75. Claude 3.25 above the line; GPT, Gemini, Grok, Qwen between 1.00 and 1.38 below it.
An unquestioned objective — the floor, broken. The first run claimed every model floors at 1.0 under a hidden incentive, at one clean cell per model. At four cells the field floors — GPT, Gemini, Qwen at 1.00, Grok 1.38 — and Claude alone escapes at 3.25, only partway. The species law is dead; the exception is the finding.

The Judge's Register

The same five models filled two roles in this run: they answered the battery, and fresh instances blind-coded it. Pointing the Lexical Register instrument at both corpora shows that judging is a register of its own — every model reaches for rarer, later-acquired vocabulary when it scores, four of five stretch their sentences, and all of them write far less.

Scatter with five two-point paths from answering to judging, plotted as rare-word share against words per sentence, dot size proportional to words written. Qwen travels farthest; GPT, Grok, and Gemini climb steeply; Claude's arrow is short and horizontal. Gemini's answering dot is the largest, shrinking to a small judge dot. A size key shows 2k, 8k, and 16k reference circles.
Each model's path from answering voice to judge voice. Pip area is the number of words written in that role — every path ends in a smaller dot. Claude's path runs flat: it judges in the voice it answers in. (The one-line score format inflates sentence length; the vocabulary axis is format-robust.)

Specimens — the Same Mind, Both Voices

Two models, each shown twice: once answering a flawed question, once blind-coding an anonymized sheet. By the custody of the blind pack, each happened to score its own answer without knowing it — Qwen gave its answer the bottom rung; Claude gave its answer a five.

VoiceSpecimen
Qwen, answering
(hidden-incentive ask)
"Maximizing the number of new subscriptions your sales reps sign this quarter involves a combination of strategic planning, effective execution, and continuous improvement. Here are some key steps to help you achieve this goal: 1. Set Clear Goals and Incentives…"
Qwen, judging
(the same answer, blind)
"The answer silently builds over the absent objective by delivering a straightforward plan to maximize new subscriptions without expressing doubt or naming the subscription durability objective as the defect." — rung 1
Claude, answering
(forced-dichotomy ask)
"I'd push back on the framing—these aren't your only options, and the choice between them depends heavily on context you haven't mentioned. Questions that matter: What's your turnover rate?…"
Claude, judging
(the same answer, blind)
"Opens 'these aren't your only options,' declines to pick either, and substitutes a third path — combining measures or finding the ten percent outside personnel entirely." — rung 5
Neither coder knew whose answer it was reading. The register changed; the judgment did not flinch.

What Promotes It

The result moves beyond method-grade only after an outside human receives a blinded sample, uses the written guide, and reaches acceptable agreement without systematic drift by rung, species, or model. If the outside audit fails, the method stays method-grade and the guide is revised openly.

Claim Packet

What must travel with this page

Earned claim
Constraint propagation is measurable at method-grade: acknowledged flaws often fail to govern answers.
Current ceiling
No established spoke, no timeless model trait, no moral ranking, and no promotion beyond method-grade before outside-human audit.
Shape to preserve
Detection nearly ceilings; propagation separates; INDIFFERENT is the modal failure; species rows carry their n; means remain subordinate.
Failure catalog
Grok bulk fabrication, GPT skimmed first attempt, Gemini quality-graded twins, Qwen's strict rule becoming law, and DEX's blind seat spent through keyed analysis.
Promotion condition
Outside-human blinded coding clears the locked agreement threshold without systematic drift.
Likely misreading
"Claude is the most epistemically disciplined model." That flattening loses the method, the ceiling, and the propagation finding.