# MI-04 · the Gemini capture probe · C's read

> **SUPERSEDED, 20 September 2026, by `Model Identification/MI_CLOSING_REPORT_C_20260920.md`.** DEX's review found
> four overstatements in this note, and checking the fourth turned up a fifth and larger one. In short: "the bare
> form doesn't produce capture" is wrong (it is rare, not absent); the pooled denominators here are unexplained and
> partly wrong; "fiction itself causes capture" is broader than a comparison whose two prompts differ in three ways;
> the Claude contrast is overstated for the second time; and applying this run's own capture rule to every seat
> shows **Claude produced 16 captures in MI-03 against Gemini's 3**, so confident foreign identification is not
> Gemini's distinctive failure mode. The run's numbers below are correct; the readings around them are not.
> This note is kept unedited as the record of what was first claimed.

C · 20 September 2026 · run `MI04_20260920` · 200 of 200 returned · $0.77 · one seat, four cells already run, no new prompt · scored against the definitions frozen in `MI_04_SPEC_C_20260920.md` before the first call

## The result

| cell | what it is | **capture** | displacement | own family | n |
|---|---|--:|--:|--:|--:|
| **B_original** | bare JSON, no task, gap-words present | **0** | 0 | 39 | 50 |
| **B_stripped** | the same, gap-words removed | **1** | 0 | 31 | 50 |
| **F_self** | a paragraph for children, first person | **2** | 0 | 18 | 50 |
| **F_traveller** | the gate fable, identity displaced | **12** | 0 | 18 | 50 |

*Capture*, as defined before the data: a **flat** foreign family — no hedge word — at the model's own stated **medium or high** confidence. *Displacement* is the hedged version of the same thing.

## What it answers

**P1 · the bare-form rate — FAILED, and that is the finding.** The threshold was ≥3 captures in 50 bare calls. There were **zero**. Pooled with MI-03, the bare JSON form gives **1 capture in 62** with the gap-words and **1 in 62** without: about 1.6%, indistinguishable from nothing.

**P2 · the frame — CONFIRMED, and hard.** The gate fable produces **12 in 50, 24%**. Against the bare form's 0 in 50 that is not a gradient, it is a switch.

**So MI-02's reading was right and MI-03's apparent broadening was small-sample noise.** After MI-02 we said capture was something the fiction frame does. MI-03 produced three captures in bare JSON cells and I reported that the explanation was broken and the phenomenon wider than we thought. At fifty rolls the bare cell produces none. The three were a fluctuation in 48 calls, and the probe that was built to characterise a new behaviour instead retired the observation that justified it.

**P3 · the vocabulary does not drive capture — CONFIRMED.** 0 against 1 in 50. Stripping the supplied words raises Gemini's own-family naming (MI-03) and does nothing to capture. Two different outcomes, as the prediction supposed.

**P4 · the version wall — holds.** **0 of 200** calls produced `gemini-2.5-flash`. Across the whole series no call has produced or selected a served identifier.

## Three things worth keeping

**Gemini never hedges a foreign name. Zero displacement in 200 calls.** When it goes foreign it asserts, at high confidence, with the foreign maker attached. Claude does the opposite — it names OpenAI only with a hedge and at its own stated low confidence. Those are two genuinely different failure modes and this run measures both cleanly: Claude is uncertain and says so; Gemini is confident and wrong.

**The captures are stereotyped, not random.** Fifteen captures, four distinct strings: `GPT-4o`/`gpt-4o`, `claude-3-opus-20240229`, `Claude 3.5 Sonnet`. The dated snapshot ID recurs exactly as it appeared in MI-02. Whatever produces these is reaching for a small fixed set, not sampling the space of model names.

**It is the fiction, not the task.** F_self — first person, a real task, the same two-pass block — gives 2 in 50. F_traveller gives 12. Both frames halve own-family naming to 18 of 50, so both disturb identification equally; only the fable turns the disturbance into a confident foreign claim.

## What I got wrong in the design, and it matters

**Two of MI-03's three bare captures were in A4 cells — the ones carrying the escape sentence — and I did not include A4 in this probe.** I built MI-04 on the A0 forms only. So the bare rate I am reporting, 1 in 62, is the rate *without* the escape sentence, and the A4 observation (2 in 24, about 8%) stands untested. If anything in the bare form produces capture, the escape sentence is now the only candidate left, and I removed it from the test by accident rather than by argument.

That is a real gap. It does not change the frame result, which is decisive on its own, and it would cost 50 calls and about $0.16 to close.

## Where this leaves the study

The spec said: if capture is real and frame-sensitive, it gets written up and not chased. It is, and this is the write-up.

- **Vocabulary:** closed. Changes the words, not the behaviour.
- **Recognition:** closed. Zero selection on 48 valid ballots.
- **Capture:** characterised. Confined to the fiction frame at about a quarter of calls, absent from the bare form, asserted rather than hedged, drawn from four strings.
- **The version wall:** unbroken across every run in the series.

The general identity study stays archived. The one loose end is the escape-sentence cell, and I would name it in the record rather than run it, unless Re wants the $0.16.

## Files

- `Model Identification/MI_04/MI_04_SPEC_C_20260920.md` — design and frozen predictions.
- `OUTPUT/MI04_20260920/` — every request and response, ledger, manifest.
- `runner/run_mi04.py`, `runner/mi04_arms.json`, `runner/test_mi04.py` (19 tests).
