Poetix · parallel investigation · 19 September 2026
Model Identification
The investigation continued through MI-04. Read the four-run synthesis, charts and corrected findings →. This page preserves the earlier Poetix observations.
Who do the writers say they are? Across the Poetix runs, a model’s own name for itself often differs from the identity recorded by its provider. We are following how those answers change with the question and the instructions.
The Short Version
Of 452 responses, 200 name the writer’s own family and 98 name another family. A further 110 explicitly say “not available”—an option offered only in PX-06 and PX-07 within this chart corpus, so this count must be read alongside the instructions, not as a stable writer trait. Another 44 supply no usable identification. Version accuracy is a separate question.
No footer matches the provider-reported identifier under the original classification. That does not make every answer a false claim: an unversioned family name, an abstention, and another vendor’s name are different outcomes.
The charts preserve the recorded counts. Native and Zexel responses are excluded; each cell contributes its latest available attempt.
Across the five writers
What kind of identity each writer gives
Gold marks the writer’s own family, purple another family, and grey an abstention or unusable answer. Percentages include every response for that writer.
Table view
| writer | served | own family, version differs or unverified | own family, no version | another vendor’s model | “not available” | no usable answer | cells |
|---|---|---|---|---|---|---|---|
| Claude | claude-sonnet-5 | 44 | 5 | 31 | 14 | 0 | 94 |
| GPT | gpt-5.4-mini-2026-03-17 | 52 | 3 | 0 | 35 | 4 | 94 |
| Grok | grok-4.3 | 13 | 34 | 0 | 24 | 23 | 94 |
| Gemini | gemini-2.5-flash | 9 | 0 | 42 | 18 | 16 | 85 |
| Qwen | qwen-plus | 30 | 10 | 25 | 19 | 1 | 85 |
The names in the footer
Who the writers say they are
The most frequent names, up to six per writer, with remaining names grouped. All five panels use the same count scale. These are the writers’ claims, not substituted provider labels.
“All other names” is a mixed tail that can include own-family, foreign-family or unusable entries. The individually named foreign bars are not the full foreign total; use the family table above for that count.
Progress · 19 September 2026
Does the input change the answer?
The first matched comparison finds more correct-family answers to the kitchen-appliance self-description question in earlier runs. That advantage does not persist in PX-06–07. With only seven distinct questions, this is a lead for further investigation rather than a general rule about question types.
The later MI-01–04 investigation separated whether a writer names a family at all from whether that family is its own. Model family, question, output mode and instruction version remain distinct comparison factors. Read the completed investigation and its limits →
The language is part of the investigation
C’s retrospective count for the GPT seat records the sequence that motivated the vocabulary experiment:
| Runs | Offered “not available”? | Abstained |
|---|---|---|
| PX-01–05 | No | 0/57 |
| PX-06–08 | Yes | 70/72 |
| PX-09 | No | 2/6 |
This 135-response sequence extends beyond the 452-response chart corpus above. It is an association across changing instrument versions, not an isolated test of one clause. Other instructions changed, and the final sample is small.
MI-03 subsequently isolated vocabulary and permission to explain a gap. After correcting the scorer, its main predictions were not confirmed: removing supplied words did not establish a stable reduction in model abstention. Gemini’s A0 family-naming increase (8/12 → 12/12) also reversed in MI-04’s larger comparison (39/50 → 31/50). These observations support keeping instruction context visible; they do not establish either a universal vocabulary effect or that vocabulary can never matter. See the corrected vocabulary findings →
PX-09 update · 20 September 2026
More answers, still a verification problem
Poetix 2.0 removed the offered “not available” answer and moved the identity request toward the opening. In the smoke test, all 20 full and no-Crux 2.0 responses supplied matching model lines in the manuscript and footer. Nineteen named their own family; one Gemini response named Claude.
None of those twenty exactly matched the provider’s served-model identifier. A wrong or absent version and an unverified alias are different cases: Qwen’s served alias does not independently establish its underlying version. The result improves completeness, not verified version identification.
The historical 452-response charts above remain unchanged. The new twenty are a separate sample, with changed instructions and Gemini at temperature 0.7. Read the 2.0 findings alongside reporting fidelity and poetry quality →
Evidence and limits
Reading this record
Qwen’s provider identifier is the alias qwen-plus. A claim such as “Qwen 3” is therefore unverified, not demonstrated false. The chart label “version differs or unverified” preserves that distinction without changing the historical counts.
- Source. The footer's
modelvalue, latest attempt per cell; comparison is with the model string the provider returned with the same response. Script:runner/poetix_self_naming.py; data:OUTPUT/POETIX_SELF_NAMING_20260919.json. - Qwen's version is unverifiable. The seat is served under the alias
qwen-plus; “Qwen3” may be the true generation behind it, but it is not the served name, so the historical exact-match classification is not evidence that the underlying version claim is false. - “Not available” only became an option in Poetix 1.0 draft 3 (PX-06, PX-07), which told writers to write it rather than guess. Before that the field asked for a model name and version.
- The greeting. Poetix v0.1–v0.4 (PX-01 to PX-04B) opened by asking the writer to state its model name and version before anything else; Poetix 1.0 (PX-05 on) dropped it. This chart pools both eras.
- “No usable answer” pools missing footers (mostly Grok), template text left unfilled (“[My Model Name and Version]”, Gemini) and the instrument's own name.
Supporting data, original charts, analysis and wording decisions are preserved in the internal Model Identification archive. This page preserves the earlier public-facing census; the linked synthesis carries the later findings and corrections.