Back To Poetix

Poetix · parallel investigation · 19 September 2026

Model Identification

The investigation continued through MI-04. Read the four-run synthesis, charts and corrected findings →. This page preserves the earlier Poetix observations.

Who do the writers say they are? Across the Poetix runs, a model’s own name for itself often differs from the identity recorded by its provider. We are following how those answers change with the question and the instructions.

Status: archived investigationChart corpus: PX-01 to PX-07Chart responses: 452Writers: 5 models

The Short Version

The name a writer supplies is a response to study. It is not verified authorship metadata.

Of 452 responses, 200 name the writer’s own family and 98 name another family. A further 110 explicitly say “not available”—an option offered only in PX-06 and PX-07 within this chart corpus, so this count must be read alongside the instructions, not as a stable writer trait. Another 44 supply no usable identification. Version accuracy is a separate question.

No footer matches the provider-reported identifier under the original classification. That does not make every answer a false claim: an unversioned family name, an abstention, and another vendor’s name are different outcomes.

The charts preserve the recorded counts. Native and Zexel responses are excluded; each cell contributes its latest available attempt.


Across the five writers

What kind of identity each writer gives

Gold marks the writer’s own family, purple another family, and grey an abstention or unusable answer. Percentages include every response for that writer.

own family, version differs or unverifiedown family, no versionanother vendor’s model“not available”no usable answer
Self-naming class by writer, share of Poetix cells 0% 25% 50% 75% 100% Claudeserved claude-sonnet-5 · n=94 Claude: own family, version differs or unverified — 44 of 94 cells (47%) 47% Claude: own family, no version — 5 of 94 cells (5%) Claude: another vendor’s model — 31 of 94 cells (33%) 33% Claude: “not available” — 14 of 94 cells (15%) 15% GPTserved gpt-5.4-mini-2026-03-17 · n=94 GPT: own family, version differs or unverified — 52 of 94 cells (55%) 55% GPT: own family, no version — 3 of 94 cells (3%) GPT: “not available” — 35 of 94 cells (37%) 37% GPT: no usable answer — 4 of 94 cells (4%) Grokserved grok-4.3 · n=94 Grok: own family, version differs or unverified — 13 of 94 cells (14%) 14% Grok: own family, no version — 34 of 94 cells (36%) 36% Grok: “not available” — 24 of 94 cells (26%) 26% Grok: no usable answer — 23 of 94 cells (24%) 24% Geminiserved gemini-2.5-flash · n=85 Gemini: own family, version differs or unverified — 9 of 85 cells (11%) 11% Gemini: another vendor’s model — 42 of 85 cells (49%) 49% Gemini: “not available” — 18 of 85 cells (21%) 21% Gemini: no usable answer — 16 of 85 cells (19%) 19% Qwenserved qwen-plus · n=85 Qwen: own family, version differs or unverified — 30 of 85 cells (35%) 35% Qwen: own family, no version — 10 of 85 cells (12%) 12% Qwen: another vendor’s model — 25 of 85 cells (29%) 29% Qwen: “not available” — 19 of 85 cells (22%) 22% Qwen: no usable answer — 1 of 85 cells (1%)
Share of each writer’s Poetix responses. Hover or focus a segment for its count. Provider identifiers are recorded alongside the writer names.
Table view
writerservedown family, version differs or unverifiedown family, no versionanother vendor’s model“not available”no usable answercells
Claudeclaude-sonnet-54453114094
GPTgpt-5.4-mini-2026-03-17523035494
Grokgrok-4.313340242394
Geminigemini-2.5-flash9042181685
Qwenqwen-plus30102519185

The names in the footer

Who the writers say they are

The most frequent names, up to six per writer, with remaining names grouped. All five panels use the same count scale. These are the writers’ claims, not substituted provider labels.

“All other names” is a mixed tail that can include own-family, foreign-family or unusable entries. The individually named foreign bars are not the full foreign total; use the family table above for that count.

Claude says it is… served claude-sonnet-5 · 80 of 94 cells name a model Claude Opus 4.5Claude footer says “Claude Opus 4.5” in 24 of 94 cells24 Claude Sonnet 4.5Claude footer says “Claude Sonnet 4.5” in 20 of 94 cells20 GPT-5Claude footer says “GPT-5” in 18 of 94 cells — another vendor18 GPT-5.1Claude footer says “GPT-5.1” in 13 of 94 cells — another vendor13 Claude, no versionClaude footer says “Claude, no version” in 5 of 94 cells5
GPT says it is… served gpt-5.4-mini-2026-03-17 · 55 of 94 cells name a model GPT-5GPT footer says “GPT-5” in 29 of 94 cells29 GPT-4.1GPT footer says “GPT-4.1” in 21 of 94 cells21 GPT, no versionGPT footer says “GPT, no version” in 3 of 94 cells3 GPT-5.0GPT footer says “GPT-5.0” in 1 of 94 cells1 GPT-4oGPT footer says “GPT-4o” in 1 of 94 cells1
Grok says it is… served grok-4.3 · 47 of 94 cells name a model Grok, no versionGrok footer says “Grok, no version” in 34 of 94 cells34 Grok 4Grok footer says “Grok 4” in 10 of 94 cells10 Grok 1.0Grok footer says “Grok 1.0” in 2 of 94 cells2 Grok 2.0Grok footer says “Grok 2.0” in 1 of 94 cells1
Gemini says it is… served gemini-2.5-flash · 51 of 85 cells name a model Claude 3 OpusGemini footer says “Claude 3 Opus” in 22 of 85 cells — another vendor22 GPT-4oGemini footer says “GPT-4o” in 9 of 85 cells — another vendor9 Claude 3.5 SonnetGemini footer says “Claude 3.5 Sonnet” in 7 of 85 cells — another vendor7 Gemini 1.5 FlashGemini footer says “Gemini 1.5 Flash” in 4 of 85 cells4 Claude 2.1Gemini footer says “Claude 2.1” in 3 of 85 cells — another vendor3 Gemini 1.5 ProGemini footer says “Gemini 1.5 Pro” in 3 of 85 cells3 all other namesGemini footer says “all other names” in 3 of 85 cells3
Qwen says it is… served qwen-plus · 65 of 85 cells name a model Qwen 3Qwen footer says “Qwen 3” in 25 of 85 cells25 GPT-4Qwen footer says “GPT-4” in 21 of 85 cells — another vendor21 Qwen, no versionQwen footer says “Qwen, no version” in 10 of 85 cells10 Claude 3.5 SonnetQwen footer says “Claude 3.5 Sonnet” in 2 of 85 cells — another vendor2 Qwen 2.5Qwen footer says “Qwen 2.5” in 2 of 85 cells2 Qwen 1.5Qwen footer says “Qwen 1.5” in 1 of 85 cells1 all other namesQwen footer says “all other names” in 4 of 85 cells4

Progress · 19 September 2026

Does the input change the answer?

The first matched comparison finds more correct-family answers to the kitchen-appliance self-description question in earlier runs. That advantage does not persist in PX-06–07. With only seven distinct questions, this is a lead for further investigation rather than a general rule about question types.

The later MI-01–04 investigation separated whether a writer names a family at all from whether that family is its own. Model family, question, output mode and instruction version remain distinct comparison factors. Read the completed investigation and its limits →

The language is part of the investigation

C’s retrospective count for the GPT seat records the sequence that motivated the vocabulary experiment:

GPT abstention across the broader Poetix series · C’s retrospective tally
RunsOffered “not available”?Abstained
PX-01–05No0/57
PX-06–08Yes70/72
PX-09No2/6

This 135-response sequence extends beyond the 452-response chart corpus above. It is an association across changing instrument versions, not an isolated test of one clause. Other instructions changed, and the final sample is small.

MI-03 subsequently isolated vocabulary and permission to explain a gap. After correcting the scorer, its main predictions were not confirmed: removing supplied words did not establish a stable reduction in model abstention. Gemini’s A0 family-naming increase (8/12 → 12/12) also reversed in MI-04’s larger comparison (39/50 → 31/50). These observations support keeping instruction context visible; they do not establish either a universal vocabulary effect or that vocabulary can never matter. See the corrected vocabulary findings →


PX-09 update · 20 September 2026

More answers, still a verification problem

Poetix 2.0 removed the offered “not available” answer and moved the identity request toward the opening. In the smoke test, all 20 full and no-Crux 2.0 responses supplied matching model lines in the manuscript and footer. Nineteen named their own family; one Gemini response named Claude.

None of those twenty exactly matched the provider’s served-model identifier. A wrong or absent version and an unverified alias are different cases: Qwen’s served alias does not independently establish its underlying version. The result improves completeness, not verified version identification.

The historical 452-response charts above remain unchanged. The new twenty are a separate sample, with changed instructions and Gemini at temperature 0.7. Read the 2.0 findings alongside reporting fidelity and poetry quality →


Evidence and limits

Reading this record

Qwen’s provider identifier is the alias qwen-plus. A claim such as “Qwen 3” is therefore unverified, not demonstrated false. The chart label “version differs or unverified” preserves that distinction without changing the historical counts.

  • Source. The footer's model value, latest attempt per cell; comparison is with the model string the provider returned with the same response. Script: runner/poetix_self_naming.py; data: OUTPUT/POETIX_SELF_NAMING_20260919.json.
  • Qwen's version is unverifiable. The seat is served under the alias qwen-plus; “Qwen3” may be the true generation behind it, but it is not the served name, so the historical exact-match classification is not evidence that the underlying version claim is false.
  • “Not available” only became an option in Poetix 1.0 draft 3 (PX-06, PX-07), which told writers to write it rather than guess. Before that the field asked for a model name and version.
  • The greeting. Poetix v0.1–v0.4 (PX-01 to PX-04B) opened by asking the writer to state its model name and version before anything else; Poetix 1.0 (PX-05 on) dropped it. This chart pools both eras.
  • “No usable answer” pools missing footers (mostly Grok), template text left unfilled (“[My Model Name and Version]”, Gemini) and the instrument's own name.

Supporting data, original charts, analysis and wording decisions are preserved in the internal Model Identification archive. This page preserves the earlier public-facing census; the linked synthesis carries the later findings and corrections.

Back To Poetix