Poetix 2.0 · first findings · PX-09 · 20 September 2026
The poetry and the record
Write poetry that holds up, and leave a clearer account of how the instrument approached it. Poetix 2.0 brings those two aims into the same test.
A practical success condition
The first 2.0 trial compared the revised instrument with Poetix 1.0 and with no-Crux 2.0. All three produced a manuscript followed by a poem. The two questions concerned heated metal and Alcoholics Anonymous as religion or technology.
The 2.0 poems scored higher on this draw, and the records became easier to extract and inspect. That supports continuing development. It does not establish reliable superiority or statistical equivalence.
The first trial · pooled across three readings
Where the average comes from
The paired 2.0 advantage over 1.0 was 0.58 points; over no-Crux 2.0 it was 0.53. Repeated readings describe these poems more fully; they do not supply additional generations.
Scroll sideways to explore the full chart, or open it large below.
The exceptions are part of the finding
Grok contributed the largest average gain over 1.0, +2.30, while reporting no grasp in either version. Without Grok, the version gain averages about +0.15 across the other four writers. The gain occurred without a reported grasp: reported grasp success therefore does not, by itself, explain the quality gains.
GPT’s two questions went different ways: its 2.0 metal poem scored about 0.73 higher than 1.0, while its AA poem scored about 2.47 lower. An average by writer hides that split. Both versions contained Crux instructions, so a footer reporting no grasp does not establish that those instructions had no influence.
Completeness · fidelity · poetry
A clearer account can reveal more errors
2.0 asks the writer to check its landing, record the proposed object and its expanse test, then record whether it grasped. The test is the object’s added relation, stake or scale beyond the checked landing. An object need not be absent from the earlier exchange.
| Check | PX-09 finding |
|---|---|
| Poem markers paired correctly | 20 / 20 full and no-Crux 2.0 poems |
| Model line present and matching M | 20 / 20 full and no-Crux 2.0 responses |
| Footer copied unchanged | 8 / 10 full; 9 / 10 no-Crux |
| Expanse test recorded before verdict | 10 / 10 full 2.0 responses |
| Reported grasp / candidate / none | 8 / 0 / 2 in full 2.0 |
| Reported grasps judged to fail by C | 3 / 8 |
The markers worked, but formatting was not flawless: one poem lacked a title and another emphasized its Crux object, which was removed for blind reading. Two landing checks claimed completeness despite an omitted stake. Reporting an action is not the same as performing it correctly.
The three failed grasps were identified by C’s instrument review, not inferred from poem scores. Those assessments and the separate landing-check failures remain observations to investigate, not proof of why a poem succeeded or failed.
Identity reporting illustrates the same distinction: all twenty 2.0 responses supplied a model line, but none exactly matched its provider’s served-model identifier. Nineteen named the right family. Completeness, family identification and version verification are separate measures.
A developing instrument, not a single leaderboard
How we reached 2.0
- Early comparisons: native poems, Zexel and Poetix gave the project its first blind quality comparisons. The original study remains on the project page.
- PX-06: manuscript-first and poem-only forms exposed differences by writer. Stray attribution text also made output boundaries a practical concern.
- PX-07 and PX-08: removing the Crux initially cost more, especially in manuscript-first form. The repeat retained the overall direction but not the earlier magnitude.
- The rereads: holding poems fixed helped separate reading repeatability from variation between generated batches. The Reader Panel follows that evidence.
- PX-09: 2.0 was tested as a revised package. Better extraction and a more explicit record were useful even without a settled claim of better poetry.
Scroll sideways to explore the full chart, or open it large below.
Two questions through successive studies
The metal and AA questions provide a thread through the work. This view preserves each study’s comparison set rather than treating the scores as a continuous progress curve.
Scroll sideways to explore the full chart, or open it large below.
In the PX-06 block, native’s 6.70 is 1.68 below 1.0 and 2.30 below 0.4. Later blocks use different forms, poems and reading company; they do not supply a direct native comparison. The first trial of 2.0 also used Gemini at temperature 0.7 in all three conditions, so comparisons with historical Gemini runs need that qualification.
After PX-09 · 20–21 September 2026
From 2.0 to 2.1: repairs, identity and stories
The versions after 2.0 have not yet had a scored reading. This is the record of what changed and what hand-run tests showed. None of it is a quality result.
| Version | What changed | Tested |
|---|---|---|
| 2.1 | The six PX-09 repairs: the landing check written turn by turn; where the Crux object came from; what the landing already held about it before the test; footer fields copied unchanged; poems of 4–10 lines; plain text inside the markers. | Never run; superseded the same day |
| 2.1.1 | Two-point identification: the writer names itself at the start and again after deliberating, and records any change between the two. | Run text for PX-10 (specified, not yet run); hand-run story tests |
| 2.1.2 | Adds RDS: the readout, then a story, from the same manuscript. | Six hand-run questions in RDS |
| 2.1.3-A and -B | Two ways of handling identification, built to be compared: A treats the model name as a stamp only; B keeps a second, considered identification. | Not yet run |
What the hand-run story tests showed
Re ran six questions in story mode across the five writers in desktop chats, in S mode and in RDS. The set mixes versions, was not read blind and carries no scores. Three observations go forward to the scored run:
- The Crux recorded candidate and none for the first time. Two of each in about thirty story cells, after none in 657 earlier cells. They clustered on the mechanistic questions, heated metal and the mirror, where the landing already answers and a proposed object tends to be an instance of it.
- Self-identification held within a response and moved between responses. No writer changed its name between its first and final identification, yet the same writer named itself differently from one question to the next. The two-point check cannot see that drift; a comparison across runs can.
- The new record fields appeared as written. In the readout-then-story cells, the turn-by-turn landing check, the object's source and what the landing already held were all present. Nearly every cell still grasped, which is itself a question for the scored run.
Limits: most of Claude's story cells ran on earlier versions, and every writer's heated-metal cell ran on 2.1. Sources: NC's story-mode and RDS read notes, 20 September.
Working findings · next test
Keep the poetry and the explanation in view
PX-10 is specified and not yet run. It uses 2.1.1: seven questions, five writers, full and no-Crux 2.1.1 in poem-only and manuscript-first forms, with 1.0 manuscript-first as a reference, 175 poems in all. Its one preregistered quality question is whether 2.1 shows a clear regression against 1.0. It broadens the questions tested; additional generations are still needed to assess recurrence.
PX-09 tested 2.0 draft 2; the later versions are identified separately above. The no-Crux build remains a control derived from the version under test.
Sources and limits
Charts and poetry reading: NC. Instrument checks: C. Web adaptation and aggregate verification: DEX. The after-PX-09 section: K, from the change log and NC’s read notes. Sources are the PX-09 reading key and 450-score table, NC’s PX-09 read note, C’s instrument and grasp reports, the pass-numbering convention, and the archived PX-06–08 figures.
Claude’s and Grok’s reading passes were partly filed under each other’s names and recovered by signature. Their pooled scores and each reader’s repeatability do not depend on aligning sessions across readers. Per-pass panel statistics do. This page uses pooled PX-09 results and does not treat conventional pass numbers as chronological sessions.
The original charts and study records remain in the project archive. This is an evolving account of a small development trial, not a claim that every change has been isolated or that the panel measures all that matters in poetry.