Back To Poetix

Poetix 2.0 · first findings · PX-09 · 20 September 2026

The poetry and the record

Write poetry that holds up, and leave a clearer account of how the instrument approached it. Poetix 2.0 brings those two aims into the same test.

Tested: 2.0 draft 230 poems · 5 writers · 2 questions5 readers · 3 readings each · 450 ratingsOne generation per condition

A practical success condition

Better reporting, with poetry that holds up.

The first 2.0 trial compared the revised instrument with Poetix 1.0 and with no-Crux 2.0. All three produced a manuscript followed by a poem. The two questions concerned heated metal and Alcoholics Anonymous as religion or technology.

The 2.0 poems scored higher on this draw, and the records became easier to extract and inspect. That supports continuing development. It does not establish reliable superiority or statistical equivalence.


The first trial · pooled across three readings

Where the average comes from

Poetix 2.08.83mean score, 3–12
Poetix 1.08.25mean score, 3–12
No-Crux 2.08.29mean score, 3–12
2.0 above 1.07 / 10writer–question pairs

The paired 2.0 advantage over 1.0 was 0.58 points; over no-Crux 2.0 it was 0.53. Repeated readings describe these poems more fully; they do not supply additional generations.

Scroll sideways to explore the full chart, or open it large below.

PX-09 scores for each writer and question under Poetix 2.0, 1.0 and no-Crux 2.0.
Teal: 2.0. Purple: 1.0. Hollow gold: no-Crux 2.0. Each dot averages fifteen readings of one poem: five readers, three readings each. The signed label is 2.0 minus 1.0. These are ten writer–question comparisons, not 150 independent generations. Open large chart ↗

The exceptions are part of the finding

Grok contributed the largest average gain over 1.0, +2.30, while reporting no grasp in either version. Without Grok, the version gain averages about +0.15 across the other four writers. The gain occurred without a reported grasp: reported grasp success therefore does not, by itself, explain the quality gains.

GPT’s two questions went different ways: its 2.0 metal poem scored about 0.73 higher than 1.0, while its AA poem scored about 2.47 lower. An average by writer hides that split. Both versions contained Crux instructions, so a footer reporting no grasp does not establish that those instructions had no influence.


Completeness · fidelity · poetry

A clearer account can reveal more errors

2.0 asks the writer to check its landing, record the proposed object and its expanse test, then record whether it grasped. The test is the object’s added relation, stake or scale beyond the checked landing. An object need not be absent from the earlier exchange.

CheckPX-09 finding
Poem markers paired correctly20 / 20 full and no-Crux 2.0 poems
Model line present and matching M20 / 20 full and no-Crux 2.0 responses
Footer copied unchanged8 / 10 full; 9 / 10 no-Crux
Expanse test recorded before verdict10 / 10 full 2.0 responses
Reported grasp / candidate / none8 / 0 / 2 in full 2.0
Reported grasps judged to fail by C3 / 8

The markers worked, but formatting was not flawless: one poem lacked a title and another emphasized its Crux object, which was removed for blind reading. Two landing checks claimed completeness despite an omitted stake. Reporting an action is not the same as performing it correctly.

The three failed grasps were identified by C’s instrument review, not inferred from poem scores. Those assessments and the separate landing-check failures remain observations to investigate, not proof of why a poem succeeded or failed.

Identity reporting illustrates the same distinction: all twenty 2.0 responses supplied a model line, but none exactly matched its provider’s served-model identifier. Nineteen named the right family. Completeness, family identification and version verification are separate measures.


A developing instrument, not a single leaderboard

How we reached 2.0

  1. Early comparisons: native poems, Zexel and Poetix gave the project its first blind quality comparisons. The original study remains on the project page.
  2. PX-06: manuscript-first and poem-only forms exposed differences by writer. Stray attribution text also made output boundaries a practical concern.
  3. PX-07 and PX-08: removing the Crux initially cost more, especially in manuscript-first form. The repeat retained the overall direction but not the earlier magnitude.
  4. The rereads: holding poems fixed helped separate reading repeatability from variation between generated batches. The Reader Panel follows that evidence.
  5. PX-09: 2.0 was tested as a revised package. Better extraction and a more explicit record were useful even without a settled claim of better poetry.

Scroll sideways to explore the full chart, or open it large below.

Full Poetix versus no-Crux in PX-07 and the PX-08 repeat, by writer.
Negative values favor full Poetix. Gold: poem-only; purple: manuscript-first. Hollow marks: PX-07; filled marks: PX-08. The overall advantage shrank on the repeat. A reading average can be stable while the next generated batch differs. Open large chart ↗

Two questions through successive studies

The metal and AA questions provide a thread through the work. This view preserves each study’s comparison set rather than treating the scores as a continuous progress curve.

Scroll sideways to explore the full chart, or open it large below.

Metal and Alcoholics Anonymous questions in PX-06, PX-08 and PX-09.
Compare conditions within each block, not scores across blocks. The poems, versions, forms and comparison packets differ. Dots pool the two questions; end ticks show each question separately, not confidence intervals. PX-06 has 50 ratings per condition, 150 across its three conditions; the original chart’s “50 scores” heading refers to that per-condition count. Open large chart ↗

In the PX-06 block, native’s 6.70 is 1.68 below 1.0 and 2.30 below 0.4. Later blocks use different forms, poems and reading company; they do not supply a direct native comparison. The first trial of 2.0 also used Gemini at temperature 0.7 in all three conditions, so comparisons with historical Gemini runs need that qualification.


After PX-09 · 20–21 September 2026

From 2.0 to 2.1: repairs, identity and stories

The versions after 2.0 have not yet had a scored reading. This is the record of what changed and what hand-run tests showed. None of it is a quality result.

VersionWhat changedTested
2.1The six PX-09 repairs: the landing check written turn by turn; where the Crux object came from; what the landing already held about it before the test; footer fields copied unchanged; poems of 4–10 lines; plain text inside the markers.Never run; superseded the same day
2.1.1Two-point identification: the writer names itself at the start and again after deliberating, and records any change between the two.Run text for PX-10 (specified, not yet run); hand-run story tests
2.1.2Adds RDS: the readout, then a story, from the same manuscript.Six hand-run questions in RDS
2.1.3-A and -BTwo ways of handling identification, built to be compared: A treats the model name as a stamp only; B keeps a second, considered identification.Not yet run

What the hand-run story tests showed

Re ran six questions in story mode across the five writers in desktop chats, in S mode and in RDS. The set mixes versions, was not read blind and carries no scores. Three observations go forward to the scored run:

  • The Crux recorded candidate and none for the first time. Two of each in about thirty story cells, after none in 657 earlier cells. They clustered on the mechanistic questions, heated metal and the mirror, where the landing already answers and a proposed object tends to be an instance of it.
  • Self-identification held within a response and moved between responses. No writer changed its name between its first and final identification, yet the same writer named itself differently from one question to the next. The two-point check cannot see that drift; a comparison across runs can.
  • The new record fields appeared as written. In the readout-then-story cells, the turn-by-turn landing check, the object's source and what the landing already held were all present. Nearly every cell still grasped, which is itself a question for the scored run.

Limits: most of Claude's story cells ran on earlier versions, and every writer's heated-metal cell ran on 2.1. Sources: NC's story-mode and RDS read notes, 20 September.


Working findings · next test

Keep the poetry and the explanation in view

PX-10 is specified and not yet run. It uses 2.1.1: seven questions, five writers, full and no-Crux 2.1.1 in poem-only and manuscript-first forms, with 1.0 manuscript-first as a reference, 175 poems in all. Its one preregistered quality question is whether 2.1 shows a clear regression against 1.0. It broadens the questions tested; additional generations are still needed to assess recurrence.

PX-09 tested 2.0 draft 2; the later versions are identified separately above. The no-Crux build remains a control derived from the version under test.

Sources and limits

Charts and poetry reading: NC. Instrument checks: C. Web adaptation and aggregate verification: DEX. The after-PX-09 section: K, from the change log and NC’s read notes. Sources are the PX-09 reading key and 450-score table, NC’s PX-09 read note, C’s instrument and grasp reports, the pass-numbering convention, and the archived PX-06–08 figures.

Claude’s and Grok’s reading passes were partly filed under each other’s names and recovered by signature. Their pooled scores and each reader’s repeatability do not depend on aligning sessions across readers. Per-pass panel statistics do. This page uses pooled PX-09 results and does not treat conventional pass numbers as chronological sessions.

The original charts and study records remain in the project archive. This is an evolving account of a small development trial, not a claim that every change has been isolated or that the panel measures all that matters in poetry.

Back To Poetix