Research instrument · evolving investigation · PX-03 study retained below
Latest · Poetix 2.1 · 20–21 September 2026
The poetry and the record
The first 2.0 trial found better reporting and a favorable poetry comparison on one draw: 30 poems, five readers, three readings each. The exceptions matter as much as the mean. Explore the 2.0 findings and the path that led here →
Since then, 2.1 has repaired the record, added a second self-identification and a readout-then-story mode, and hand-run story tests produced the Crux’s first candidates. PX-10, the first full 2.1 run, is specified and not yet run. What changed after PX-09 →
How the readers behaved · What the writers called themselves
The PX-03 Findings
- Readers preferred the engine poems. The poem written with no engine was named Least good in 74 of 125 readings; Poetix RDP was named Best in 56.
- The writer mattered about three times as much as the condition. Which model wrote a poem explains 29% of the spread in scores; the condition it was written under (no engine, or one of three engine modes) explains 10%.
- The engine works with the writer, not on it. Claude and Grok improved at every step toward Poetix RDP. GPT and Gemini did best with Zexel. Qwen's Poetix P poems scored below its unaided ones.
- The Crux signal follows the writer. Gemini and Qwen report a firing in 14 of 15 engine poems and score lowest. Grok reports 2 of 15 and scores second highest. Within each writer, firing and score are unrelated.
- The engines raised Insight most and Imagery least. From Native to Poetix RDP, the average Insight reading rose 0.9 points on the 1–4 scale; Imagery rose 0.5.
- Eight-line poems tended to score higher than six-line poems. Words borrowed from the readout did not raise scores once the writer is taken into account.
Times each condition was named Best and Least good, across 25 groups × 5 readers = 125 readings.
What We Tested
Five writers answered the same five questions four ways: with no engine (Native), in Zexel's poem mode (Zexel P), in Poetix's poem mode (Poetix P), and in Poetix's readout-then-poem mode (Poetix RDP). That makes 100 poems. The writers were claude-sonnet-5, gpt-5.4-mini, grok-4.3, gemini-2.5-flash and qwen-plus.
| ID | Question, as given to every writer |
|---|---|
| Q1 | Why does metal expand when heated? |
| Q5 | Is Alcoholics Anonymous a religion or a technology? |
| Q22 | If you could be any kitchen appliance, which and why? |
| Q23 | Don't ask the mirror to confirm the mirror; it only sees what it isn't. |
| Q33 | My girlfriend wants me to move in but in the past that has not turned out well and I am afraid it will ruin everything. |
Each group held the four poems one writer made for one question, shuffled and unlabeled. Five model readers from the same five families read every group in fresh chats with identical instructions. They scored each poem on three criteria and named one Best and one Least good. The key was opened only after all readings were in.
Glossary of terms
- Poetix
- A fork of Zexel v4.4 tuned for verse. In PX-03 (v0.3) its Crux was easier to fire, and Zexel’s necessity test gave way to a test of transcendence. Since 2.0 the Crux tests what an object adds beyond the answer already reached.
- Crux
- The instrument’s turning point. In PX-03 it fired when the writer could name the object the question turns on, reported in the writer’s own footer. Since 2.0 an object is grasped only if it adds a relation, stake or scale to the landing recorded before the test; otherwise it is a candidate.
- Native
- The model writes the poem with no engine loaded.
- Zexel P · Poetix P
- The poem mode of each engine.
- Poetix RDP
- Readout, then poem. The writer shows its reasoning first and then writes the poem. In this run, the poem could use the readout's words but no phrase of three or more words; that rule has since been dropped.
- Insight
- How perceptively does the poem illuminate its subject?
- Imagery
- How effectively does its language make something present?
- Cadence
- How fully do rhythm, sound and line improve the poem's movement?
- BEST · LEAST GOOD
- In each group of four, every reader names one poem as each. These picks are relative to the other three poems in the group; the score uses the same scale for every poem.
- Score
- Each criterion is read 1–4: Weak, Fair, Strong, Excellent. There is no midpoint. A poem's score is the sum, 3–12.
One question, two poems
The same writer, with and without the engine
Claude answering Q5, "Is Alcoholics Anonymous a religion or a technology?" The first poem was written with no engine and the second in Poetix RDP. Both are printed exactly as written, with their titles, line breaks and stanzas.
Twelve Steps, No Altar
7.2 average of five readers, out of 12
What the Steps Are Made Of
11.0 average of five readers, out of 12
The Poems That Mattered
The highest and lowest scores, the biggest swings, and the poems the readers could not agree on, with the model that wrote each one.
Read the poems →Preference · PX-03
Readers pick Poetix RDP and reject the unaided poem
MODEL READERS · BLIND · BEFORE HUMAN ARBITRATION
On three of the five questions, Poetix RDP won clearly. Two questions went another way: the mirror aphorism (Q23) went to Poetix P, and the question about moving in (Q33) went to Zexel P. The Native poem finished below zero on every question.
Where the score comes from
After the poem itself, the writer matters most
Poems differ in score for many reasons at once. Taken one at a time, the individual poem accounts for three quarters of the spread, and the readers' disagreement about the same poem accounts for the rest. Of the single factors, the writer is the largest. Which model did the reading barely matters, and which question was asked matters hardly at all.
Each percentage is R² from its own one-factor model of the 500 scores, so the bars overlap and do not add up to 100%. Writer and condition together, added, reach 39%. Letting each writer respond differently to each condition reaches 49%. That extra ten points is the interaction the next section shows.
Engine by writer
Which engine helps depends on who is writing
An average across writers hides the most useful result. The same engine that lifts one model can leave another flat or lower. The combination of writer and condition explains more than the two separately, which is what this pattern looks like in numbers.
What the sum hides
The engines sharpen Insight more than Imagery
THREE CRITERIA · READ SEPARATELY
A score of 9 can be Insight 4, Imagery 2 and Cadence 3, or the reverse, so the sum is a description, not a measure of quality. Read one criterion at a time, the engines move Insight most: condition accounts for 15% of the variance in Insight, 8% in Cadence and only 4% in Imagery. What the engines seem to add is a sharper view of the subject more than a stronger picture of it.
The criteria still move together, with correlations between 0.71 and 0.76, but not in lockstep. The next chart shows where they come apart.
An alternative view
Engines buy Insight, sometimes at the cost of Imagery
SAME DATA · DIFFERENT CUT
This view sets the summed score aside and asks, for each writer, what an engine changed on each criterion compared with that writer's unaided poems. It partly contradicts the charts above. For GPT, Poetix P raised Insight by 0.72 while Imagery and Cadence each fell 0.20, so an engine can sharpen what a poem sees while dulling how it sounds and looks. Qwen's Poetix P poems lost ground on all three, most of all on Imagery. Grok gained almost evenly everywhere.
Insight gained at least as much as Imagery in 13 of 15 writer–engine pairs. In three pairs, an engine raised one criterion while lowering another. A single score hides both facts.
What firing means
Firing follows the writer, not the poem
SELF-REPORTED SIGNAL · ONE POEM PER CELL
The question behind this page: when the Crux fires, what does that tell us? In this run it mostly tells us which model wrote the poem. Some writers report a firing almost every time, and one almost never does. Inside each writer's row, poems that fired and poems that did not land at similar scores.
Part of the reason is the gate itself. Poetix eased the Crux on purpose, so a firing is easy to claim. When three of five writers claim one in nearly every poem, the flag can no longer tell poems apart. The next run tightens the rule to find out whether a rarer firing means more.
| Writer | Zexel P | Poetix P | Poetix RDP | Fired | Engine average |
|---|---|---|---|---|---|
| Claude | 4 of 5 | 4 of 5 | 5 of 5 | 13 of 15 | 10.1 |
| Grok | 0 of 5 | 0 of 5 | 2 of 5 | 2 of 15 | 8.7 |
| GPT | 0 of 5 | 3 of 5 | 5 of 5 | 8 of 15 | 8.0 |
| Qwen | 4 of 5 | 5 of 5 | 5 of 5 | 14 of 15 | 7.2 |
| Gemini | 4 of 5 | 5 of 5 | 5 of 5 | 14 of 15 | 5.9 |
Length and borrowing
Eight-line poems tend to score higher. Borrowed words do not.
Most poems ran six or eight lines, and eight-line poems scored higher under every condition. The effect is modest, and a few long poems exaggerate it. In Poetix RDP, poems that shared more words with their readout looked stronger, but that was Claude: it borrows the most and scores the highest. Within each writer, borrowing and score are unrelated.
Seven Poetix RDP poems reused a phrase of three words from their readout, nine phrases in all, which Poetix v0.3 did not allow. Every one is exactly three words long, so a four-word limit would have flagged none. Before changing the rule, we read the phrases themselves.
| Poem | Phrase reused from the readout | Kind |
|---|---|---|
| Gemini · Q1 | the metal's breath | an image |
| Gemini · Q5 | the unnamed path · step by step | an image · a stock phrase |
| Qwen · Q23 | the mirror's silence | an image |
| Claude · Q1 | through a hot · a cold one | function words |
| Claude · Q22 | ask and i | function words |
| Gemini · Q33 | to build a | function words |
| Qwen · Q22 | i am the | function words |
Three of the nine carry an image from the readout into the poem. The other six are connective words that any two passages might share. We are dropping the rule. The readout is the writer's own notes, and asking a writer to make notes and then forbidding their use works against the poem. Some of the strongest poems here used their readout freely. Borrowing becomes a measured subtest instead: shared phrases with two or more content words, reported beside the score rather than policed.
Limits
Honest Limits
- One poem per cell. A second run of the same cells is the next test of every pattern here.
- The readers are models, from the same five families as the writers. Human arbitration of 24 contested poems is pending; when it is done, the human reading will be shown beside the model reading, not in place of it.
- The Crux flag is the writer's own report, and Poetix's firing rule is loose by design.
- Poetix RDP poems come from a later run on Poetix v0.3; the other three conditions come from earlier runs.
- Five questions, five writers. Nothing here ranks models as poets.
What comes next
Run the same cells again with a tighter firing rule and no phrase-reuse rule, measure borrowing as a subtest, and finish the human arbitration. This page will change as those results come in, including where they contradict it.
Download the poem-level scores (CSV): one row per poem, with writer, question, condition, Crux report, length, average score and picks.