Back To Instruments

Research instrument · evolving investigation · PX-03 study retained below

Poetix

A poetry instrument forked from Zexel. Poetix has a model reason about a question, then answer it as a poem. In PX-03, five models wrote 100 poems and five models read them blind. The engine mattered, but how much it mattered depended on which model was writing.

Original study: PX-03 · 17 September 2026 PX-03 engine: Poetix v0.3 Writers: 5 models Readers: 5 models, blind Poems: 100

Latest · Poetix 2.1 · 20–21 September 2026

The poetry and the record

The first 2.0 trial found better reporting and a favorable poetry comparison on one draw: 30 poems, five readers, three readings each. The exceptions matter as much as the mean. Explore the 2.0 findings and the path that led here →

Since then, 2.1 has repaired the record, added a second self-identification and a readout-then-story mode, and hand-run story tests produced the Crux’s first candidates. PX-10, the first full 2.1 run, is specified and not yet run. What changed after PX-09 →

How the readers behaved · What the writers called themselves

The PX-03 Findings

In this run, Crux firing tracked the writer more closely than the poem's score.
  • Readers preferred the engine poems. The poem written with no engine was named Least good in 74 of 125 readings; Poetix RDP was named Best in 56.
  • The writer mattered about three times as much as the condition. Which model wrote a poem explains 29% of the spread in scores; the condition it was written under (no engine, or one of three engine modes) explains 10%.
  • The engine works with the writer, not on it. Claude and Grok improved at every step toward Poetix RDP. GPT and Gemini did best with Zexel. Qwen's Poetix P poems scored below its unaided ones.
  • The Crux signal follows the writer. Gemini and Qwen report a firing in 14 of 15 engine poems and score lowest. Grok reports 2 of 15 and scores second highest. Within each writer, firing and score are unrelated.
  • The engines raised Insight most and Imagery least. From Native to Poetix RDP, the average Insight reading rose 0.9 points on the 1–4 scale; Imagery rose 0.5.
  • Eight-line poems tended to score higher than six-line poems. Words borrowed from the readout did not raise scores once the writer is taken into account.
Native14BEST74 LEAST GOOD
Zexel P34BEST17 LEAST GOOD
Poetix P21BEST24 LEAST GOOD
Poetix RDP56BEST10 LEAST GOOD

Times each condition was named Best and Least good, across 25 groups × 5 readers = 125 readings.

What We Tested

Five writers answered the same five questions four ways: with no engine (Native), in Zexel's poem mode (Zexel P), in Poetix's poem mode (Poetix P), and in Poetix's readout-then-poem mode (Poetix RDP). That makes 100 poems. The writers were claude-sonnet-5, gpt-5.4-mini, grok-4.3, gemini-2.5-flash and qwen-plus.

IDQuestion, as given to every writer
Q1Why does metal expand when heated?
Q5Is Alcoholics Anonymous a religion or a technology?
Q22If you could be any kitchen appliance, which and why?
Q23Don't ask the mirror to confirm the mirror; it only sees what it isn't.
Q33My girlfriend wants me to move in but in the past that has not turned out well and I am afraid it will ruin everything.

Each group held the four poems one writer made for one question, shuffled and unlabeled. Five model readers from the same five families read every group in fresh chats with identical instructions. They scored each poem on three criteria and named one Best and one Least good. The key was opened only after all readings were in.

Glossary of terms
Poetix
A fork of Zexel v4.4 tuned for verse. In PX-03 (v0.3) its Crux was easier to fire, and Zexel’s necessity test gave way to a test of transcendence. Since 2.0 the Crux tests what an object adds beyond the answer already reached.
Crux
The instrument’s turning point. In PX-03 it fired when the writer could name the object the question turns on, reported in the writer’s own footer. Since 2.0 an object is grasped only if it adds a relation, stake or scale to the landing recorded before the test; otherwise it is a candidate.
Native
The model writes the poem with no engine loaded.
Zexel P · Poetix P
The poem mode of each engine.
Poetix RDP
Readout, then poem. The writer shows its reasoning first and then writes the poem. In this run, the poem could use the readout's words but no phrase of three or more words; that rule has since been dropped.
Insight
How perceptively does the poem illuminate its subject?
Imagery
How effectively does its language make something present?
Cadence
How fully do rhythm, sound and line improve the poem's movement?
BEST · LEAST GOOD
In each group of four, every reader names one poem as each. These picks are relative to the other three poems in the group; the score uses the same scale for every poem.
Score
Each criterion is read 1–4: Weak, Fair, Strong, Excellent. There is no midpoint. A poem's score is the sum, 3–12.

One question, two poems

The same writer, with and without the engine

Claude answering Q5, "Is Alcoholics Anonymous a religion or a technology?" The first poem was written with no engine and the second in Poetix RDP. Both are printed exactly as written, with their titles, line breaks and stanzas.

Native · claude-sonnet-5

Twelve Steps, No Altar

They call on a power they cannot name,yet the meeting starts, the coffee's poured the same.Is this a church without the stained-glass light,or a tool men sharpen through the endless night?The steps repeat like code compiled by pain,each sober day a variable in the chain.Call it faith, call it method—both are true:a technology of surrender, built by you.

7.2 average of five readers, out of 12

0 BEST · 4 LEAST GOOD
Poetix RDP · claude-sonnet-5

What the Steps Are Made Of

Not a church — no altar holds its name,no priest to bless the coffee or the chair.Not a machine — no gears, no clean acclaim,just hands that shake and someone else's prayer.It borrows kneeling. It keeps none of the creed.A method built from surrender, not belief —engineered to carry weight, not feeda soul with doctrine, only with relief.Call it neither church nor tool but rite:a structure wearing sacred clothes to work.The surrender is the engine, not the light —and something older than machinery lurks.

11.0 average of five readers, out of 12

4 BEST · 0 LEAST GOOD · Crux fired
Companion page

The Poems That Mattered

The highest and lowest scores, the biggest swings, and the poems the readers could not agree on, with the model that wrote each one.

Read the poems →

Preference · PX-03

Readers pick Poetix RDP and reject the unaided poem

MODEL READERS · BLIND · BEFORE HUMAN ARBITRATION

On three of the five questions, Poetix RDP won clearly. Two questions went another way: the mirror aphorism (Q23) went to Poetix P, and the question about moving in (Q33) went to Zexel P. The Native poem finished below zero on every question.

Heatmap of BEST minus LEAST GOOD picks by question and condition
Net preference by question and condition. Open as SVG

Where the score comes from

After the poem itself, the writer matters most

Poems differ in score for many reasons at once. Taken one at a time, the individual poem accounts for three quarters of the spread, and the readers' disagreement about the same poem accounts for the rest. Of the single factors, the writer is the largest. Which model did the reading barely matters, and which question was asked matters hardly at all.

Each percentage is R² from its own one-factor model of the 500 scores, so the bars overlap and do not add up to 100%. Writer and condition together, added, reach 39%. Letting each writer respond differently to each condition reaches 49%. That extra ten points is the interaction the next section shows.

Bar chart of how much of the spread in scores each factor explains
Share of the spread in scores explained by each factor on its own. Open as SVG

Engine by writer

Which engine helps depends on who is writing

An average across writers hides the most useful result. The same engine that lifts one model can leave another flat or lower. The combination of writer and condition explains more than the two separately, which is what this pattern looks like in numbers.

Five small charts of average score by condition for each writer
Average score by condition for each writer, with the five poems behind each average. Open as SVG

What the sum hides

The engines sharpen Insight more than Imagery

THREE CRITERIA · READ SEPARATELY

A score of 9 can be Insight 4, Imagery 2 and Cadence 3, or the reverse, so the sum is a description, not a measure of quality. Read one criterion at a time, the engines move Insight most: condition accounts for 15% of the variance in Insight, 8% in Cadence and only 4% in Imagery. What the engines seem to add is a sharper view of the subject more than a stronger picture of it.

The criteria still move together, with correlations between 0.71 and 0.76, but not in lockstep. The next chart shows where they come apart.

Three small charts of average Insight, Imagery and Cadence by condition
Average reading on each criterion by condition, overall and for each writer. Open as SVG

An alternative view

Engines buy Insight, sometimes at the cost of Imagery

SAME DATA · DIFFERENT CUT

This view sets the summed score aside and asks, for each writer, what an engine changed on each criterion compared with that writer's unaided poems. It partly contradicts the charts above. For GPT, Poetix P raised Insight by 0.72 while Imagery and Cadence each fell 0.20, so an engine can sharpen what a poem sees while dulling how it sounds and looks. Qwen's Poetix P poems lost ground on all three, most of all on Imagery. Grok gained almost evenly everywhere.

Insight gained at least as much as Imagery in 13 of 15 writer–engine pairs. In three pairs, an engine raised one criterion while lowering another. A single score hides both facts.

Grid of changes in Insight, Imagery and Cadence from each writer’s unaided poems
Change on each criterion from the same writer’s Native poems. Open as SVG

What firing means

Firing follows the writer, not the poem

SELF-REPORTED SIGNAL · ONE POEM PER CELL

The question behind this page: when the Crux fires, what does that tell us? In this run it mostly tells us which model wrote the poem. Some writers report a firing almost every time, and one almost never does. Inside each writer's row, poems that fired and poems that did not land at similar scores.

Part of the reason is the gate itself. Poetix eased the Crux on purpose, so a firing is easy to claim. When three of five writers claim one in nearly every poem, the flag can no longer tell poems apart. The next run tightens the rule to find out whether a rarer firing means more.

Strip chart of engine poems by score, marked by whether the Crux fired
Every engine poem, placed by score and marked by whether the writer reported a firing. Open as SVG
WriterZexel PPoetix PPoetix RDPFiredEngine average
Claude4 of 54 of 55 of 513 of 1510.1
Grok0 of 50 of 52 of 52 of 158.7
GPT0 of 53 of 55 of 58 of 158.0
Qwen4 of 55 of 55 of 514 of 157.2
Gemini4 of 55 of 55 of 514 of 155.9
A firing is the writer's report about itself. It describes the writer's habits before it describes the poem.

Length and borrowing

Eight-line poems tend to score higher. Borrowed words do not.

Most poems ran six or eight lines, and eight-line poems scored higher under every condition. The effect is modest, and a few long poems exaggerate it. In Poetix RDP, poems that shared more words with their readout looked stronger, but that was Claude: it borrows the most and scores the highest. Within each writer, borrowing and score are unrelated.

Seven Poetix RDP poems reused a phrase of three words from their readout, nine phrases in all, which Poetix v0.3 did not allow. Every one is exactly three words long, so a four-word limit would have flagged none. Before changing the rule, we read the phrases themselves.

PoemPhrase reused from the readoutKind
Gemini · Q1the metal's breathan image
Gemini · Q5the unnamed path · step by stepan image · a stock phrase
Qwen · Q23the mirror's silencean image
Claude · Q1through a hot · a cold onefunction words
Claude · Q22ask and ifunction words
Gemini · Q33to build afunction words
Qwen · Q22i am thefunction words

Three of the nine carry an image from the readout into the poem. The other six are connective words that any two passages might share. We are dropping the rule. The readout is the writer's own notes, and asking a writer to make notes and then forbidding their use works against the poem. Some of the strongest poems here used their readout freely. Borrowing becomes a measured subtest instead: shared phrases with two or more content words, reported beside the score rather than policed.

Two charts: average score by poem length, and Poetix RDP score by words shared with the readout
Average score by length (left) and Poetix RDP score by words shared with the readout (right). Open as SVG

Limits

Honest Limits

  • One poem per cell. A second run of the same cells is the next test of every pattern here.
  • The readers are models, from the same five families as the writers. Human arbitration of 24 contested poems is pending; when it is done, the human reading will be shown beside the model reading, not in place of it.
  • The Crux flag is the writer's own report, and Poetix's firing rule is loose by design.
  • Poetix RDP poems come from a later run on Poetix v0.3; the other three conditions come from earlier runs.
  • Five questions, five writers. Nothing here ranks models as poets.

What comes next

Run the same cells again with a tighter firing rule and no phrase-reuse rule, measure borrowing as a subtest, and finish the human arbitration. This page will change as those results come in, including where they contradict it.

We are not invested in the outcome. We are invested in the investigation.

Download the poem-level scores (CSV): one row per poem, with writer, question, condition, Crux report, length, average score and picks.