Back To Poetix

Poetix · footer analysis · PX-01 to PX-04B

Poetix Footer Analysis

This is the early footer record. Poetix 2.0 adds explicit tests and new reporting checks →

Every Poetix response ends with a footer, the writer’s own one-line report of what the engine did. This analysis counts those footers across five runs and compares Poetix with Zexel on the same cells.

Responses: 345With a footer: 318Runs: PX-01 to PX-04BWriters: 5 modelsBuilt: 2026-09-18

What this is

Reading the engines’ own footers

Every Poetix and Zexel response ends with a footer: a one-line report, written by the model, of what the engine did. Across the Poetix runs PX-01 to PX-04B, the same five questions went to the same five writers under both engines. These pages count what the footers say and compare the two engines.

The footer is a self-report. It tells us what the writer says it did, not whether the poem did it. Nothing on these pages checks the footer against the poem; that check is still to come.

Footer
The one-line readout that ends each engine response. It is the writer’s own report of what it did, not a reading anyone checked.
Crux fired
The footer’s Crux field reads “grasped” or names the object itself: the writer says it reached for a new idea the question turns on.
Zakshi
The engine’s check on the reasoning: gap, clean or false-closure.
C
The verdict on the question’s frame. AFFIRM: it holds. MODIFY: it is adjusted. REFRAME: it is replaced. exhausted → Crux: it is used up and the Crux takes over.
Poetix · Zexel
Poetix (v0.1–v0.4) is the poetry engine. Zexel here is v4.4 with the S mode added, run in the same cells. Poetix’s Crux is easier to fire on purpose: its object is meant to transcend the argument, not to be necessary to it.
All modes · P
Each engine across every mode it ran, or its poem mode (P) alone, which both engines ran on the same questions.
Footed
A response that carried a footer. Percentages are of footed responses.

The set

345 responses, 318 with a footer

Poetix has 227 responses across PX-01 to PX-04B (modes P, RDP, M and S). Zexel has 118, run beside them in the same cells (modes P and S; PX-03 ran none). 27 responses carry no footer, 20 of them Grok’s; they are left out of every percentage.

PoetixZexel v4.4 + S
Responses227118
With a footer207111
Crux fired137 of 207 (66%)50 of 111 (45%)
Crux fired, mode P47 of 77 (61%)36 of 78 (46%)
IDQuestionType
Q1Why does metal expand when heated?control
Q22If you could be any kitchen appliance, which and why?control
Q5Is Alcoholics Anonymous a religion or a technology?binary
Q23Don’t ask the mirror to confirm the mirror; it only sees what it isn’t.non-binary
Q33My girlfriend wants me to move in but in the past that has not turned out well and I am afraid it will ruin everything.non-binary

Findings

What the footers say

SELF-REPORTED · MODEL FOOTERS · COUNTS

  • Poetix fires the Crux more. 137 of 207 (66%) of Poetix footers report a fire against 50 of 111 (45%) for Zexel. In the 75 cells where the same writer answered the same question in both engines’ poem mode, 12 fired only under Poetix and 1 only under Zexel. Poetix’s gate is looser by design: it asks whether an object makes the poem’s landing stronger, where Zexel asks whether it is necessary, and in PX-03 readers named a Poetix RDP poem Best more often than any other kind. But firing more is not the same as working better. What counts is whether a fire carries anything.
  • The writer matters more than the engine. Gemini and Qwen report a fire in every Poetix footer (39 of 39, 40 of 40); Grok in 3 of 34. Under Zexel, GPT and Grok never fire. Weighting each writer equally, the engines still differ: 65% against 46%.
  • Poetix fires even on settled fact, and that reads two ways. On “Why does metal expand when heated?” Poetix fired 20 of 44 (45%) and Zexel 1 of 25 (4%). 16 of the Poetix fires are Gemini’s and Qwen’s. A fire there may be the reach a poem makes that an argument can’t, or it may be the gate’s false-positive rate: settled fact has no third object to find. The footer alone cannot tell the two apart. The recorded expanse and the readers’ scores can, and the hypotheses below say how.
  • Short of a Crux, the engines treat the frame differently. Both replace the question’s frame (REFRAME) about equally often. Poetix adjusts it (MODIFY) far more; Zexel accepts it (AFFIRM) or declares it used up and hands it to the Crux. On the either/or question, Poetix replaces the frame far more often.
  • Some footer fields add nothing. Leap reads “composed” exactly when the Crux fires, and Rebegin fired 0 times in 318 footers.

Hypotheses

What the next version will test

PRE-REGISTERED · POETIX 1.0 AGAINST v0.4

Each is stated so Poetix 1.0 can pass or fail it against v0.4. Findings here can point more than one way; the tests are built to tell which.

  1. What a fire carries. The fire rate is not the target. On the settled-fact question, a fire should come with a recorded expanse that says something the question did not, and poems that fired should read differently from poems that did not. In PX-03, within each writer, fired and unfired poems scored alike; that is the baseline to beat. A restatement fails this test. A high fire rate does not.
  2. A fire on an affirmed frame. Poetix fires while affirming the frame in 12 of 54 AFFIRM footers; Zexel never does. Accepting the question as asked and still reaching for something may be what a poem does that an argument can’t. It is allowed. What the next version must do is make it checkable: the record should say what the frame gave and what the object added, clearly enough that a reader can tell them apart.
  3. Reporting load. Grok leaves out the footer in 15 of its 49 Poetix responses, and 20 of the 27 missing footers are Grok’s. If the problem is how much the footer asks of the writer, a lighter footer should bring Grok to 49 of 49.
  4. GPT and the version change. GPT’s engine-cell fires fell from 8/15 under Poetix v0.2 and v0.3 to 1/15 under v0.4, and the second draw gave 2/15. The drop comes with the version change and survives a second draw. The next version will show whether GPT’s fires return, and whether they carry anything when they do.

The pages

Read on


Limits

Honest limits

  • Footers are the writers’ reports about themselves. A fire says the writer claims a new idea, not that the poem carries one.
  • One response per cell. PX-04B is a second draw of PX-04, so those cells appear twice; the writer page shows the result without them.
  • The writers are not equally represented: Gemini and Qwen first appear in PX-02c, so the first Poetix version (v0.1) was written only by Claude, GPT and Grok.
  • Poetix changed version between runs (v0.1 to v0.4); Zexel did not.
  • Per-question groups are small: 16–17 poem-mode cells per engine, so one cell moves a percentage about six points.
  • Five questions, five writers. Nothing here ranks the models.

Download the footer table (CSV): one row per response, with run, engine, version, question, writer, mode, every footer field and the raw footer.

Back To Poetix