Back To Poetix

Poetix · the evaluating models · 20 September 2026

The Reader Panel

Before trusting an average, look at what it averages. Five models read the same poems. Their scores differ, their judgments often move together, and their average is more repeatable than an individual reading.

Panel: ChatGPT · Claude · Gemini · Grok · QwenEvidence: PX-07–09 & repeated readings

Five readers, one poem

Readers see the poems without the manuscripts or production records. Each scores insight, imagery and cadence from 1 to 4, for a total of 3–12, and names the strongest and weakest poem in its group. The reported panel score is the average of the five readers.

Here the models are evaluators. A generous score does not tell us how well that model writes poetry. These are tendencies observed in this panel and these reading sessions.


175 poems · 875 scores · PX-08

Every poem, every reader

The average gives each poem one number. The dots show what that number hides: five distinct readings, sometimes several points apart.

Scroll sideways to explore the full chart, or open it large below.

Five individual reader scores for each of 175 PX-08 poems, ordered by their panel mean.
Each column is one poem; each colored dot is one reader. The dark line is the five-reader mean. Small vertical offsets separate tied dots; scores are whole numbers from 3 to 12. The line rises because poems are sorted by their mean, not because this is a time series. Open large chart ↗

The highest and lowest scores for the same poem are 3.0 points apart on average in PX-08. For 15 of the 175 poems, the spread is five or six points. ChatGPT and Grok score higher on average; Claude, Gemini and Qwen lower. That pattern does not remove disagreements over individual poems.


Different score levels · related judgments

How the readers differ

ChatGPT has the highest mean and Grok the second-highest in all three sittings. The order of Claude, Gemini and Qwen changes. The panel shows recurring differences in scoring level, rather than five interchangeable judges.

Scroll sideways to explore the full chart, or open it large below.

Reader mean scores and pairwise Pearson score correlations in PX-07, PX-08, and the PX-08 reread.
Top: each reader’s mean score. Matrices: Pearson correlations between readers’ scores. PX-07 has 174 poems; PX-08 is a different batch of 175 poems, shared with its reread. A correlation of 1 means a perfect linear relationship, not necessarily identical scores. Open large chart ↗

Across the three sittings, pairwise score correlations range from 0.68 to 0.83. Readers tend to score the same poems higher or lower, but they neither give identical scores nor always choose the same favorite. All five chose the same strongest poem in 10 of 35 PX-07 groups, 11 of 35 PX-08 groups, and 9 of 35 reread groups.

Compare each reader’s mean score
ReaderPX-07 · 174 poemsPX-08 · 175 poemsReread · same 175
ChatGPT9.108.779.02
Claude7.537.027.22
Gemini6.957.106.95
Grok8.328.438.50
Qwen7.397.437.30

Same poems · fresh sessions

What holds when they read again?

PX-08’s 175 poems were read again by the same five-reader panel, using packets identical apart from the signature line. An individual reader’s score changed by 0.68 points on average; the five-reader mean changed by 0.38. The correlation between the two sets of panel means was 0.97.

Scroll sideways to explore the full chart, or open it large below.

Full versus no-Crux contrasts for the same poems in two reading sessions, grouped by writer.
Hollow marks are the first reading; filled marks are the reread. Gold is poem-only; purple is manuscript-first. Values are no-Crux minus full Poetix, so negative values favor full Poetix. The pooled contrasts move by less than 0.1 point; individual writer contrasts can move more. Open large chart ↗

The panel’s averages were repeatable in this reread. The picture is less uniform within each writer: Grok’s manuscript-first contrast changes from −0.11 to +0.51. A steady overall mean can coexist with a changing local judgment.

Compare each reader’s change on the same poems

Mean absolute change, in points, over 175 poems per reader. This measures reread consistency, not accuracy.

ReaderMean absolute score change
ChatGPT0.62
Claude0.43
Gemini0.68
Grok0.55
Qwen1.11

What the average can tell us

Averaging can reduce individual reading fluctuations; it does not automatically cancel systematic bias or establish an objective measure of poetic quality. A repeatable panel may still share preferences or blind spots.

For context, matched writer–question–condition cells changed by about 1.52 points on average between PX-07 and PX-08, when poems were generated and read again. The reread isolates a narrower change: the poems stay fixed. This supports investigating variation in generation, while leaving the precise division of variance conditional on the study’s assumptions.


PX-09 · a new packet · three readings

Each reader, three times over

The 2.0 smoke trial supplied another repeatability check: the same 30 poems read three times by each evaluator, in fresh sessions. This view follows each reader’s condition averages.

Scroll sideways to explore the full chart, or open it large below.

Each PX-09 evaluator’s condition averages over three readings of the same poems.
Teal: 2.0. Purple: 1.0. Hollow gold: no-Crux 2.0. Three marks show each reader’s three condition averages; labels are their pooled means. Overlapping marks may look like fewer than three. Each average covers ten poems; this chart does not show the full variation on individual poems. Open large chart ↗

Four readers put 2.0 above 1.0 on each of their readings. Gemini’s difference was small and changed sign. Claude had the smallest mean three-reading range on individual poems (0.53); Qwen the largest (1.57). These ranges are not the same statistic as the two-reading absolute changes reported above.

Two returns were filed under the wrong reader’s name and reassigned by signature. The three readings can be compared within a reader, but their chronological alignment across readers is not recoverable. Pooled means and individual-reader repeatability survive that ambiguity; per-pass panel means depend on a numbering convention.

See what these readers found in the 2.0 poems →


An evolving record

Sources & what comes next

Charts and initial web copy: NC. Web adaptation and numerical checks: DEX. This page begins with the available charts and reading records; it can grow with the fuller reader-panel report and subsequent runs.

The figures retain NC’s plotted data and colors. Titles and explanatory captions have been adapted for the web. The original charts, reader responses and score tables remain in the project archive.

  • PX-07: 870 ratings of 174 poems, five readers per poem.
  • PX-08: 875 ratings of 175 poems.
  • PX-08 reread: another 875 ratings of those same 175 poems.
  • Scope: these models, poems, instructions and reading sessions; the charts do not establish general evaluator accuracy.

Instrument and writer comparisons belong to the main Poetix investigation. Models’ claims about their identities are tracked separately in Model Identification.

Back To Poetix