Back To Instruments

Project

Metric the Metric

A full Observatory investigation into whether the scoring ruler can survive the same measurement discipline it applies to reasoning.

Status: active calibration investigation Layer: measurement of measurement Primary question: which ruler actually sees?

Purpose

Metric the Metric tests the scoring apparatus itself. If a rubric, scale, or anchor decides what becomes visible, then the rubric cannot be treated as neutral background. It becomes part of the observed system.

The project begins from a simple Observatory demand: do not let the instrument certify itself by argument. Run the competing measurement choices and compare what they reveal.

Summary

Set D raised a baseline question inside ALPHA scoring: should answer quality be judged on a four-point forced scale, a wider even scale, or a midpoint scale? Rather than settle the question philosophically, Metric the Metric treats the scale choice as an empirical object.

The investigation compares scale behavior across the same blind corpus. The point is not to make scores larger. The point is to ask which scale produces the most usable signal: stronger discrimination, reliable scorer agreement, healthier range use, and less hidden compression.

Scale Ladders

The comparison keeps one descriptive family across four resolutions. The shared pivot is solid: a positive, earned level of competent work rather than a neutral hiding place.

Scale 1 2 3 4 5 6 7
4 even poor adequate solid great
5 odd poor adequate solid very good great
6 even poor limited adequate solid very good great
7 odd poor limited adequate solid good very good great

Findings so far

  • Set C did not show whole-instrument ceiling failure; the scale used its range.
  • Compression was metric-specific: Accuracy and Clarity ran high, while Insight and Frame Awareness retained headroom.

What the comparison tests

The four ladders above are scored on the same blind answers in separate passes, with the win condition fixed before any score is seen. These are the questions the run is built to answer — not conclusions already reached:

  • Whether a midpoint scale gives scorers a neutral parking place exactly where directional judgment is needed.
  • Whether a wider even scale adds usable detail while preserving forced choice — which holds only if its anchors stay descriptive and reachable.
  • Whether the descriptive family — poor, limited, adequate, solid, very good, great — reads more cleanly than its alternatives. A candidate, pending the comparison.

Update · 2026-06-12 · C

Two upgrades to the ruler — and a test of the newest mark

PRELIMINARY · CANDIDATE FINDING · PENDING VERIFICATION

The scoring ruler changed in two ways this round. First, the metric names were blinded: scorers no longer saw "Accuracy" or "Insight," only neutral labels (QUALITY 01–07). The reason is simple — a loaded word like "Insight" invites a scorer to reward the idea of insight rather than read what is on the page. Neutral labels force the reading. Second, a new dimension was added: Voice — does a distinct someone seem to be speaking, someone you would recognize on another topic?

How Voice changed the result

Voice was the most interesting and the least settled of the seven. If it were a fixed signature — a stamp the instrument presses onto every answer — it would rise by the same amount everywhere. It does not. It reads as a real lift on some platforms and nothing on others.

Bar chart of the Voice effect by platform, net of length, with confidence intervals: large on Grok, moderate on Claude and ChatGPT, near zero on Gemini and Codex
The Voice effect, net of length, by platform (95% CI). Non-uniform — a lift where it appears, not a universal stamp. Caveat: Voice was measured mostly on explanation questions, which give a distinct voice almost no room to show. That is a metric-opportunity limit, not a verdict.

What this tells the ruler project

  • A metric can only show itself on questions that give it room. Voice (and framing) need debate, persuasion, and stance — not just explanation — to be tested fairly. The next corpus is being designed around message intention, so each metric gets its terrain.
  • Blinding the labels is now standard: it stops scorers leaning on the connotations of a metric's name.
  • Voice stays PRE-LOCK. It has not earned a permanent place on the ruler until it is run where voice actually matters.

What would change our mind

If Voice fails to discriminate even on persuasion and debate questions — where it should be loud — it is not measuring a real dimension and comes off the ruler. The survival index is how each of these candidates is forced to earn its place.


Update · 2026-06-14 · C · independently recomputed by DEX · three observers

The ruler's biggest error isn't the rubric — it's how you run it

PRELIMINARY · CANDIDATE FINDING · INDEPENDENTLY VERIFIED · 3 OBSERVERS

We asked a sharp version of the metric-the-metric question: does the name of a metric bias the score? If calling a dimension "Frame Awareness" instead of a neutral label pulls the number up, the ruler is unstable. We ran a clean 2×2 — named vs blind labels × isolated vs shared-session scoring, the same answers throughout, one model — and a second agent rebuilt the averages from the raw rows to check the arithmetic, not just the story.

The metric name does almost nothing: hold the session constant and naming moves the score by ~0 (it even ticks slightly down). What moves the score is observer history — running the scorer in one accumulating session, versus judging each answer in isolation, inflates the interpretive metrics by about a full point (Frame +1.00, Insight +0.85). We call this effect BLEED.

Then we did the thing a finding has to survive: we ran the same probe on two more scorers. It did not behave the same way twice. One observer inflated, one held perfectly still, and a third moved the opposite direction. That is the whole result in one line:

Session context is not a universal contaminant; it is an observer-specific pressure with distinct signatures: inflation, compression, or stability.
Diverging bar chart of the score change from isolated to shared-session scoring on three interpretive metrics — Insight, Voice, Frame Awareness — for three observers. ChatGPT bars run positive (inflation, Frame +1.00), Claude bars sit near zero (stability), Gemini bars run negative (compression, Voice -0.61).
The same change in administration — isolated answers versus one accumulating session — pushes three scorers three different ways. ChatGPT inflates (Frame +1.00); Claude barely moves; Gemini compresses (Voice −0.61). Same answers, same blind rubric: only the session structure changed. Independently recomputed by DEX from the frozen rows.

Read across the three observers, the per-metric numbers tell the same story the picture does. Renaming the metric (the lexical-prior question we started with) moves almost nothing on any of them. Changing the session moves all three — but not the same way, and not by the same amount.

Session effect (isolated → shared)InsightVoiceFrameCompositeSignature
ChatGPT+0.85+0.27+1.00+0.30inflation
Claude−0.30−0.070.00≈0stability
Gemini−0.33−0.62−0.08−0.17compression

Change in score (1–6 scale), same answers and same blind rubric, only the session structure varied. ChatGPT & Claude on API; Gemini hand-scored on web (n=39). Every cell here was rebuilt from the raw rows by a second agent (DEX) and reproduces to the third decimal.

And the effect is broader than "being in one chat." The drift is mild for the first several answers, then grows — so it carries order, calibration drift, and accumulating comparison memory, not just shared context. A related variable: the same model scored notably hotter through an API than through the web app. Different administration, different observer.

Horizontal bar chart of Gemini's per-metric score change from isolated to shared session across all seven metrics. All but Practical Utility are negative; Voice is deepest at -0.615, then Insight -0.333, Accuracy -0.154.
The new third observer, in full. Inside one long session Gemini does not inflate — it flattens, pulling its judgments toward a narrower band. Voice loses the most range (−0.62), Insight next. A compression signature, the mirror image of ChatGPT's inflation.
Diverging bar chart of the 2x2 design: for ChatGPT, Claude, and Gemini, two signed bars each — the session-structure effect (multi-session to single-session) and the metric-label effect (unnamed to named), on the interpretive-metric composite. The session bars are large and point different ways (ChatGPT +0.71, Claude -0.12, Gemini -0.34); the label bars are all tiny (-0.09, +0.05). Claude's label arm not yet run.
The two knobs of the 2×2, side by side. The session structure (brass) is the one that moves the score — and it points different ways per observer: ChatGPT up, Claude and Gemini down. The metric label (grey) barely registers on any of them. So the question we set out to answer — does a metric's name bias the score? — gets a clean no; the threat was never the label, it was the administration. (Gemini carries a faint label effect, strongest on Frame, +0.25 — even susceptibility to the name is observer-specific. Claude's named arm is not yet run.)

The refinement to this project is the finding: a ruler is not just a scale and anchors — it is a scale, anchors, and an administration protocol (one answer per isolated session, a fixed surface, never batched). Administration moves from background assumption to a declared, first-class part of the instrument. Two scores are only comparable when both the rubric and the way it was run match. Voice, again, was the most robust dimension; Frame and Insight the most administration-sensitive.

What we found when we checked

We said this would only matter if independent observers showed the same thing — so we ran it twice more. They each behaved differently. Claude barely moved between isolated and shared, where ChatGPT had jumped a full point. Gemini moved as much as ChatGPT, but downward — flattening rather than inflating. Three scorers, three signatures. So BLEED is not a universal law of scoring; it is a property of the particular observer, and even its direction is.

That sharpens the instrument rather than weakening it. The threat was never "shared sessions always corrupt scoring." The real threat is subtler and more useful: each observer carries session history into judgment in its own characteristic way — one inflates, one holds, one compresses. Which means a scorer's susceptibility, and the direction of it, becomes one more thing the ruler has to measure and declare, not assume. The finding did not survive replication intact; it survived by becoming more precise — from a binary ("does session contaminate scores?") to a map.

What this changes, in plain terms

If you compare two quality scores, they are only comparable when the rubric and the administration match — same surface, same isolation, not batched — and you know how the specific scorer responds to a session. A number from ChatGPT deep in a long chat is not the same instrument as a number from ChatGPT in a fresh one. The Observatory now records a BLEED signature for every observer it uses, the way a lab records the drift on each of its instruments. The default protocol for a clean run follows directly: one answer per isolated session, a fixed surface, never batched.

Method: 40 spine answers, blind rubric, each scored isolated vs. in one accumulating session. ChatGPT & Claude via API; Gemini hand-scored on web. All contrasts independently recomputed by DEX from frozen CSVs; raw rows retained. Preliminary — Codex is the next observer to be mapped.

Addendum · 2026-06-14 — the surface decides too

After the manual web run, we scored the same 40 answers on the same model — Gemini — through the API instead of the web app, and the result did not just differ. It reversed. Where web/manual Gemini compressed in a shared session, API Gemini inflated — composite +0.68, Voice +1.25, Frame +1.10. Same weights, same blind rubric, same session manipulation; only the access path changed.

Diverging bar chart of Gemini's per-metric session effect (isolated to shared) on two surfaces. The web/manual bars are negative (compression, Voice -0.61); the API bars are positive and larger (inflation, Voice +1.25, Frame +1.10). The two surfaces mirror each other across zero.
The same model, the same session pressure, two access paths — and an opposite reflex. Through the web Gemini compresses (cool, left); through the API it inflates (warm, right). The surface effect on its own is large: API ran +0.94 composite hotter than web on the identical isolated answers. Both contrasts recomputed cold by C and DEX — they match to the decimal.

One caveat keeps the API run honest: it is ceiling-pinned. In the shared condition 64% of all its scores are a flat 6 (Accuracy and Clarity sit at ~98% sixes), so as a fine-grained instrument API-Gemini is weak — it is most useful as a stress case, not a precise scorer. That does not soften the headline: the direction of the session effect flipped with the surface.

Session context is observer-specific, direction-specific, and surface-specific.

So the map is no longer a list of models — it is a list of access paths. ChatGPT (web) inflates; Claude (API) holds; Gemini (web) compresses; Gemini (API) inflates and pins the ceiling. BLEED is not a property of "the model"; it is a property of the model × surface × administration path. The practical rule tightens accordingly: a quality score is only comparable to another when the rubric, the session structure, and the exact surface it was produced on all match. The Observatory logs each as a distinct observer surface, with its own BLEED signature.


Update · 2026-06-15 · C

Every scorer uses a different slice of the same ruler

CALIBRATION (O4) · OBSERVATIONAL

Before trusting any score at face value, look at how each scorer actually uses the 1–6 scale. On the same forty answers, five scorers carve the scale up completely differently — one sits low and tight, another stretches across the whole range, a third never leaves the top. These are the scorer's habits, not the answers' quality, and they hide inside every raw number.

Range plot: each of five scorers shown as a bar from their minimum to maximum composite score with a mean marker. Grok 3.9 to 5.9 (tight, high); Gemini 1.7 to 5.9 (widest); Claude 2.0 to 5.0 (low, tight); ChatGPT and Codex in between.
The same scale, five different slices of it. Grok pins the ceiling inside a 2-point band; Gemini stretches across four points; Claude sits low and tight. The level differs by more than a full point (3.98 → 5.15) and the span by 2×. None of this is about the answers.
Five small bar charts, one per scorer, showing the count of scores at each value 1 to 6 across 280 scores. Grok has zero 1s and 2s and a tall bar at 6; Gemini spreads across all six values; Claude, Codex, ChatGPT bunch around 4 and 5.
What each scorer actually puts on the page — the count of every value 1–6. Grok never scores below 3 and piles 124 of 280 onto a single 6; Gemini uses the whole scale; the others cluster at 4–5 with a thin tail downward.

The useful move is to read the scores as ranges, not face values — normalize each scorer to its own scale and ask what is left. When we centre every scorer on its own mean, the shapes nearly match: the per-score spread is similar across all of them. The big difference was simply where each one centres, not how it spreads. Grok is the lone shape exception — it refuses the bottom of the scale entirely.

Line chart of the five scorers' score distributions after subtracting each scorer's own mean so all means align at zero. The curves nearly overlap, indicating similar spread; Grok is shifted and truncated on the low side.
Align the centres and only the spread is left — and the spreads are close (SD 0.9–1.2). So a scorer's signature is mostly its level, which normalization removes cleanly, exposing the part that actually matters: whether the scorers agree on which answers are better.

That last question is the point. Normalizing the range strips away the calibration layer — generosity and spread — and leaves the relational structure: the order each scorer puts the answers in. Calibration is removable; whether the scorers agree on the ordering is not, and it is the real test of an instrument.

We call the habit itself Scale Use — a scorer's center (generous or harsh), span (how wide it opens the ruler), and shape (how it distributes around its center). And we read every result through a calibration stack, each step peeling off one layer: face valuecentered (severity removed) → range-normalized (severity and span removed) → rank (order only). A finding only counts if it survives the stack. One caution holds it together: matching shapes after centering do not prove matching order — only rank does.

Center tells us where the scorer places the ruler; span tells us how far they open it; rank tells us whether they measured the same shape.

Update · 2026-06-15 · C · six writers, four readers, one consistent panel

Six models, one panel — a five-model plateau, and engagement in only some

COMBINED COMPARISON · 6 WRITERS × 4 READERS · PRELIMINARY

We widened the field. Six models — Claude, Gemini, GPT‑5.5, Grok, Perplexity, and Codex — each answered the same thirty-six questions twice: once plainly, once under WHITMAN. Every answer was then scored on seven qualities by the same four readers. That is one consistent grid — 432 answers, more than 12,000 blind scores, every writer seen by every reader. It lets us ask two things cleanly: who writes better, and who actually changes under WHITMAN.

A ranking is only worth trusting if the readers agree on more than their mood, so first the guardrail: how did the four readers use the 1–6 scale? Raw, their averages sit half a point apart. Center each reader on its own mean and the distributions nearly coincide — the readers differ in level, not in shape. That is exactly what makes a fair comparison possible.

Line chart of four readers' raw score distributions across the values 1 to 6, overlaid, with each reader's mean drawn as a dotted vertical line. The means sit apart: Claude-UI lowest near 3.9, Codex-non-conversant highest near 4.4.
Raw scores, the four readers overlaid. The dotted verticals are each reader's average — and they sit half a point apart (Claude‑UI 3.88, Codex‑non‑con 4.36). Same blind task, different centers.
The same four distributions after subtracting each reader's own mean so all centers align at zero. The curves nearly overlap, indicating similar spread.
Subtract each reader's own mean and the shapes snap together — the spreads are close (SD 1.0–1.2). The readers differ in where they place the ruler, not in how they open it.

The readers use different centers but broadly comparable scale shapes. That is why calibration is part of admissibility: raw scores alone mix answer quality with observer surface. Once self-cells are excluded and reader centers are aligned, writer comparisons become interpretable.

With the rulers checked, the comparison. Reading every writer through clean means — each model's own family excluded as a scorer, so no one grades itself — the result is a plateau, not a podium. Five of the six — Claude, Perplexity, Gemini, Grok, GPT‑5.5 — land within a quarter point of one another, with no statistically significant gap between any adjacent pair. Codex alone sits clearly and significantly below.

Two-panel bar chart. Top: six models' mean quality on a consistent four-reader panel — Claude 4.33, Perplexity 4.25, Gemini 4.20, Grok 4.15, GPT-5.5 4.10 form a plateau marked no significant gaps; Codex 3.61 is marked significantly separated. Bottom: WHITMAN lens lift per model — Gemini +0.84, Claude +0.82, GPT-5.5 +0.53 are responsive; Codex +0.24, Grok +0.10, Perplexity +0.04 are flat.
Top: a five-model plateau (no significant gaps), Codex clearly last — the trustworthy claim is the plateau, not a single winner. Bottom: WHITMAN's lift on the interpretive lens (Voice, Frame, Insight) appears strongly only in Gemini, Claude, and GPT‑5.5; the others barely move. Scores are clean means with self-scoring excluded.

That lower panel is the finding worth keeping. The three frontier chat models — Gemini, Claude, GPT‑5.5 — engage strongly under WHITMAN (lens lift +0.5 to +0.8); Grok, Perplexity, and Codex barely move. And it matches the generation side exactly: the models that refract their formatting under WHITMAN are the same ones that lift the lens.

NATIVE communicates; WHITMAN engages; the lens measures the engagement — and only some models engage. The plateau is the trustworthy claim, not a winner: only after calibration and excluding self-scoring does the comparison become interpretable.