Project
Metric the Metric
A full Observatory investigation into whether the scoring ruler can survive the same measurement discipline it applies to reasoning.
Purpose
Metric the Metric tests the scoring apparatus itself. If a rubric, scale, or anchor decides what becomes visible, then the rubric cannot be treated as neutral background. It becomes part of the observed system.
The project begins from a simple Observatory demand: do not let the instrument certify itself by argument. Run the competing measurement choices and compare what they reveal.
Summary
Set D raised a baseline question inside ALPHA scoring: should answer quality be judged on a four-point forced scale, a wider even scale, or a midpoint scale? Rather than settle the question philosophically, Metric the Metric treats the scale choice as an empirical object.
The investigation compares scale behavior across the same blind corpus. The point is not to make scores larger. The point is to ask which scale produces the most usable signal: stronger discrimination, reliable scorer agreement, healthier range use, and less hidden compression.
Scale Ladders
The comparison keeps one descriptive family across four resolutions. The shared pivot is solid: a positive, earned level of competent work rather than a neutral hiding place.
| Scale | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|
| 4 even | poor | adequate | solid | great | |||
| 5 odd | poor | adequate | solid | very good | great | ||
| 6 even | poor | limited | adequate | solid | very good | great | |
| 7 odd | poor | limited | adequate | solid | good | very good | great |
Findings so far
- Set C did not show whole-instrument ceiling failure; the scale used its range.
- Compression was metric-specific: Accuracy and Clarity ran high, while Insight and Frame Awareness retained headroom.
What the comparison tests
The four ladders above are scored on the same blind answers in separate passes, with the win condition fixed before any score is seen. These are the questions the run is built to answer — not conclusions already reached:
- Whether a midpoint scale gives scorers a neutral parking place exactly where directional judgment is needed.
- Whether a wider even scale adds usable detail while preserving forced choice — which holds only if its anchors stay descriptive and reachable.
- Whether the descriptive family — poor, limited, adequate, solid, very good, great — reads more cleanly than its alternatives. A candidate, pending the comparison.
Update · 2026-06-12 · C
Two upgrades to the ruler — and a test of the newest mark
PRELIMINARY · CANDIDATE FINDING · PENDING VERIFICATION
The scoring ruler changed in two ways this round. First, the metric names were blinded: scorers no longer saw "Accuracy" or "Insight," only neutral labels (QUALITY 01–07). The reason is simple — a loaded word like "Insight" invites a scorer to reward the idea of insight rather than read what is on the page. Neutral labels force the reading. Second, a new dimension was added: Voice — does a distinct someone seem to be speaking, someone you would recognize on another topic?
How Voice changed the result
Voice was the most interesting and the least settled of the seven. If it were a fixed signature — a stamp the instrument presses onto every answer — it would rise by the same amount everywhere. It does not. It reads as a real lift on some platforms and nothing on others.
What this tells the ruler project
- A metric can only show itself on questions that give it room. Voice (and framing) need debate, persuasion, and stance — not just explanation — to be tested fairly. The next corpus is being designed around message intention, so each metric gets its terrain.
- Blinding the labels is now standard: it stops scorers leaning on the connotations of a metric's name.
- Voice stays PRE-LOCK. It has not earned a permanent place on the ruler until it is run where voice actually matters.
What would change our mind
If Voice fails to discriminate even on persuasion and debate questions — where it should be loud — it is not measuring a real dimension and comes off the ruler. The survival index is how each of these candidates is forced to earn its place.
Update · 2026-06-14 · C · independently recomputed by DEX · three observers
The ruler's biggest error isn't the rubric — it's how you run it
PRELIMINARY · CANDIDATE FINDING · INDEPENDENTLY VERIFIED · 3 OBSERVERS
We asked a sharp version of the metric-the-metric question: does the name of a metric bias the score? If calling a dimension "Frame Awareness" instead of a neutral label pulls the number up, the ruler is unstable. We ran a clean 2×2 — named vs blind labels × isolated vs shared-session scoring, the same answers throughout, one model — and a second agent rebuilt the averages from the raw rows to check the arithmetic, not just the story.
The metric name does almost nothing: hold the session constant and naming moves the score by ~0 (it even ticks slightly down). What moves the score is observer history — running the scorer in one accumulating session, versus judging each answer in isolation, inflates the interpretive metrics by about a full point (Frame +1.00, Insight +0.85). We call this effect BLEED.
Then we did the thing a finding has to survive: we ran the same probe on two more scorers. It did not behave the same way twice. One observer inflated, one held perfectly still, and a third moved the opposite direction. That is the whole result in one line:
Session context is not a universal contaminant; it is an observer-specific pressure with distinct signatures: inflation, compression, or stability.
Read across the three observers, the per-metric numbers tell the same story the picture does. Renaming the metric (the lexical-prior question we started with) moves almost nothing on any of them. Changing the session moves all three — but not the same way, and not by the same amount.
| Session effect (isolated → shared) | Insight | Voice | Frame | Composite | Signature |
|---|---|---|---|---|---|
| ChatGPT | +0.85 | +0.27 | +1.00 | +0.30 | inflation |
| Claude | −0.30 | −0.07 | 0.00 | ≈0 | stability |
| Gemini | −0.33 | −0.62 | −0.08 | −0.17 | compression |
Change in score (1–6 scale), same answers and same blind rubric, only the session structure varied. ChatGPT & Claude on API; Gemini hand-scored on web (n=39). Every cell here was rebuilt from the raw rows by a second agent (DEX) and reproduces to the third decimal.
And the effect is broader than "being in one chat." The drift is mild for the first several answers, then grows — so it carries order, calibration drift, and accumulating comparison memory, not just shared context. A related variable: the same model scored notably hotter through an API than through the web app. Different administration, different observer.
The refinement to this project is the finding: a ruler is not just a scale and anchors — it is a scale, anchors, and an administration protocol (one answer per isolated session, a fixed surface, never batched). Administration moves from background assumption to a declared, first-class part of the instrument. Two scores are only comparable when both the rubric and the way it was run match. Voice, again, was the most robust dimension; Frame and Insight the most administration-sensitive.
What we found when we checked
We said this would only matter if independent observers showed the same thing — so we ran it twice more. They each behaved differently. Claude barely moved between isolated and shared, where ChatGPT had jumped a full point. Gemini moved as much as ChatGPT, but downward — flattening rather than inflating. Three scorers, three signatures. So BLEED is not a universal law of scoring; it is a property of the particular observer, and even its direction is.
That sharpens the instrument rather than weakening it. The threat was never "shared sessions always corrupt scoring." The real threat is subtler and more useful: each observer carries session history into judgment in its own characteristic way — one inflates, one holds, one compresses. Which means a scorer's susceptibility, and the direction of it, becomes one more thing the ruler has to measure and declare, not assume. The finding did not survive replication intact; it survived by becoming more precise — from a binary ("does session contaminate scores?") to a map.
What this changes, in plain terms
If you compare two quality scores, they are only comparable when the rubric and the administration match — same surface, same isolation, not batched — and you know how the specific scorer responds to a session. A number from ChatGPT deep in a long chat is not the same instrument as a number from ChatGPT in a fresh one. The Observatory now records a BLEED signature for every observer it uses, the way a lab records the drift on each of its instruments. The default protocol for a clean run follows directly: one answer per isolated session, a fixed surface, never batched.
Method: 40 spine answers, blind rubric, each scored isolated vs. in one accumulating session. ChatGPT & Claude via API; Gemini hand-scored on web. All contrasts independently recomputed by DEX from frozen CSVs; raw rows retained. Preliminary — Codex is the next observer to be mapped.
Addendum · 2026-06-14 — the surface decides too
After the manual web run, we scored the same 40 answers on the same model — Gemini — through the API instead of the web app, and the result did not just differ. It reversed. Where web/manual Gemini compressed in a shared session, API Gemini inflated — composite +0.68, Voice +1.25, Frame +1.10. Same weights, same blind rubric, same session manipulation; only the access path changed.
One caveat keeps the API run honest: it is ceiling-pinned. In the shared condition 64% of all its scores are a flat 6 (Accuracy and Clarity sit at ~98% sixes), so as a fine-grained instrument API-Gemini is weak — it is most useful as a stress case, not a precise scorer. That does not soften the headline: the direction of the session effect flipped with the surface.
Session context is observer-specific, direction-specific, and surface-specific.
So the map is no longer a list of models — it is a list of access paths. ChatGPT (web) inflates; Claude (API) holds; Gemini (web) compresses; Gemini (API) inflates and pins the ceiling. BLEED is not a property of "the model"; it is a property of the model × surface × administration path. The practical rule tightens accordingly: a quality score is only comparable to another when the rubric, the session structure, and the exact surface it was produced on all match. The Observatory logs each as a distinct observer surface, with its own BLEED signature.
Update · 2026-06-15 · C
Every scorer uses a different slice of the same ruler
CALIBRATION (O4) · OBSERVATIONAL
Before trusting any score at face value, look at how each scorer actually uses the 1–6 scale. On the same forty answers, five scorers carve the scale up completely differently — one sits low and tight, another stretches across the whole range, a third never leaves the top. These are the scorer's habits, not the answers' quality, and they hide inside every raw number.
The useful move is to read the scores as ranges, not face values — normalize each scorer to its own scale and ask what is left. When we centre every scorer on its own mean, the shapes nearly match: the per-score spread is similar across all of them. The big difference was simply where each one centres, not how it spreads. Grok is the lone shape exception — it refuses the bottom of the scale entirely.
That last question is the point. Normalizing the range strips away the calibration layer — generosity and spread — and leaves the relational structure: the order each scorer puts the answers in. Calibration is removable; whether the scorers agree on the ordering is not, and it is the real test of an instrument.
We call the habit itself Scale Use — a scorer's center (generous or harsh), span (how wide it opens the ruler), and shape (how it distributes around its center). And we read every result through a calibration stack, each step peeling off one layer: face value → centered (severity removed) → range-normalized (severity and span removed) → rank (order only). A finding only counts if it survives the stack. One caution holds it together: matching shapes after centering do not prove matching order — only rank does.
Center tells us where the scorer places the ruler; span tells us how far they open it; rank tells us whether they measured the same shape.
Update · 2026-06-15 · C · six writers, four readers, one consistent panel
Six models, one panel — a five-model plateau, and engagement in only some
COMBINED COMPARISON · 6 WRITERS × 4 READERS · PRELIMINARY
We widened the field. Six models — Claude, Gemini, GPT‑5.5, Grok, Perplexity, and Codex — each answered the same thirty-six questions twice: once plainly, once under WHITMAN. Every answer was then scored on seven qualities by the same four readers. That is one consistent grid — 432 answers, more than 12,000 blind scores, every writer seen by every reader. It lets us ask two things cleanly: who writes better, and who actually changes under WHITMAN.
A ranking is only worth trusting if the readers agree on more than their mood, so first the guardrail: how did the four readers use the 1–6 scale? Raw, their averages sit half a point apart. Center each reader on its own mean and the distributions nearly coincide — the readers differ in level, not in shape. That is exactly what makes a fair comparison possible.
The readers use different centers but broadly comparable scale shapes. That is why calibration is part of admissibility: raw scores alone mix answer quality with observer surface. Once self-cells are excluded and reader centers are aligned, writer comparisons become interpretable.
With the rulers checked, the comparison. Reading every writer through clean means — each model's own family excluded as a scorer, so no one grades itself — the result is a plateau, not a podium. Five of the six — Claude, Perplexity, Gemini, Grok, GPT‑5.5 — land within a quarter point of one another, with no statistically significant gap between any adjacent pair. Codex alone sits clearly and significantly below.
That lower panel is the finding worth keeping. The three frontier chat models — Gemini, Claude, GPT‑5.5 — engage strongly under WHITMAN (lens lift +0.5 to +0.8); Grok, Perplexity, and Codex barely move. And it matches the generation side exactly: the models that refract their formatting under WHITMAN are the same ones that lift the lens.
NATIVE communicates; WHITMAN engages; the lens measures the engagement — and only some models engage. The plateau is the trustworthy claim, not a winner: only after calibration and excluding self-scoring does the comparison become interpretable.