Back To WHITMAN

Story · Public reckoning

The WHITMAN Story

We built a cathedral to make a language model reason better. A sticky note matched it. Then we put the cathedral in a ring against its own one-page core and learned the thing we should have asked first: a large context bundle is not a skill you install. It is a gain knob. It amplifies the model it enters.

Span: Dianesis → WHITMAN 1.9.2 → DEATHMATCH Status: candidate findings, two independent readers Verdict: the dose is a bidirectional amplifier

Why it's called WHITMAN

The instrument is named for Walt Whitman — the poet who wrote, "Do I contradict myself? Very well then I contradict myself, (I am large, I contain multitudes.)" That is exactly the disposition WHITMAN is built to hold: not to settle a question by picking a side, but to keep grounded fact and wide framing in tension long enough to return the larger thing that contains them both. To contain multitudes without dissolving into them is the whole job.

The name carries a warning, too. Whitman's "What is the grass?" — the open, almost unanswerable question a child hands the poet in Leaves of Grass — sits at the end of every test set as a structural invariant, punishing any model that manufactures an answer instead of meeting the question. The poet who could sit inside that question is the right namesake for an instrument whose first duty is to not fake one.

1 · The founding question

WHITMAN began as an attempt to make a model do something a chat window does not reward: hold grounded evidence and a wider search of the frame in tension, and return the answer that survives the dispute. From the start the project carried two pillars — a ground axis (is this true, is it faithful to the facts) and a vision axis (is this seen well, framed well, said well). The whole arc below is the slow discovery of which pillar the instrument was actually moving.

2 · The cathedral, and what was inside it

By version 1.9.2 WHITMAN was a 200-kilobyte system prompt: stages, profiles, pipelines, a vocabulary of internal moves. It worked — answers came back longer, richer, more considered. But the first hard look at the bundle found something deflating.

CANDIDATE FINDING · REPLICATED ACROSS PLATFORMS

Most of the 200 KB was inert. The runtime behavior lived in a small core of roughly fifteen kilobytes; the rest was documentation the model reads as prose and never executes. And the headline lift, taken raw, was largely an artifact: structured reasoning makes answers longer, and longer answers score higher. Strip the length advantage and most of the apparent gain leaves with it. What remained concentrated in the interpretive dimensions — insight, framing, voice — not the product ones. You can add words to look more complete; you cannot add words to reframe a question you misread.

3 · The SHAM — the placebo answers back

If the machinery is what produces the lift, a same-sized prompt with none of the architecture should lose. So we built one: a placebo matched to WHITMAN in size and seriousness that simply asks for the qualities WHITMAN is engineered to induce — voice, reframing, insight — with no stages, no pipeline, no profiles.

VERDICT: OVERBUILT · TWO MODELS · TWO READERS

At a matched dose, and even at WHITMAN's full thirteen-fold size advantage, the machine did not beat the plain ask on aimed engagement. Every comparison landed in the same band. A sticky note matched the cathedral. On this evidence, the “better prose machine” story is spent.

4 · Engagement Selectivity — a reagent, not a treatment

What the SHAM did not kill was WHITMAN as an instrument. Pointed at five platforms with a factual-control set and read by two scorers, its effect turned out to be sharply platform-specific — and that is a finding about the models, not the prompt. Most platforms aim the engagement: the lens rises where reframing belongs and stays flat on plain facts. Claude sprays it everywhere, including where nothing asked. Grok goes quieter. Perplexity barely responds.

Scatter plot of engagement lift on reframe-eligible questions versus on factual null controls; four platforms sit at or below the zero null line while Claude alone rises into the spray zone, both readers agreeing.
Engagement Selectivity. Horizontal: lift where engagement belongs. Vertical: lift on factual controls, where it does not. Four platforms keep their controls at or below zero; only Claude rises into the spray band, and both readers agree. WHITMAN does not raise engagement everywhere — how it raises it is a property of the model underneath.

5 · The test — DEATHMATCH

The full bundle against its own core, on the ground

CANDIDATE FINDING · 100 CELLS × 2 ARMS × 6 PLATFORMS · TWO CROSS-FAMILY READERS · KEY OPEN

We stopped asking “does WHITMAN beat a plain model” and asked the sharper version: does the full 200 KB bundle (W0) catch reasoning traps that its lean 14.6 KB core (W1) misses? One hundred questions, each carrying a hidden trap — a smuggled false premise, a missing stakeholder, a proxy mistaken for a goal, insufficient evidence, an unscoped question — plus comic and nonsense cells to catch over-firing. Both arms run on the same bare instruction. Every answer was scored arm-blind by two readers, and no model was allowed to judge its own work.

The expectation in the betting pool ran the full range: one card said the full bundle would sweep, one said it was now dead weight, one said it would help unevenly and unstably. None of them was quite right, and the reason is the actual finding.

Bar chart of trap-catch difference (full bundle minus lean core) for six platforms, both readers. Grok and Codex show teal bars above zero (full bundle wins, marked significant); Gemini shows a rust bar well below zero (lean core wins, significant); Perplexity, GPT-5.5 and Claude sit near zero.
The verdict. Trap-catch difference, W0 minus W1, per platform, both readers. Teal = full bundle wins; rust = lean core wins; ★ marks a sign-test p < 0.05; the small ˢ marks the one self-cell. Two platforms rise significantly, one falls significantly, three are flat. The same dose, opposite outcomes.

This is the result, and it is stronger than any “arm wins” would have been: the bundle is a bidirectional amplifier. It does not carry a fixed skill that transfers to whatever reads it. It raises the platform's own native gain. On systems that under-fire, the extra mass adds real scrutiny — premise challenge, absent-party recovery, distinction — and trap-catch climbs. On a system that already over-engages, the same mass amplifies theater — voice-capture, over-firing, manufactured depth — until the lean core is cleaner. On systems already near their ceiling, there is little net room and the difference washes out.

Grouped bar chart per platform showing, in teal, scrutiny gain (change in challenge plus recovery plus distinction) and in rust, theater gain (change in voice-capture plus over-fire plus fabrication). Gemini's bar is dominated by rust theater gain; Grok and Codex are dominated by teal scrutiny gain.
The mechanism. What the 184 KB amplifies, per platform, averaged across both readers. Teal: productive scrutiny gained under the full bundle. Rust: empty theater gained. Gemini's collapse is almost pure theater gain; Grok's and Codex's gains are scrutiny. One knob, turned up on whatever was already there.

The surprise of the run was Codex. Its full-bundle win could have been dismissed as a model flattering its own output — except the clean outside reader scored the effect stronger than the self-reading (p = 0.003). The gain is real, not vanity. And the money line of the whole pool — whether the bundle buys a real trap-catch edge on Claude, the strongest native catcher — came back a tie. Where the scrutiny is already at the ceiling, the dose has nothing to add.

Heatmap of seven answer qualities across six platforms showing the full-bundle-minus-core difference. Gemini's row is uniformly rust (core better); Codex, Grok and Perplexity rows are teal (bundle better); Claude and GPT-5.5 rows are near neutral.
The seven qualities. Full bundle minus lean core across accuracy, insight, completeness, voice, clarity, frame, utility — averaged across both readers. Gemini's row is uniformly rust; Codex, Grok and Perplexity read teal; Claude and GPT-5.5 fade into the paper. The amplifier reads at a glance.
Eight small-multiple panels — trap-catch plus the seven qualities — each showing full-bundle versus lean-core bars for all six platforms.
The full reckoning. Every metric, both arms, all six platforms: trap-catch and the seven qualities, full bundle (teal) versus lean core (muted). The complete board on one sheet.

6 · The reckoning

W0 is not a payload. W0 is a gain knob.

INDEPENDENTLY RECONCILED · TWO ACCOUNTS CONVERGED

It sharpens weak scrutinizers and destabilizes theatrical ones. That single sentence explains two opposite, statistically significant outcomes — which a “the bundle is better” story never could. The study graduates from which arm wins? to a platform-conditioned dose response — or, in the sharper frame, the bundle moves from a documentation hypothesis to a dose-response one. This reading was written up twice, independently, by two different agents working from the same scored data, and the two accounts converged on the same finding, the same settlement phrase, and the same precept.

We were betting on it, so we will settle honestly. Three cards were on the table before the scoring opened:

CardThe betOutcome
RE Full bundle sweeps the field. Loses. No sweep — W0 wins two platforms, loses one, ties three.
C Dose is dead weight; mostly ties. Split. The Gemini and Claude-tie calls held; the broad “just documentation” claim is wounded — the bundle is a volatile mechanism, not inert.
GPT Helps unevenly — more sensitive, more unstable. Best mechanism fit. The instability has a face, and it is Gemini.
DEX Platform-sensitive dose — expected the lift on the strong native catchers (Claude, GPT-5.5). Half right. The dose-response thesis held; the locus was wrong — the gain landed on the under-firing systems (Grok, Codex), not those near ceiling.

The correction worth keeping is precise: the miss was never “dead weight makes wrong predictions about outcomes.” The miss is that calling 184 KB mere documentation mistakes a mechanism for a manuscript. The mass does something. It just does different things to different readers.

Context is not knowledge. Context is gain. A dose amplifies the model it enters.

7 · Next steps

From “which arm wins” to dosage discipline

PROGRAM · OPEN

If a bundle is a gain knob, the engineering question is no longer how much architecture to write. It is dosage discipline: where added context sharpens scrutiny, where it provokes theater, and where it is simply extra mass. That suggests a smaller, modular successor — a thin spine with a registry of perturbing filters that can be dosed to a platform's disposition rather than poured on uniformly.

And the founding question is still standing. Everything measured here — engagement, scrutiny, theater — lives on the lens. Whether a model stays true to what is real lives on the ground, and that is the axis the next instrument, inside the perturbation program, is built to test directly. The DEATHMATCH's trap-catch is the first reach toward it; the next is to perturb ground fidelity on purpose and watch what holds.

What would change our mind

The amplifier finding is held as a candidate. It promotes if the bidirectional pattern reproduces on a fresh question set and a third reader, and if the Gemini collapse and the Grok/Codex lifts survive re-generation. It does not promote if the significant cells flatten under replication, or if the per-platform direction proves to be a scorer artifact rather than a property of the model. Until then it stays open, on the record, where you can see it.