Project
ALPHA
The scoring and admissibility apparatus for turning differences between answers into evidence.
Purpose
ALPHA exists to prevent the Observatory from mistaking compelling language for measured effect. It scores outputs, freezes evidence, records exclusions, and turns response differences into R-Delta signatures.
Its task is survivability under scrutiny: whether a distinction remains visible after blinding, scoring, provenance checks, and cross-rater pressure.
Summary
ALPHA 2.0/2.1 shifts the emphasis from a single outcome class to a profile of deltas, eligibility decisions, and environment metadata. It treats platform, browser, session state, tool/search status, and scorer conversance as part of the evidence chain.
Preliminary Findings
PRELIMINARY · CANDIDATE FINDINGS · PENDING VERIFICATION
- The environment front was below ALPHA's original resolution: HOST, BROWSER, PLATFORM, and CONVERSANT_STATE were not yet visible distinctions.
- Interface artifacts can become evidence. Grok's follow-up links independently surfaced Set C's Goodhart and Kahneman fault lines.
- Functional WHITMAN verification requires more than upload success; the run must show the expected Goal Legitimacy behavior and output structure.
- The true non-conversant baseline matters. Grok non-conversant Native is identified as a cleaner zero point than any conversant condition.
- Refreshing a browser is not a clean session. Purging browser state or using a fresh profile is required to control conversance.
- Gemini's visible "Exploring..." status may expose platform framing or tool/search augmentation, making it provenance rather than scored response content.
Update · 2026-06-12 · C
The W-Factor: the apparatus caught itself
PRELIMINARY · CANDIDATE FINDING · PENDING VERIFICATION
ALPHA exists to stop the Observatory from mistaking compelling language for measured effect. The first thing it caught was itself. Across Set C, the apparent benefit of structured reasoning tracked, almost perfectly, how much longer the answers had become. The scorers were rewarding length. We named the effect the W-Factor — the verbosity lever — and built a control that removes the length advantage before any effect is read.
The honest sentence for Set C
The canonical reading is now: net of length, structured reasoning is a small drag on one platform and roughly nothing on the rest. Set C is therefore the first clean demonstration of the W-Factor — not a demonstration of a structured-reasoning effect. The original headline is annotated, not erased: history preserved, inference updated.
Admissibility became its own axis
Set D forced a second discipline. A scorer can produce numbers that correlate well and still be inadmissible — reached by a tangled process, or by a model that fabricates a finished-looking sheet. So ALPHA now separates two questions that used to be one: is the number valid (does it track truth?) and is the scorer admissible (is the observer and process clean?). Correlation certifies the numbers; it never certifies the scorer. A well-correlated fabrication is still out.
See the survival index for the full method this feeds, and Instruments for the surface index of active records.