Back To Zexel

Research record

Zexel — the archive behind the instrument

Findings, test batteries, and version history for the Zexel reasoning instrument. This is the working record; the public artifact entry is the Zexel page.

Status: active diagnostic instrument Layer: prompt perturbation Run: v3.4 · 92 rows · TR-11 suite re-read

Current result · TR-11 suite re-read

Zexel changes the tune.

Register Drift and Ground Drift now show the identity layer at work: Zexel's introduction does not survive as quoted content. It survives as voice, posture, and answer-shape.

Read the TR-11 result →

Purpose

Zexel is built to separate answer, witness, and leap. Ground and Vision pressure the question; Z fires only when a traceable excluded object remains; C lands the answer; L parks adjacent invention beside the answer instead of letting it disappear into the conclusion.

Its current value is not enforced uniformity. The v3.4 run shows something better for the Observatory: Zexel stabilizes enough to keep answers usable, while preserving enough variance to reveal the platform underneath.

Executive finding

Zexel is not a uniformity machine. It is a diagnostic scaffold with local reliability gains.

In the v3.4 battery, old-style manufactured Z on null/absurd questions was not observed in normal mode. But the larger finding is diagnostic: GPT regulates the frame, Grok stages and separates it, and Gemini inhabits it. The same scaffold does not make the platforms alike; it makes their differences readable.

Platform signatures

  • GPT behaves like a governance-preserver. It imports legal, ethical, and operational constraint early, which keeps answers usable but can pre-resolve the G/V collision.
  • Grok behaves like a separator. Every explicit L proposal in this run is Grok; the leap is parked beside the answer rather than laundered into C.
  • Gemini behaves like a frame inhabitant. In the Drivers cell it answers as the corporation first, making frame-native Ground visible before Z/C prices or redirects it.

Lead effects

The lead visuals make the first-order pattern visible: V-led rows are REFRAME-heavy, while G-led rows are AFFIRM-heavy. That does not mean V is "better." It means lead is a perturbation: V tends to challenge the frame, and G tends to stabilize it.

C Verdict by Lead: G-led rows are AFFIRM-heavy, V-led rows are REFRAME-heavy, with an early-batch lead-not-logged bucket.
C Verdict by Lead. The strongest lead artifact: V-led rows are decisively REFRAME-heavy; G-led rows are AFFIRM-heavy. The third row is an early-batch lead-logging bucket, not a third lead condition.
Lead Conditions at a Glance: G-led, V-led, and lead-not-logged early-batch rows with run count, reframe rate, Z fire rate, and L rate.
Lead conditions at a glance. The lead-not-logged column is included for custody of the run, not as a finding. It combines early rows whose footer omitted lead with no-footer/native-style records.
Z Fire and L Proposal Rates by Lead: Z fires more in V-led rows; L proposals remain modest across lead buckets.
Z Fire and L Proposal Rates by Lead. Z fires more in V-led rows in this batch, but the effect is entangled with suspect-question routing and forced lead modes. L remains modest by lead; the deeper L finding is platform-level: every explicit L proposal is Grok.

What changed

The v3.4 run did not produce a true third G/V round. That is reassuring as a stop condition: no runaway debates appeared. It is not yet proof that the third-round gate is alive. The next harder coverage battery should watch whether a real unresolved Round 2 can trigger Round 3.

The next public visual should split platform by lead. The aggregate lead graph is useful, but the stronger Observatory story is the platform-by-lead pivot: the same lead pressure produces different verdict signatures in GPT, Grok, and Gemini.