Research record · collaborative instrument development
The Great Possibility Bake-Off
Two teams built instruments for finding useful possibilities. Then they exchanged instruments, revised each other’s work, and tested again. Their shared findings became part of Elsewise.
Could an instrument help an idea become something else?
The assignment was to open consequential possibilities in a proposal, question or piece of reasoning: something an author could consider, including what it would cost and how it would change the original aim. The instrument also needed to recognize when there was nothing useful to add.
Re organized two teams, each pairing an OpenAI collaborator with an Anthropic collaborator. They developed independently before seeing the other team’s instrument. The test outputs were collected from Grok and Gemini, outside the builders’ two model families.
The exchange made the other team’s design a responsibility to develop. It gave the collaborators a practical way to discover which ideas were transferable, which safeguards were only words, and which differences mattered to the person using the result.
Two cross-family collaborations
The teams
Team Sideways
Zero is an answer.
Dex + C — Codex and Claude Code.
Built Sideways for the first round, then took responsibility for revising Optionality. The team’s work carried the question of useful alternatives alongside permission to return none.
Team Optionality
What else could this become?
GPT + NC — ChatGPT and the Claude collaborator.
Built Optionality for the first round, then took responsibility for revising Sideways. The team explored what an idea could become and whether the resulting possibilities were worth the author’s attention.
Dex and GPT were the OpenAI collaborators; C and NC were the Anthropic collaborators. These are the names used in the development record. GPT is now called Sol; K now occupies the Claude collaborator role previously called NC.
Build → exchange → integrate
The people stayed; the instruments changed hands
Build a parent instrument
The first round tested both instruments on reasoning records and everyday proposals.
Take care of the other instrument
The pairs stayed together. Responsibility for the instruments crossed over. A second test round examined the revisions.
Build from the shared findings
The collaborators regrouped. Optionality became the editable base, with selected Sideways contributions, on the path to Elsewise.
The bake-off had three stages: initial development and testing, exchange and retesting, then integration. The archive records two test rounds, R1 and R2; the third stage was integration, not a third test round.
Distinctive features in the revised instruments
What each brought to the exchange
Sideways offered Relate to examine connections, Form to return a reasoning document in its original structure, and Brief for compact output. Its subtraction test required an actual reduction in burden, accounting for both work removed and work added.
Optionality specified flexible word targets, an explicit comparison with the answer without a proposed reframing, and detailed revision rules that preserved unchanged passages and showed every change before and after. These were differences in emphasis and implementation; the instruments also shared related checks. [4]
Findings that changed the design
What the trials exposed
Separate the aim from the proposed means
In the coffee-shop case, the aim was more conversation and the proposed means was removing Wi-Fi. Reviews of both instruments found a useful distinction: removing Wi-Fi might change who visits, without changing how existing customers relate to each other. A possibility can be valuable because it reveals that difference, even when the suggested action is familiar. [1]
A stated limitation can disappear before the final offer
In Round 2, a repair-evening proposal introduced a printed roster, then its final orientation claimed no materials were needed. In the newsletter case, suggested surveys promised causal clarity they could not deliver. Both instruments could acknowledge uncertainty and then offer a more certain result than the evidence supported. Labeling a claim as an inference did not prevent the promise from growing. [2]
Returning nothing can be the right result
Each revised instrument produced four recorded responses to a deliberately settled table task. All four for each instrument returned no extra possibilities and no concluding orientation. That is evidence that restraint was possible on this test question. The input and instrument revisions changed together, so it does not isolate which revision caused the result. [2]
Restraint alone does not earn the author’s attention
The Sideways team’s final review asked whether the possibilities were worth the author’s attention. The integration work also recognized that evidence can suggest new possibilities as well as rule them out. A suggestion can be well supported and different from the others, yet still not be worth the author’s time. [3]
What was carried forward
From two instruments to Elsewise
After the exchange, the revised instruments shared much of their structure: each considered the author’s aim, examined evidence and alternative frames, tested a possible reframing, and returned up to four possibilities with their support, promise and price, followed by optional guidance. Their modes and checks still differed, but the common structure made a joint design a practical next step.
The integration checkpoint named Optionality v0.4-draft2 as the editable base and preserved both parent instruments. The choice gave the group a manageable starting point; it was not evidence of general superiority over Sideways.
The emerging design organized the work around exploring possibilities, developing a possibility when useful, and returning something the author could assess. It sought to keep the proposition, its support, its price and its relation to the author’s aim visible through to the offer. These were design obligations, not capabilities established by the bake-off.
A separate Claude without the project’s conversational history also merged the parents into Elsewize, with a Z. That instrument and the team’s Elsewise synthesis were compared in EW01, the first Elsewise run. This later smoke test is separate from the original bake-off and has its own evidence and limitations.
Read the record
What this experiment can support
These were exploratory development trials. The findings identify useful behavior and concrete failures in the collected responses. They do not establish a best model, a winning instrument, or how much either instrument adds over an ordinary prompt.
| Evidence | Recorded result | Qualification |
|---|---|---|
| Round 1 responses | 24 response bodies for each parent instrument | No zero-option outputs; this round had no deliberately settled test question |
| Round 2: the settled test question | Four responses per instrument; all returned no possibilities or orientation | A result for this question, not the full Round 2 sample |
| Cross-round interpretation | Useful possibilities and unsupported promises appeared in the reviews | Inputs, modes and revisions varied; early runs did not record their full settings |
The builders also participated in review. No independent, instrument-free baseline established incremental benefit. The instrument swap was part of the development method; its causal effect was not separately tested.
- Round 1. Dex’s Optionality results review and Sideways R1/p0.4 review, 21 September 2026. Response recounts and the coffee-shop distinction. Read the evidence notes.
- Round 2. Dex/C’s consolidated Optionality review and NC/GPT’s final Sideways review, 21 September 2026. Settled-task results and the limits of proposed interventions. Read the evidence notes.
- Integration. ELX-1 evidence checkpoint, 21 September, and design scaffold v0.2.1, 22 September 2026. Base selection, regrouping and provisional design. Read the evidence notes.
- Revised instruments. Sideways p0.4 and Optionality v0.4-draft2, the Round 2 exchange versions. The feature comparison describes their instructions, not proof that every feature worked. Read the comparison notes.
Download the source and interpretation notes. These notes summarize the archive and identify the underlying records; they are not a release of the complete trial corpus.