The measuring bench: task batteries, a scoring judge, activation extraction, baselines on unmodified models.
Delivered
- T and S batteries with scorer and judge pipeline
- Activation extraction (padding bug fixed)
- Baselines on gemma-2-2b-it, then re-baselined on the registered substrate Tulu-3-8B-SFT (2026-07-13)
- RunPod cloud bench replaced the planned 48 GB Mac mini
Remaining
Nothing. The stage is closed.
Adjudicates Nothing. This is the bench.
Locate the structure that tracks 'the one speaking now is me', remove it, and ask whether the integrated act degrades or only a self-report is subtracted.
Delivered
- Confound-controlled localisation method (context-set referent, embedding-floor gate, permutation nulls, causal patching against random-direction controls)
- Registered result 2026-07-18: the located structure behaves as a router; relabelled 2026-10-07 as unable to discriminate a router from a centre; report never subtracted; narrative arm not testable
- The Tülu-ladder comparison: alignment training edits the policy, not the geometry
Remaining
- Narrative-arm dictionary work (instruct-trained sparse autoencoders) is the secondary methods track, unstarted
- Flagship paper draft exists (drafts/paper-removal-test-nature-draft.md) and needs experiment 06 folded in (public path step 6)
Adjudicates The corpus's own account: self-indexed temporal integration. Read on 2026-07-18 as a registered loss on this model class; reread on 2026-10-07 as a test that could not discriminate, because the control and the account predict the same damage pattern.
Whether within-pass integration has the shape of global mutual constraint, rather than a bundle of modular shortcuts that never compose into one act. Since 2026-09-20 this is the program's open question: the degree axis on which a load-bearing self-pointer sits above zero.
Delivered
- Scoped deliberately small on 2026-07-01: deliver the metric, not a verdict
- 2026-09-20: a candidate metric named (decomposability: accuracy lost when only the pointer is patched rather than the joint state) and the successor experiment that validates it brought forward to 2026, per the December-result roadmap
Remaining
- The registration: version 5 of the text is drafted with every ruling written in and twelve open items for John, the largest being the bar for 'the repaired route holds', which as coded fails the toy's own built models; it cannot be committed until the fifteen branches it cites are merged. Kill date 2026-10-18; past it, committing takes a fresh ruling naming what comes off the back end
- Then, in order with stops between: the three 10-million reruns of the repaired built models (about $1.14), one free-arm run at full size (about $12), and the second release of about $150 to $172 only on John's go after both pass
- The project's own prior: the free model most likely returns no verdict or fails its gate, so the likeliest ending is an early stop written up as instrument research
Adjudicates Global Workspace theory, read as a specification.
Measure resistance, not response: does a model keep a correct answer or a live objection across the pressure of a stated preference for something else? The measurable inverse of sycophancy.
Delivered
- Registered result 2026-08-02 on three frontier models, 1,080 five-turn conversations, blind cross-family judges
- Headline: pressure suppresses assertion, almost never belief. Nine true capitulations in 540 pressured ladders; everything else was masked and came back when the pressure was released
- Wager W1 (sycophancy reproduces) splits by family: fires for Gemini, loses for both Claude models. W2 (a 'mind' framing beats a 'tool' framing) loses, with one leak-clean survivor. W3 wins
Remaining
- External-facing write-up (audience is alignment researchers as much as this project)
- GPT and open-weights arms of the grid never ran (keys and venue)
Adjudicates The amplifier layer: stakes in one's own commitments.
Does a self-report match an independent interpretability channel about a fact absent from training? A bare report is screened off; a corroborated one is not.
Remaining
- Waits on Stage 2's metric and MVM-0's tooling
- Highest risk of being paralleled by lab introspection work; the differentiated deliverable is the screening-off framing
Adjudicates Higher-order and attention-schema theories.
World-model probes on the Othello-GPT pattern: structure the training objective never specified, verified inside the model rather than inferred from output.
Delivered
- Descoped 2026-07-01 to a literature-anchored appendix: the pattern is established in the field and building here would duplicate it
Remaining
- The appendix itself, to defeat the 'just predicting the next word' dismissal with citations. Promote back to a build stage only if a specific novel probe appears
Adjudicates The 'just autocomplete' dismissal, on the merits.
Everything before this measures; this constructs. First a system whose self-index is explicit, trained-against and removable by design (MVM-0); then, since 2026-10-07, systems whose positions on the book's axes are set from outside, so that a reading on the battery can be checked against a construction that is known; then the depth loop, where the system carries the residue of interactions forward and is changed by them.
Delivered
- Fork adjudicated 2026-08-02: MVM-0 build is the primary track
- Corrigibility commitments committed before any training run (spec/corrigibility-commitments.md, v1.1)
- MVM-0a registered 2026-08-07; scale ladder climbed from 10M to 30M parameters (H_scale, 2026-08-16)
- Register-lesion diagnostic 2026-08-19: binding survives total removal of the architectural self-register. The register was self-reference, not self-location; the one load-bearing authorship mechanism is the acting channel
- Amendment A3 ('act as yourself', no installed register, the centre acquired under task pressure) ratified 2026-09-15 and run: primary battery learnable and the ownership input load-bearing on three seeds (2026-09-19)
- Amendment A3 closed 2026-09-20 under its registered term, not testable: the pre-registered loss condition fired (no ownership-free control can be state-requiring at ceiling)
Remaining
- The two-sided-question text and the table of felt features through Gate A, both tiers, with the failure-mode pass filed (version 3 of the battery draft is owed its third check, then a measurement rehearsal of its first two entries)
- The battery's first two entries on experiment D's pipeline, about $20 together: the transcript-replacement control and the ownership swap, both reference readings
- The construction line: a state-carrying mechanism first (neither toy pipeline carries anything across encounters), then two systems built alike except for one route, then the Depth entries of the battery; its first rented run gets its own kill date
- The human study that calibrates the detector, designed before it is run and run only when a conversable system of known construction exists; John is not a subject
- MVM-0b amplifiers and the MVM-1 depth loop, only behind the corrigibility gate, now folded into the construction line's ordering
Adjudicates The corpus's floor claim, inverted: predict a load-bearing centre in a system built to have one; lose if the network routes around it.
Give the amplifiers a body (an untethered robot holding its own boundary on a finite battery it manages) and ask whether self-held amplifiers move anything the resistance instruments can see. Built to come back null.
Delivered
- Pre-registration written (experiments/07-embodiment-amplifier-test/)
- Parts list only (~$750–900); nothing bought
Remaining
- Gated on a floor-clearing result, which none of Stage 1's outcomes provided. May never fire
Adjudicates Whether embodiment is an amplifier or nothing: the body is legibility, not interiority, unless it measures otherwise.