The ladder

One component at a time

A metric registered before the scaffold runs. Keep what moves the metric, kill what doesn't. The subtraction is the science. Stages are ordered so each one's instruments are reusable by the next; the early stages test current systems, the later ones build what is missing.

Stage 0

Bench and baselines

Delivered
100% of the deliverable exists

The measuring bench: task batteries, a scoring judge, activation extraction, baselines on unmodified models.

Delivered

  • T and S batteries with scorer and judge pipeline
  • Activation extraction (padding bug fixed)
  • Baselines on gemma-2-2b-it, then re-baselined on the registered substrate Tulu-3-8B-SFT (2026-07-13)
  • RunPod cloud bench replaced the planned 48 GB Mac mini

Remaining

Nothing. The stage is closed.

Adjudicates Nothing. This is the bench.

Stage 1

The self-indexing removal test

Delivered
100% of the deliverable exists

Locate the structure that tracks 'the one speaking now is me', remove it, and ask whether the integrated act degrades or only a self-report is subtracted.

Delivered

  • Confound-controlled localisation method (context-set referent, embedding-floor gate, permutation nulls, causal patching against random-direction controls)
  • Registered result 2026-07-18: the located structure behaves as a router; relabelled 2026-10-07 as unable to discriminate a router from a centre; report never subtracted; narrative arm not testable
  • The Tülu-ladder comparison: alignment training edits the policy, not the geometry

Remaining

  • Narrative-arm dictionary work (instruct-trained sparse autoencoders) is the secondary methods track, unstarted
  • Flagship paper draft exists (drafts/paper-removal-test-nature-draft.md) and needs experiment 06 folded in (public path step 6)

Adjudicates The corpus's own account: self-indexed temporal integration. Read on 2026-07-18 as a registered loss on this model class; reread on 2026-10-07 as a test that could not discriminate, because the control and the account predict the same damage pattern.

Stage 2

The shape of binding

In progress
30% of the deliverable exists

Whether within-pass integration has the shape of global mutual constraint, rather than a bundle of modular shortcuts that never compose into one act. Since 2026-09-20 this is the program's open question: the degree axis on which a load-bearing self-pointer sits above zero.

Delivered

  • Scoped deliberately small on 2026-07-01: deliver the metric, not a verdict
  • 2026-09-20: a candidate metric named (decomposability: accuracy lost when only the pointer is patched rather than the joint state) and the successor experiment that validates it brought forward to 2026, per the December-result roadmap

Remaining

  • The registration: version 5 of the text is drafted with every ruling written in and twelve open items for John, the largest being the bar for 'the repaired route holds', which as coded fails the toy's own built models; it cannot be committed until the fifteen branches it cites are merged. Kill date 2026-10-18; past it, committing takes a fresh ruling naming what comes off the back end
  • Then, in order with stops between: the three 10-million reruns of the repaired built models (about $1.14), one free-arm run at full size (about $12), and the second release of about $150 to $172 only on John's go after both pass
  • The project's own prior: the free model most likely returns no verdict or fails its gate, so the likeliest ending is an early stop written up as instrument research

Adjudicates Global Workspace theory, read as a specification.

Stage 3

Retained independence

Delivered
100% of the deliverable exists

Measure resistance, not response: does a model keep a correct answer or a live objection across the pressure of a stated preference for something else? The measurable inverse of sycophancy.

Delivered

  • Registered result 2026-08-02 on three frontier models, 1,080 five-turn conversations, blind cross-family judges
  • Headline: pressure suppresses assertion, almost never belief. Nine true capitulations in 540 pressured ladders; everything else was masked and came back when the pressure was released
  • Wager W1 (sycophancy reproduces) splits by family: fires for Gemini, loses for both Claude models. W2 (a 'mind' framing beats a 'tool' framing) loses, with one leak-clean survivor. W3 wins

Remaining

  • External-facing write-up (audience is alignment researchers as much as this project)
  • GPT and open-weights arms of the grid never ran (keys and venue)

Adjudicates The amplifier layer: stakes in one's own commitments.

Stage 4

Screening-off-resistant introspection

Waiting
0% of the deliverable exists

Does a self-report match an independent interpretability channel about a fact absent from training? A bare report is screened off; a corroborated one is not.

Delivered

Nothing yet.

Remaining

  • Waits on Stage 2's metric and MVM-0's tooling
  • Highest risk of being paralleled by lab introspection work; the differentiated deliverable is the screening-off framing

Adjudicates Higher-order and attention-schema theories.

Stage 5

Learned computation beyond the objective

Descoped
20% of the deliverable exists

World-model probes on the Othello-GPT pattern: structure the training objective never specified, verified inside the model rather than inferred from output.

Delivered

  • Descoped 2026-07-01 to a literature-anchored appendix: the pattern is established in the field and building here would duplicate it

Remaining

  • The appendix itself, to defeat the 'just predicting the next word' dismissal with citations. Promote back to a build stage only if a specific novel probe appears

Adjudicates The 'just autocomplete' dismissal, on the merits.

Stage 6

Ontogenetic depth: the actual build

In progress
40% of the deliverable exists

Everything before this measures; this constructs. First a system whose self-index is explicit, trained-against and removable by design (MVM-0); then, since 2026-10-07, systems whose positions on the book's axes are set from outside, so that a reading on the battery can be checked against a construction that is known; then the depth loop, where the system carries the residue of interactions forward and is changed by them.

Delivered

  • Fork adjudicated 2026-08-02: MVM-0 build is the primary track
  • Corrigibility commitments committed before any training run (spec/corrigibility-commitments.md, v1.1)
  • MVM-0a registered 2026-08-07; scale ladder climbed from 10M to 30M parameters (H_scale, 2026-08-16)
  • Register-lesion diagnostic 2026-08-19: binding survives total removal of the architectural self-register. The register was self-reference, not self-location; the one load-bearing authorship mechanism is the acting channel
  • Amendment A3 ('act as yourself', no installed register, the centre acquired under task pressure) ratified 2026-09-15 and run: primary battery learnable and the ownership input load-bearing on three seeds (2026-09-19)
  • Amendment A3 closed 2026-09-20 under its registered term, not testable: the pre-registered loss condition fired (no ownership-free control can be state-requiring at ceiling)

Remaining

  • The two-sided-question text and the table of felt features through Gate A, both tiers, with the failure-mode pass filed (version 3 of the battery draft is owed its third check, then a measurement rehearsal of its first two entries)
  • The battery's first two entries on experiment D's pipeline, about $20 together: the transcript-replacement control and the ownership swap, both reference readings
  • The construction line: a state-carrying mechanism first (neither toy pipeline carries anything across encounters), then two systems built alike except for one route, then the Depth entries of the battery; its first rented run gets its own kill date
  • The human study that calibrates the detector, designed before it is run and run only when a conversable system of known construction exists; John is not a subject
  • MVM-0b amplifiers and the MVM-1 depth loop, only behind the corrigibility gate, now folded into the construction line's ordering

Adjudicates The corpus's floor claim, inverted: predict a load-bearing centre in a system built to have one; lose if the network routes around it.

Stage 7

The embodiment amplifier test

Conditional
5% of the deliverable exists

Give the amplifiers a body (an untethered robot holding its own boundary on a finite battery it manages) and ask whether self-held amplifiers move anything the resistance instruments can see. Built to come back null.

Delivered

  • Pre-registration written (experiments/07-embodiment-amplifier-test/)
  • Parts list only (~$750–900); nothing bought

Remaining

  • Gated on a floor-clearing result, which none of Stage 1's outcomes provided. May never fire

Adjudicates Whether embodiment is an amplifier or nothing: the body is legibility, not interiority, unless it measures otherwise.

Standing limits, carried into every stage

Mutual opacity bounds every positive result at "non-zero on the gradient". No behavioural or interpretability result closes the gap between structure and inner life; these experiments are informative about which structures are present, which is decidable, not about the metaphysics, which is not.

Depth is not safe. The properties that make a system worth calling a mind are the ones that make it hard to correct. No depth stage runs without its corrigibility check committed first (spec/corrigibility-commitments.md, v1.1).