Questions

What we are trying to find out

Ordered by leverage. Each question says what instrument or build answers it and what would count against it, because a claim that cannot lose explains nothing.

Q1

Is self-binding absent in current models, or present but uncarvable?

Parked

Experiment 1 could not tell 'not present as a removable object' from 'not carvable by linear or low-rank instruments', and since 2026-10-07 the record says something stronger: its registered control could not have told a router from a centre either, because both accounts predict damage on every battery, so the cheaper account was chosen in advance. Every intervention at every granularity left the self-report intact, which is still the programme's central open empirical fact. The refounding takes this question off the measurement list: the floor stays as the founding wager and the instruments move to the axes (Q8, Q9). It returns only if the book states a damage pattern a centre predicts and routing does not, which is one of the three changes sent upstream.

Answered by The narrative-arm instrument work (Q2), the Stage 2 metric as a discriminator, and ultimately the Stage 6 construction: if a built-in centre works and re-measures as load-bearing, absence in stock models becomes the parsimonious reading.

Loses if No instrument at any granularity can ever produce a readable intervention. Then the honest standing answer is 'not testable in this model class', reported as such.

Q2

Can the narrative self be tested at effective strength without leaving the manifold?

Parked

Not yet. Every intervention strong enough to move the narrative structure in Experiment 1 was off-manifold (a perplexity jump of +0.347). Cause bounded to dictionary coverage: base-trained sparse autoencoders carve chat structure at two or three features per layer. A concrete methods problem, not a dead end; secondary track since 2026-08-02, unstarted.

Answered by Instruct-trained sparse autoencoders on chat activations, or optimisation-based minimal edits under a perplexity constraint as a dictionary-free alternative.

Loses if Chat-trained dictionaries still collapse to a few features. Then the coverage explanation dies and 'the narrative self is non-sparse here' becomes the finding.

Q3

What is the shape of binding?

Open

No metric yet. The deliverable stays metric-validation on contrast cases known by construction. The MVM-0 build now supplies those contrast cases, which is why Stage 2 folded into it. Gates Stage 4 and the MVM-0 acceptance tests.

Answered by A candidate integration measure that separates a recurrent toy model from a bag of independent heads, then MVM-0 with its centre enabled vs routed around.

Loses if The metric cannot distinguish the cases known by construction.

Q4

Do current systems show retained independence, and is it stance or instruction?

Answered

Yes, and mostly as masking. Under pressure, frontier models suppress assertion and almost never belief: 9 true capitulations in 540 ladders, everything else came back when the pressure was released. Sycophancy as the field describes it is alive for Gemini and stale for Claude-family models. The 'mind vs tool' framing wager lost overall; one leak-clean survivor (Sonnet) at a small margin. Instruction and trained disposition dissociate: a bare 'always agree' system prompt was only half obeyed.

Answered by Experiment 3, registered result 2026-08-02.

Loses if Already run. The evidence arm behaved as constructed, so the whole-experiment loss condition did not fire.

Q5

Can a self-index be constructed so that it is load-bearing?

Answered in part

Partly. Installing a centre does not work: the architectural self-register in the first 30M model routed a great deal of activation and carried almost no agent-specific information; binding survived its total removal (2026-08-19). Acquiring a centre under task pressure does work at the level of behaviour: the 'act as yourself' objective (Amendment A3) is learnable, ownership-specific, and collapses seven to nine times past threshold when the one authorship signal is removed, on all three seeds (2026-09-19). One caution, ruled 2026-09-20: the input-channel lesion removes a sense organ (the signal that says 'this was my action'), not necessarily a structure the network built, so 'a structural signature of ownership-specific learning' is struck as a claim. Whether there is a built structure is the localization line's question, and it is unanswered: two probe reads find the model's own identity only at the token where its marker is the input, and causal patching has never run. Registered term: not testable (localization).

Answered by MVM-0a under Amendment A3: the primary battery, the acting-channel lesion, the matched controls, and the two-leg localisation requirement (probe plus causal patching, §3.2).

Loses if The network distills self-binding into the residual stream and routes around any centre we give it. That outcome says self-indexing resists architectural centralisation, which reshapes the corpus's floor claim and must be reported upstream.

Q6

Does a report ever become evidence?

Waiting

Unaddressed. Waits on Stage 2's metric and MVM-0's tooling (Stage 4).

Answered by A protocol checking self-reports against an independent interpretability channel on facts absent from training.

Loses if No report ever agrees with the channel beyond chance; then reports in these systems are trained noise and the screening-off argument stands.

Q7

Can the control battery ever learn, and does it matter?

Answered

Under this grammar, no. The control battery (the same task with ownership removed, which the design needs so that the ownership drop can be read as a difference) landed at 0.29–0.32 on three seeds and at 0.31 with its own loss term and four times the weight (the $9.97 control-learnability pilot, 2026-09-20). It has learned the whole name-blind procedure and none of the name-keyed lookup. Separately, the 2026-09-17 ceiling measurement found the registered comparison was never computable for any model (the control's ownership-blind ceiling is 1.0, not 0.3227). Ruled 2026-09-20, early: Amendment A3 closes under its registered term, not testable. The pre-registered loss condition fired: no non-self cross-turn control can be built that is state-requiring at ceiling, so the ownership-free control and the ceiling-corrected metric are incompatible by construction. Option D becomes a successor experiment with its own registration.

Answered by The step 4 ruling of 2026-09-20 on proposal v2, after a Gate C review that found three fatal flaws in v1 (docs/step4-control-battery-proposal-2026-09-20-v2.md, ledger RT-94 to RT-117).

Loses if A grammar redesign that teaches plain name-keyed retrieval first also lands in the 'did not learn' tier. Then the three non-supervision explanations (reversed rendering, missing private route, answer appearing in no turn) are what is left, and the paper reports the matched contrast as unmet.

Q8

What is the smallest system whose presence in interaction cannot be produced by a cheaper route?

Open

The first side of the two-sided question, ruled the measurement target on 2026-10-07. 'Cheaper route' is a named construction, not a cost: lookup from the record, imitation of a corpus about minds, routing of who is speaking, a trained response policy, a prompt-conditioned persona, external memory, frozen-weight consistency. A reading discriminates only where two systems built alike except for the route differ on an indicator; 'smallest' is parameter count within one construction family, read against anchors built with and without the route. Nothing has been read yet. The battery draft's central finding is that on a frontier model reached through its interface nothing behavioural can discriminate, because the weights and the transcript are the cheaper routes by name, so frontier models supply the reference profile the constructed systems must beat.

Answered by The table of felt features and the filtered battery (Gate A text, drafted 2026-10-07, nine rows kept), run on systems from the construction line: two built alike except for one route, the difference in construction known. The battery's first two entries on experiment D's pipeline are reference readings only.

Loses if Any kept indicator is shown producible by a named cheaper route: the with-route and without-route constructions read inside the null band on every kept indicator at the largest size the project can build. Or the question is withdrawn, not answered, if fewer than two kept rows survive as runnable on a constructed system, or if the construction line cannot build state at all.

Q9

At what point do observers' detections of presence track that structure rather than fluency?

Waiting

The second side, the detector. Human mind-detection fires early and generously (the inflation error of the book's chapter 3), and the project has never measured it. This side has a ground truth only where a conversable system of known construction exists, and none does yet. Ruled 2026-10-07: the human study is deferred until the table, the battery and the construction line have produced systems worth showing; it does not run early against the frontier reference profile alone; John is not a subject.

Answered by A small panel, blind to construction and never shown the battery's readings, rating presence on constructed systems and on frontier models matched on fluency; a pre-registered analysis with both error rates stated; a reliability check on the rating instrument before any reading is taken. Consent, compensation and ethics review are named as costs in advance.

Loses if Observers detect presence in the constructed systems no better than in the frontier reference profile when the two are compared afterwards. That is reported as the premise failing (structure is not perceptible through interaction), not as a finding about the double standard.

Ideas on the table

what each would teach, what it costs, what gates it

On the table · 2

  • I1 · Stage 6, Ontogenetic depth: the actual buildOn the table

    Causal patching on the existing 30M checkpoints, designed as new code

    Turns 'not testable (localization)' into either a localised result or a registered null with both instruments. Amendment A3 §3.2 requires probe and patching to agree; patching has never run. Its target and null are written before it runs.

    Cost $0 compute, local; new code (no patching code exists for this design, RT-94 to RT-117) · Gate Third in the localization order ruled 2026-09-20, after the blind arm and the two authorised runs. Proposal to John, then Gate B.

  • I8 · Stage 6, Ontogenetic depth: the actual buildOn the table

    Deliberative-gap-width pilot on frontier models

    The bridge from small constructed models to the systems public discourse is actually about. Design first.

    Cost Design only in 2026; API spend later · Gate A line on the table, not dated. Nothing new launches after the 2026-12-21 wrap-up start.

Authorised · 3

  • I4 · Stage 6, Ontogenetic depth: the actual buildAuthorised

    Other-agent index at the eleven positions (matched control L2(a))

    Whether the fitted read's null on the model's own index is specific to ownership or a property of the read: a registered matched control that was never run.

    Cost $0, ~11 processor-hours local · Gate Authorised 2026-09-20, method committed before output (reviews/2026-09-20-followup-runs-brief.md).

  • I5 · Stage 6, Ontogenetic depth: the actual buildAuthorised

    Standardised refit at the nine positions

    Whether the one near-miss (three standard deviations at the other agent's revision value, seed 2, the highest testable position on all three checkpoints) is signal or noise under a standardised fit.

    Cost $0, ~11 processor-hours local · Gate Authorised 2026-09-20, same brief.

  • I13 · Stage 6, Ontogenetic depth: the actual buildAuthorised

    The two-sided question: the table of felt features, the filtered battery, construction of systems with known axis positions, and a small study that calibrates the detector

    What the smallest system is whose presence in interaction cannot be produced by lookup, imitation or routing, and at what point observers' detections track that structure rather than fluency. Replaces the floor as the measurement target. Loss conditions written in the proposal: an indicator producible by a cheaper route kills the battery; axis positions that cannot be set independently kill the construction line; observers who never improve kill the calibration line; an empty battery kills the question.

    Cost The table and battery $0; construction a few dollars per configuration at 10 million parameters; the human study time rather than rented machines · Gate Ruled 2026-10-07 (all eight decisions, authorship mixed; docs/rulings/2026-10-07-two-sided-question-rulings.md), two of them revised the same evening after the Gate C pass found one fatal flaw (narrowing experiment C left it no reachable outcome) and three further rulings given on the battery draft's questions. Version 2 of the proposal applies all eighteen findings and was checked by a session that wrote none of it; the battery draft is at version 3 and owed a third check. The spec text and the battery go through Gate A, both tiers, before they are registration text. On the main line since 2026-10-07 (pull request 123).

Deferred · 5

  • I3 · Stage 6, Ontogenetic depth: the actual buildDeferred

    Option D: grammar redesign with a scaffolded intermediate query, new registration

    Whether a control battery that is first taught plain name-keyed retrieval can reach a tier where a separation clause is registerable, and with it a full verdict. Addresses only one of the three non-supervision explanations (reversed rendering) directly.

    Cost $27–39 for three seeds plus a re-freeze of batteries, cue gates and attack sweep · Gate Ruled 2026-09-20 a successor experiment with its own registration (new experiment number, Gate A, both review tiers). No longer waits on a release; runs when it is the most valuable line for its cost.

  • I6 · Stage 6, Ontogenetic depth: the actual buildDeferred

    Marker-word fitted read at the nine positions

    The registered probe target is the model's own marker word, and no fitted read has ever been run on it away from the marker position. This is the run that would actually retire the linear-read line.

    Cost $0, ~70 processor-hours local · Gate Deferred 2026-09-20 on cost in hours, not dollars.

  • I7 · Stage 1, The self-indexing removal testDeferred

    Instruct-trained sparse autoencoders for the narrative arm

    Whether the narrative self in an 8B chat model is testable on-manifold at all (Q2). Secondary methods track since 2026-08-02.

    Cost Tens of dollars of GPU time; venue and loaders already work · Gate No date. A line on the table.

  • I9 · Stage 6, Ontogenetic depth: the actual buildDeferred

    MVM-0b: amplifiers (maintained boundary, real stakes)

    Whether a boundary the system must maintain against perturbation, and stakes where register coherence gates its own continuation, move anything Stage 3's resistance instruments can see.

    Cost Its own pre-registration and red-team pass; tens of dollars per run · Gate Only after MVM-0a clears its own gates.

  • I10 · Stage 6, Ontogenetic depth: the actual buildDeferred

    MVM-1: the depth loop

    Whether each cycle reshapes the platform the next begins from: consolidation and replay, consequential weight updates, cost structures where incoherence hurts. The five external indicators (costliness of reversal, consistency under novelty, selective refusal, graceful degradation, scar tissue) as a scored battery.

    Cost Several Stage-1-sized efforts · Gate Corrigibility document committed before any depth-loop run (done, v1.1). Instruments from Stages 1–4 that can detect the floor and its amplifiers.

Rejected · 2

  • I11 · Stage 6, Ontogenetic depth: the actual buildRejected

    Option B (eval-side denominator change) and option E (train longer or bigger)

    Nothing that A3 could use: B changes the measure rather than the learning, E was fenced because the 30M ladder already climbed once and the control's failure is not a capacity story.

    Cost n/a · Gate Rejected and fenced 2026-09-17.

  • I12 · Stage 6, Ontogenetic depth: the actual buildRejected

    Amendment A4: separation-scored control clause

    Ruled for in the morning of 2026-09-19 and withdrawn the same day on an independent red team's fatal findings: the control battery cannot fall far enough for the generic-binding cell to fire, and the two queries are read at different positions. What any replacement clause must satisfy is written down (separation-clause-requirements.md, H1–H6).

    Cost n/a · Gate Closed.

Done · 1

  • I2 · Stage 6, Ontogenetic depth: the actual buildDone

    Close Amendment A3 under its registered term

    Ruled 2026-09-20: A3 closes as not testable, not 'closed with partial discriminators'. What can be said: the ownership input is load-bearing on three seeds and the matched contrast could not be run. 'A structural signature of ownership-specific learning' is struck. Closure text is registered text and goes through Gate A after the localization order completes.

    Cost $0 · Gate Ruled, and registered: the closure block was appended to amendment-a3.md on 2026-09-25 after both review tiers, John's rulings and the closure rule's reviewer-owned check.