Where we are · updated Oct 8, 2026

On 2026-10-07 the project was refounded, and on 2026-10-08 the refounding reached the main line and the site was brought up to date with it. What changed: a brief asking whether the project was measuring the right thing was checked against the record and put, by API and in two rounds, to Gemini 3.1 Pro, GPT and Claude Opus 5.5, every reply saved word for word. All three agreed that experiment 1's router control could not tell bookkeeping from a centre (a check the same day reached the same conclusion from the registration text: both accounts predict damage on every battery, and the control fired by a margin inside the noise), that the project's reframe added nothing beyond the identity bet, and that no second release of the degree experiment should be funded now. John then named what he perceives from: what it is like to be among minds that are present or shallow on the book's axes. That perception is of the axes, not of the floor. He ruled all eight decisions of the refounding proposal, two of them revised the same evening after the Gate C pass: the two-sided question is the measurement target; experiment A is relabelled as unable to discriminate; experiment C keeps both releases, the second conditional on the first; experiment D is promoted to the pipeline and reference profile of a filtered battery; the two kill dates stand and no new calendar dates are set; the table and battery are the next paper work; the human study is deferred; the book changes go upstream. Pull request 123 carried all of it to the main line at his word. Two drafts wait on checks: version 5 of the registration text, with twelve open items (the largest: the bar for 'the repaired route holds' is unruled and, as coded, fails the toy's own built models), and version 3 of the battery draft. About $229.62 of the $450 ceiling is spent; nothing has been rented since the four development runs of 2026-10-04.

Two things run side by side to the end of the year. The first is the degree experiment (experiment C), which the December-result roadmap still binds: the registration committed by 2026-10-18 with the $1.14 repair written in, the repaired built models verified (a failure stops C there), one free-arm run at full size (a gate or floor miss stops C there), and the second release of about $150 to $172 asked for only on those figures, a pass being necessary and not sufficient. The project's own stated prior is that the free model most likely returns no verdict or fails its gate, so the likeliest ending for C is an early stop written up as instrument research, which is a registered outcome and not a schedule failure. The second is the new question, in task order with no step on a date: the two-sided-question text and the table of felt features through Gate A; the battery's first two entries on experiment D's pipeline by API, about $20 together, as reference readings of what the cheaper routes produce; then the construction line, small systems whose positions on the axes are set from outside, which needs a state-carrying mechanism neither toy pipeline has, and is where every test that can discriminate lives; last, and designed before it is run, the human study that calibrates the detector. By 2026-12-21, realistically: C run to a registered outcome, the new question in the spec, the table and the first two battery entries registered and run, the relabelling of experiment A in the paper draft, and the packet of three changes sent to the book. Not this year: a degree reading of a free model, a discriminating depth reading, or an observer-side measurement. Loss conditions are written for every new line. Claims are whatever the registered results support: a position on the axes, the route named, the degree measured, never a verdict. Never a conscious machine.

The next most valuable steps

and what each can teach us
N4Pending

Causal patching: built once, for the successor experiment's grammar, not for the A3 checkpoints

Ruled 2026-09-20: patching the A3 checkpoints would build a tool for a model it cannot answer on (no matched-role comparator, target never localized). The A3 closure no longer waits on it.

with the successor implementation, which starts once the registration is committed · Claude Code session · $0 compute; new code

N7Pending

Refresh explainer.md

The plain-language companion last touched 2026-07-01 predates every registered result since. It carries the founding wager once that lands.

after the 2026-09-30 book text lock, no date · Cowork session · $0

N18Pending

Small things owed from Weekend 1: the correction note beside the old review's key count, the launcher's deadline gaps and acknowledgement wait, the settled billing figures in the ledger (the old agent working folders, also on this list, were removed on 2026-10-03)

Nothing new. Each is ruled or already identified; none is done.

any quiet hour; before the first full-size launch for the launcher items · Claude Code session · $0

N24Pending

Check the frozen code and the development runs' results, by a session that wrote neither

Whether the port of the toy scripts into one procedure is faithful, whether the outcome rule and the tripwire do what the rulings say, and whether the trainer ran the loop it claims on the rented machine.

once the development runs are home · a Claude Code session that did not write the code · $0

Waiting on John

  1. Merge the fifteen checked branches the registration text cites, so that every citation resolves at a main-line commit (before the registration can be committed; the sharpness branch (pull request 118) is owed its check first)
    Nothing new. A registration may not rest on a record its reader cannot open at a main-line commit (failure 4 of the known-failure list), and fifty of version 5's citations point at unmerged branches today.

The question

Two sides, ruled 2026-10-07 (spec text pending Gate A): what is the smallest system whose presence in interaction cannot be produced by a cheaper route (lookup, imitation of a corpus about minds, routing of who is speaking), and at what point do observers' detections of presence begin to track that structure rather than fluency? The earlier question, whether the smallest self-centred act of integration (the floor) is present in a system, stays in the spec as the founding wager and stops being a measurement target.

Take the consciousness account developed at Sentient Horizons and The Calibration Problem as a specification, build and measure against it, and let every claim carry a stated loss condition.

The bet under everything

the presupposition, not a hypothesis

Sufficiently deep self-indexed temporal integration is experience, seen from within. A system whose architecture binds past, present and anticipated state into one act centred on itself is a bearer of experience to the degree that it does so. The functional agent with all of that architecture and no one home is declined: it can be imagined, it predicts nothing, licenses no experiment and sorts no system, so it pays no rent.

Its status A bet that cannot lose, so by the project's own rule it is not a hypothesis. It enters as a presupposition, the way induction presupposes that nature is uniform: the condition under which structural evidence is evidence about experience. Reject it and every result here still stands as a claim about which structures are present; it just stops bearing on experience.

Why still never a verdict The wager makes the evidence evidential without making it conclusive. Two limits remain, and neither is the zombie: the identity is a bet that correlational evidence cannot carry, and the instruments can be wrong about structure before any question of experience arises. That is what mutual opacity names here.

The Calibration Problem, ch. 5 and Appendix A; spec §The Founding Wager (proposed 2026-09-20, Gate A pending)

Three nested goals

kept separate so expectations stay honest
G1Delivered

Instruments that can lose

Turn the floor claim (consciousness is self-indexed temporal integration, adjudicated by the removal test) into pre-registered, falsifiable instruments. Achievable regardless of what the instruments find, and where most of the durable value lives: the field has position papers; this repo builds tests that can fail.

Where it stands A fully registered removal test ran on 2026-07-18 and its own pre-committed control voided the headline signature when it appeared. The discipline (register, red-team, lock, run, report verbatim) has since carried Stage 3 and the whole MVM-0a programme, with an outside-review protocol in force since 2026-09-19.

G2In progress

Measure current systems, and say which route produced what they show

Run the instruments on models that exist. Since the refounding of 2026-10-07 the target is the axes an observer perceives (does reversal cost the system anything; did its history change it or only its record; is anything at stake), not the floor. On a frontier model reached through its interface, the trained weights and the transcript are the cheaper routes by name, so what these models give is a reference profile of what lookup, imitation and routing produce. The output is bounded at 'non-zero on the gradient' or 'not testable here', never a verdict.

Where it stands Experiment 1 (8B instruction-tuned model): a self-tracking structure was located and removed, and the registered control could not tell a router from a centre (relabelled 2026-10-07); the self-report was never subtracted. Experiment 3: frontier models mostly mask rather than capitulate under pressure (nine true capitulations in 540 ladders). Both are registered results. Next on this goal: experiment D's transcript-replacement control and the ownership swap, about $20 together, which read which cheaper route produced the retained-independence result.

G3In progress

Build what is missing and re-measure

Construct small systems whose structure is known, because that is the only place a reading can discriminate: two systems built alike except for one route, read on the same indicator. MVM-0a, the constructed self-index, was the first step. The degree experiment (C) is the second: built models at both ends of a scale and a free model read between them. The construction line of the two-sided question is the third: systems whose positions on the axes are set from outside, the smallest of which that passes the battery is the minimum viable mind, within its construction family.

Where it stands MVM-0a exists as three 30-million-parameter checkpoints; the ownership input is load-bearing on three seeds, and Amendment A3 closed 2026-09-20 as not testable because its matched contrast could never be run. The degree experiment's code is frozen and reproduces the toy record; at 10 million parameters two of three built models switched their built-in route off, so a $1.14 repair is written into the registration and verified before any full-size run. The construction line needs a state-carrying mechanism neither pipeline has; nothing of it exists yet beyond the battery draft's entries 3 to 8.

  1. Stage 0Delivered Bench and baselines
    100%
  2. Stage 1Delivered The self-indexing removal test
    100%
  3. Stage 2In progress The shape of binding
    30%
  4. Stage 3Delivered Retained independence
    100%
  5. Stage 4Waiting Screening-off-resistant introspection
    0%
  6. Stage 5Descoped Learned computation beyond the objective
    20%
  7. Stage 6In progress Ontogenetic depth: the actual build
    40%
  8. Stage 7Conditional The embodiment amplifier test
    5%

On the table right now

all questions and ideas →