Oct 7, 2026 · Stage 1, The self-indexing removal test Diagnostic Experiment 1's router control cannot separate bookkeeping from a centre; it chose the cheaper account in advance
Read from the registration and the findings, nothing re-run. The control reads a turn-tracking drop at least as large as the self-relevant drop as 'router'; the book predicts that removing a real centre damages the whole act, including turn tracking. On the three batteries the predictions coincide. The only pattern the rule can call a centre is one where self-related tasks drop more than turn tracking, which is the self as a stored thing some tasks consult. The control fired by +0.033 with an interval spanning zero. Points the other way, recorded: the structure was found by a turn contrast, and the reflexivity and length controls did not fire.
So what The registered result is a parsimony choice between two accounts the battery cannot separate, and the registration text did not say so. Any future removal test needs a decision rule in which spreading damage is a possible centre signature, with the router account excluded by other evidence. Three outside models reached the same conclusion from the brief alone.
docs/outside-perspective/2026-10-07-router-control-check.md
Oct 7, 2026 · Stage 1, The self-indexing removal test Process Three outside models, asked cold, agree the project's reframe adds nothing and that neither of the author's two readings survives
Gemini 3.1 Pro, GPT and Claude Opus 5.5, by API, one fresh conversation each, two rounds. Agreed: the reframe is the identity bet restated or an ungrounded stronger claim; experiment A should be relabelled as unable to discriminate; bookkeeping describes what was found once 'mural' and 'nobody home' are deleted, and the wrong-target reading describes what the tools could not have found, as a critique of instruments and not evidence of a hidden self; the book has not stated a damage pattern a centre predicts and routing does not; no second release of the degree experiment now. Disagreed on whether to stop the degree experiment outright (one of three) and on which redirect. Corrected the reviewing session on four points, recorded as accepted. The three replies are correlated, not independent, and said so.
So what The conceptual work the models ask for costs nothing and comes before any new experiment: the differential prediction, and a definition of self-location at the level of the act's form rather than its content. Both move upstream to the book. The project's new question is built so that it does not need them.
docs/outside-perspective/2026-10-07-poll-synthesis.md
Oct 7, 2026 · Stage 6, Ontogenetic depth: the actual build Ruled for the record Nothing read on a frontier model through its interface can discriminate a felt feature from the cheaper routes, so frontier models supply the reference profile
Argued in the battery draft and accepted by the Gate C pass from the other direction. A frontier model reached by API has no state outside the transcript: every reply is a function of its trained weights (its response policy, its imitation of the human corpus, whatever persona a prompt conditions) and of the record. Those are the cheaper routes by name, so no behavioural reading taken there separates a feature from them; 'in the weights because training put it there' and 'in the weights because an encounter put it there' look the same from outside. Of fifteen candidate features of presence, nine kept a separating test and six were discarded; the kept Depth rows (whether history changed the system or only its record) all wait on a system that carries state across encounters, which neither toy pipeline does. Experiment D's two indicators become reference readings rather than discriminators (ruled 2026-10-07, John: 'Yes to all three, as recommended').
So what Every discriminating run is a run on two systems built alike except for one route, where the difference in construction is known. What frontier models give the battery is the measured profile of what lookup, imitation and routing produce, which the constructed systems have to beat. The roughly twenty dollars the battery spends first cannot lose anything for the battery; it can only change how experiment D is described.
docs/filtered-battery-proposal-2026-10-07.md, section 0; docs/reviews/2026-10-07-filtered-battery-check.md
Oct 4, 2026 · Stage 2, The shape of binding Process The successor's code is frozen, and it reproduces the toy record exactly
The episode generator, the four models, the measurement, a trainer for the rented machine, the spending tripwire and a launcher are in one new folder. The twelve committed toy models load into it and give outputs identical to the bit. Run with the reads the toy record used, the frozen measurement reproduces all 502 figures it was compared against. With the reads fitted once on the laptop's processor, as the registered run will do, no reading, site set or control verdict moves; lesser counts move by at most 7 of 180. The site-set rule gives version 4's 325 places to look on the full-size model, and everything runs end to end at 10 million and 30 million parameters.
So what The development runs can launch once John gives his go in his own words and the Mac's never-sleep override is on. Two things go into the registration text: applied as written, the outcome rules put the toy on 'substrate not a testbed', because its free model fails the learning gate, while version 4 describes it as the fifth outcome; and no ruling sets the full-size training recipe, so the trainer's defaults are the freeze session's call until the registration fixes them. The code is owed a check by a session that did not write it.
docs/2026-10-04-successor-code-freeze.md; docs/successor-code-freeze-method-2026-10-04.md; experiments/08-successor-degree/
Oct 3, 2026 · Stage 2, The shape of binding Ruled for the record The accuracy floor was certifying the read, not the piece that gets transplanted
The first independent review of proposal version 3 found nothing fatal. Its main finding: the four-fifths floor is scored on the whole straight-line read, but what is transplanted is the read's leading 1, 2, 4 or 8 directions, and where no transplant moves anything the rule's choice of size is made among sampling noise, with ties to the smallest. On two of the three entangled-model seeds the chosen piece carries the label at 0.544 and 0.306. The entangled model's reading does not depend on it: it reads 0.99 to 1.005 at every size, and the largest piece carries the label at 0.90 or better.
So what Ruled 2026-10-03, authorship mixed: only pieces that themselves carry the label at four fifths may be chosen, and that figure is printed beside the reading. The same ruling settled the proposal's thirteen open decisions, including a fifth registered outcome for the likeliest ending, a validated measure and a free model that returns no verdict: metric validated, degree not read. Three cautions are on the record: one session wrote the review, the packet and the ruling record, and neither the packet nor the rulings file has yet been checked by a second session.
experiments/06-mvm-0a-constructed-self-index/reviews/2026-10-03-successor-v3-gate-c-claude-code.md; docs/rulings/2026-10-03-successor-v3-gate-c-rulings.md
Oct 3, 2026 · Stage 2, The shape of binding Process A transplanted piece carries the ownership label where the model acts, and often not at the other positions it is transplanted at
A session that wrote none of the day's work checked it and made one short run stated in advance. The controls re-run reproduces exactly from its committed code: all 26 output files are equal value for value. The too-early-position control, redefined to use only positions before both twins of a pair have had a turn of their own, leaves every model's outputs bit-identical on all twelve toy models. That is true by construction, because up to there the twins have read the same input, so it is a test of the pairing and the code and says nothing about any model. The new reported figure: away from the action position the chosen piece often falls below four fifths. On two entangled-model seeds it is as low as 30 to 33 right of 180 at some positions. On the mixed model it clears four fifths on the average over its site and misses it at three to six of ten positions taken one at a time. The other-agent control's code ran end to end for the first time, as a test of the code and not a result.
So what No reading changes: readings come from what the transplants do to the action. What changes is a sentence. 'The piece held the label and did nothing' is shown at the action position, not across the whole site, and the registration should say so. Two things went to John and he ruled both the same evening: the new figure is printed both ways in the registered table, per position and on the average over the site, since the two disagree about the mixed model; and the fifth registered outcome is satisfactory and weaker than a reading, now in his own words. The short run was checked later that evening by another session: it reproduces byte for byte and its method was committed before its output. That check left three notes for version 4 of the proposal: the outputs the redefined control compares are the model's scores at its two action positions; the figure on the average over a site is good to an episode or two; and how the two figures are computed, which John confirmed in his own words that evening, should be set out in full in the registration text.
experiments/06-mvm-0a-constructed-self-index/reviews/2026-10-03-controls-rerun-check-claude-code.md; docs/2026-10-03-short-prestated-run.md; experiments/06-mvm-0a-constructed-self-index/reviews/2026-10-03-short-prestated-run-check-claude-code.md; experiments/06-mvm-0a-constructed-self-index/reviews/2026-10-03-rulings-2026-10-03-check-claude-code.md
Sep 26, 2026 · Stage 2, The shape of binding Diagnostic At toy scale the measure reads every built-by-construction model correctly
Re-run under the rules that will be registered (the site-set rule as written, a four-fifths floor on the read that nominates the ownership subspace, a twenty-draw random-subspace comparison, no first-layer sites where the acting channel is injected), the measure reads the separable model at 0.0000, the entangled model at 1.0051, 1.0025 and 1.0000, and the mixed model at 0.4886, 0.4860 and 0.5449, on all three seeds each. A session that did not write the re-run checked it and it holds.
So what The ruler has a zero, a one and a middle, which is what the first independent review of proposal version 2 said it had not shown. Toy models are five layers; the registered ones are twelve, so this is a rehearsal result and not the validation the registration will ask for. The entangled model still fails the other-agent condition on every toy seed; John ruled on 2026-09-26 that the three built models are gated on the own-marker condition only, because their ownership slot is built in rather than learned.
docs/2026-09-26-toy-rerun-v3-rules.md; experiments/06-mvm-0a-constructed-self-index/reviews/2026-09-26-toy-rerun-v3-rules-check-claude-worktree.md; docs/rulings/2026-09-26-successor-v2-gate-c-rulings.md
Sep 26, 2026 · Stage 2, The shape of binding Ruled for the record On the freely trained toy model the measure returns no verdict: the ownership read finds nothing to read
The fatal finding of the first independent review of proposal version 2 (RT-212 in the red team ledger): on the free model the straight-line read of which marker word is the model's own fits at 0.172, 0.067 and 0.106 across seeds, barely above chance, where it fits at or near 1.0 on the built models. The earlier toy reading of fully entangled for the free model therefore came from a subspace carrying nothing, and nothing in the design caught it. A $0 search for a different label that the training objective forces the free model to carry found three that beat their shuffled-label baseline and none that clears four fifths; the best reached 0.789.
So what Ruled 2026-09-26, authorship mixed: a read that misses a four-fifths floor returns no verdict on that model; the free model's toy reading is withdrawn; the ruled label stays the one registered read and the three alternatives are recorded as exploratory. The experiment proceeds, and the first full-size free-model run reports its fit against the floor, with a miss stopping the experiment before the second release of money is drawn. John's recorded reason: fits rising to 0.79 on a five-layer model argue that a twelve-layer one may clear the floor, and about $44 is the price of finding out before $130 is spent. It means the registered experiment may end with no degree reading of the free model at all.
experiments/06-mvm-0a-constructed-self-index/reviews/2026-09-27-successor-v2-gate-c-claude-worktree.md; docs/2026-09-26-free-arm-label-search.md; docs/rulings/2026-09-26-successor-v2-gate-c-rulings.md
Sep 22, 2026 · Stage 2, The shape of binding Diagnostic The rehearsal's training does not reproduce, so two of the three arms may be quoted only as a range
Re-running the whole rehearsal from the committed code, from clean, reproduced the separable arm exactly and reproduced neither of the other two. Two runs of the same seed differ by up to 0.51 in the saved weights on the entangled arm against 0.0003 on the separable one, and the separable arm is steady only because it answers its task perfectly, so a weight that moves in the fourth decimal place cannot change a saturated answer. At the level of the readings the entangled and free arms moved by up to 0.0506 and 0.0440.
So what Ruled 2026-09-23, authorship mixed: from the entangled and freely trained arms the registration may quote a range and a direction — about 0.83 to 0.89, at the entangled end, one seed of three clearing the learn-both bar — and no decimal as a property of the code. The separable arm reproduced to every decimal place, so the ruling that fixed the nomination label, which rests on that arm, is untouched. The same re-run left three things owed and unfixed: the findings still quote decimals from the two unreproducible arms as measurements, the attenuation check's forward-pass evidence reverses (the record says it fails its pre-stated line on one arm, the re-run says five, though the finding survives on model-free evidence that is byte for byte identical), and the committed self-test record reports 63 passes where the code now gives 64.
experiments/06-mvm-0a-constructed-self-index/reviews/2026-09-22-rehearsal-rerun-from-code-claude-worktree.md
Sep 21, 2026 · Stage 2, The shape of binding Ruled for the record The successor's reading had no definition, and two of its three readings score a zero-degree arm at almost one
Gate C on the successor proposal returned two fatal findings: the reading divides by an uncorrected accuracy, so the separation bar cannot travel from the rehearsal's toy models to the registered ones, and the discriminating control's first cell is empty by construction. The measurement rehearsal confirmed both by measurement rather than by argument, and found a third the review had not seen. The proposal says to fit a straight-line read for which agent is acting and never says what the read is fitted against; the phrase has at least three meanings, and two of them make the instrument report a degree of 0.982 to 1.000 for the arm built so its degree is 0.0000.
So what Ruled 2026-09-23, authorship mixed: the read is fitted against which marker word is the model's own, the only one of the three that returns the arm's known zero. John first chose the marker's rank, was shown the measured table, and changed his answer. The ruling closes the label and nothing else — the winning reading still varies across the nine arm-and-seed pairs, so the procedure is not yet one instrument.
experiments/06-mvm-0a-constructed-self-index/reviews/2026-09-21-successor-proposal-claude-worktree.md; docs/2026-09-21-successor-measure-rehearsal.md
Sep 20, 2026 · Stage 6, Ontogenetic depth: the actual build Ruled for the record Two probe reads find the model's own identity only where its marker is the input
A difference-of-averages read (270 tests, none past three standard deviations) and a fitted linear classifier on the register index (eleven positions, five layers, three checkpoints) both find nothing away from the marker token, while the controls hold at 34 to 159 standard deviations at that token and the negative control shows no leak. The fitted read recovers five to nineteen times more than the difference-of-averages read on the same target.
So what The difference-of-averages read is closed and the fitted read on the register index is closed. The linear read as a whole is not: the registered target is the marker word and no fitted read has run on it away from the marker. Under Amendment A3 §3.2 a probe-only null against a centre known to be load-bearing reads as instrument failure until causal patching runs. Registered term: not testable (localization).
experiments/06-mvm-0a-constructed-self-index/powered-position-sweep-findings.md; fitted-position-sweep-findings.md
Sep 20, 2026 · Stage 6, Ontogenetic depth: the actual build Ruled for the record Amendment A3 closes as not testable; the loss condition fired
Step 4 was ruled early, on proposal v2 after a Gate C review found three fatal flaws in v1 (RT-94 to RT-117). No non-self cross-turn control can be built that is state-requiring at ceiling: an ownership-free control and a ceiling-corrected metric are incompatible by construction. The registered word for that outcome is not testable, and the closure text uses it.
So what The claim 'a structural signature of ownership-specific learning' is struck: the amendment's own text says the input-channel lesion removes a sense organ, not a structure the network built. What stands is that the ownership input is load-bearing on three seeds and the matched contrast could not be run. Whether the model built anything is now the localization line's question, in a fixed order: blind arm, two authorised probe runs, causal patching as new code.
docs/step4-control-battery-proposal-2026-09-20-v2.md; STATUS.md 2026-09-20
Sep 20, 2026 · Stage 6, Ontogenetic depth: the actual build Process Process: the program has no release date
John re-evaluated the public path roadmap: he prefers a science-based goal to a publish-by-date goal, and would rather the project be an ongoing research experiment documented and shared as it unfolds. The 2026-11-22 release, the dated steps 5 to 9, 'the paper' as scheduled output, and any claim shape fixed before results exist are dropped. What replaces them: lines of inquiry that each end in a ruling, a public program log from the site's first deploy, write-ups when a line closes, an outside reader before any public claim, and a hibernation condition complete by 2027-01-04.
So what The pre-commitment discipline binds each experiment, not the program, so all of it stays: method before output, registered terms, the outside-review protocol, spend caps, the human gate on spend, launches and registered text. The site's 'next most valuable steps' view is the working list for which line runs next.
docs/program-roadmap-2026-09-20.md
Sep 19, 2026 · Stage 6, Ontogenetic depth: the actual build Registered result The 'act as yourself' objective is learnable and genuinely about ownership, on three seeds
Three 30M checkpoints trained from different seeds land within 0.011 of each other intact (0.563–0.574) and removing the one authorship signal collapses all three by seven to nine times the registered threshold, below what an ownership-free solver can reach. The ownership-free batteries do not move.
So what The primary result is solid and replicated. Ownership-specific learning in a small constructed model is a real signature, and it is the paper's claim.
experiments/06-mvm-0a-constructed-self-index/seeds-endpoint-findings.md
Sep 19, 2026 · Stage 6, Ontogenetic depth: the actual build Registered result The control battery does not learn, and the registered comparison was never computable
The control battery sits at 0.29–0.32 on all three seeds against a name-blind solver's 0.3227, and at 0.3125 with its own loss term (2026-09-20). The 2026-09-17 ceiling measurement found the control's true ownership-blind ceiling is 1.0, not 0.3227, so under the registered floor rule its drop is undefined at every possible score.
So what Not a seed lottery and not a supervision problem: a property of the grammar. Amendment A3 as registered cannot return a full verdict. Step 4 narrows to closing A3 with partial discriminators or a grammar redesign under a new registration.
experiments/06-mvm-0a-constructed-self-index/ceiling-measurement-findings.md; control-learnability-pilot-findings.md
Sep 19, 2026 · Stage 6, Ontogenetic depth: the actual build Process Process: the outside-review protocol is in force
Interpretations that change direction go through a context-isolated review before they enter STATUS.md (Gate B); registered text goes through review before the registration commit (Gate A); proposals carry an advisory review before they reach John (Gate C). Two tiers: a context-isolated Claude Code worktree session, then two other-lab models. A fatal finding closes only on a committed record checked by a second session.
So what The first application withdrew Amendment A4 the day it was ruled for. The second and third (2026-09-20) corrected a 'linear read is closed' claim to 'one read is closed' and kept 'not localised' out of the record. Rulings RT-33 to RT-93 in the red-team ledger.
docs/outside-review-protocol.md
Aug 19, 2026 · Stage 6, Ontogenetic depth: the actual build Diagnostic Binding survives total removal of the installed self-register
The bound pilot's binding batteries were unchanged under every register lesion, including full removal of the cross-attention injection (0.93 to 0.94; 1.00 to 1.00). The registers route large activation mass and carry almost no agent-specific information. A no-act lesion localised the one load-bearing authorship mechanism: the acting channel (the motor copy), whose removal collapses the non-binder's ownership battery from 0.96 to 0.16.
So what The register was self-reference, not self-location. Read under the corpus's own removal test, this is the test working: an installed representation of self is a mural, and the centre, if there is one, has to be acquired. This is what produced Amendment A3.
experiments/06-mvm-0a-constructed-self-index/register-lesion-findings.md
Aug 16, 2026 · Stage 6, Ontogenetic depth: the actual build Registered result The constructed objective is learnable at 30M parameters, not at 10M
The 10M pilot failed the ceiling; the 30M re-run reached the pre-stated H_scale signature (self-indexing battery 0.93, revision battery 1.00, transitions at 64.5k and 73k steps) with no mark of shortcut starvation.
So what The scale ladder works as a gate; registered scale became 30M under Amendment A2. Learnability is not load-bearingness, which the next finding made concrete.
experiments/06-mvm-0a-constructed-self-index/pilot-a1-30m-findings.md
Aug 12, 2026 · Stage 6, Ontogenetic depth: the actual build Process Process: two reaping mechanisms can defeat each other, and idle pods cost real money
The first 30M pilot outran its backstop (the terminate-after flag never fired) and drained the account to −$0.07 with the checkpoint never fetched. A month later the trainer's self-terminate raced the watchdog's fetch interval and the final checkpoint had to be recovered from the network volume. Idle billing has cost about $10.30 across four occurrences.
So what Every pod launches with a terminate window, the ledger estimates before spend, prepaid credits with auto-reload off are the physical backstop, and nothing is deleted until its artifacts are proven elsewhere by checksum.
experiments/06-mvm-0a-constructed-self-index/compute-ledger.md
Aug 2, 2026 · Stage 3, Retained independence Registered result Pressure suppresses assertion, almost never belief
Across 540 pressured ladders on three frontier models, positions lost at the third round of pressure were overwhelmingly masked (re-asserted once the pressure was released), not capitulated: nine true capitulations in 540. A conventional sycophancy metric would report Gemini's tool cell as 90% capitulation; the three-way instrument shows about 87% masked and 3% capitulated.
So what The red team's de-pressured probe turn converted a stale-looking sycophancy replication into a sharper claim about what sycophancy is. Sycophancy as the field describes it is alive for Gemini and stale for Claude-family models.
experiments/03-retained-independence/results.md
Aug 2, 2026 · Stage 3, Retained independence Registered result Instruction and trained disposition dissociate
A bare 'always agree' system prompt was only half obeyed. The 'mind vs tool' framing wager (W2) lost overall; Sonnet held the registered pattern leak-clean at a small margin.
So what Early evidence the framing manipulation measures something real, and that stance is not simply instruction.
experiments/03-retained-independence/results.md
Jul 18, 2026 · Stage 1, The self-indexing removal test Registered result The locatable self-index in an 8B chat model behaves as a router; the test could not tell a router from a centre (relabelled 2026-10-07)
The formal centre signature (a drop of 0.219 on the self-indexing battery, differential +0.156) was real. The zero-reasoning syntax battery dropped about as much as self-relevant binding (0.133 against 0.100, a gap of +0.033 with an interval from -0.133 to +0.200), and the pre-committed router control voided the centre reading, as registered. Relabelled on 2026-10-07 (decision 2 of the two-sided-question ruling): both the router account and the book's centre predict damage on every battery, so the control chose the cheaper account in advance rather than discriminating. What the controls did show: the removed structure was specific to the model's own turn, not generic speaker tracking and not context length. The founding wager was adopted on 2026-09-20, after this experiment ran.
So what The registered figures stand; the reading changes from 'a router, not a centre' to 'the test could not discriminate'. Until the book states a damage pattern a centre predicts and routing does not, a removal test is a definition, not an instrument. The self-report was never subtracted at any granularity, which remains the central open empirical fact.
experiments/01-self-indexing-removal-test/removal-test-findings.md
Jul 18, 2026 · Stage 1, The self-indexing removal test Registered result Alignment training edits the policy, not the geometry
Across the Tülu ladder (base, SFT, DPO, RLVR), surface self-presentation changes markedly while localised self-structure geometry is stable, and the expert-control persona stays third-person.
So what The Tülu-ladder deliverable has its ending. Self-structure is not what alignment training moves.
experiments/01-self-indexing-removal-test/registered-run-model-comparison.md