Story
The story so far
One row per entry in the handoff record (STATUS.md), newest first. The full prose of each entry, with its numbers and caveats, lives in the repo; this is the shape of the road.
October 2026
- Oct 8, 2026Process
The refounding is on the main line; the registration text and the battery draft wait on checks; the site is brought up to date
Pull request 123 merged at John's word, carrying the poll, the router-control check, the refounding proposal with its reviews and eleven rulings, and the battery draft. Two drafts written by fresh sessions on 2026-10-07 wait on their checks: version 5 of the registration text (39 ruled changes written in; twelve open items, the largest the unruled bar for 'the repaired route holds', which as coded fails the toy's own built models; fifty citations to fifteen unmerged branches) and version 3 of the battery draft. A session read the whole record and gave John its assessment: the diagnosis behind the refounding is sound; everything that can discriminate depends on a construction line that does not exist yet; the likeliest ending for experiment C this year is an early stop recorded as instrument research. He agreed to three recommendations: rule the verification bar first, restate the December result under the refounding in one page, and bring the site current so a Sentient Horizons post can be written from it. The site's data, the weekend re-plan and the plain-language page were rewritten the same day. Later the same day John ran the first deploy from his machine and the site went live at its subdomain; the two Cloudflare secrets were set the same afternoon and the deploy workflow ran green, so every merge now publishes the site on its own. The project name's .com, .org and .net are registered to someone else until at least April 2027, so the subdomain stays. Nothing spent.
- Oct 7, 2026Ruling
The project is refounded on the two-sided question; three outside models polled; experiment A relabelled, C kept with a conditional second release, D promoted
A brief for outside models was checked against the record and revised, then put by API to Gemini 3.1 Pro, GPT and Claude Opus 5.5 in two rounds, every reply saved word for word. All three agreed the reframe adds nothing beyond the identity bet, that experiment 1's router control cannot tell bookkeeping from a centre, that neither of John's two readings survives as written, and that no second release of the degree experiment should be funded now. A check of the registration text reached the same conclusion about the control. John then named what he perceives from, the axes rather than the floor, and ruled all eight decisions of the refounding proposal before the Gate C pass, at his own instruction: the two-sided question is the measurement target; A is relabelled; C keeps its second release conditional (struck by the first ruling, restored after the Gate C pass found the narrowing left no reachable outcome); D is promoted to the seed of a filtered battery; the kill dates stand and no new dates are set; the table and battery are next; the human study is deferred; the book changes go upstream. Seven commits on one unmerged branch; two checks running; $0 rented.
- Oct 4, 2026Build
The successor's code is frozen and tested at $0; the go packet for the development runs is ready
A new folder holds the episode generator, the four models, the measurement, a trainer for the rented machine, the spending tripwire and a launcher. The frozen code reproduces all 502 toy figures compared, gives the 325 places to look on the full-size model, and runs end to end at 10 and 30 million parameters. Two things go to the registration text: the outcome rules as written put the toy on 'substrate not a testbed', and no ruling sets the full-size training recipe. The four development runs' go packet and ledger rows are written; nothing launched, nothing spent. John also ruled that no step is scheduled on a future date any more.
- Oct 3, 2026Ruling
The seven questions were ruled twice in one evening; John settles which record stands
Two sessions each put version 4's seven questions to John within minutes and each recorded his agreement. The records differed on four points. He ruled that the fuller one stands: library versions pinned, the registration says what the order of the piece rule can miss, the sampling band printed at the floor, and the separation between the built models taken as the lowest entangled reading minus the highest separable one. Both records carry a dated note and version 4 is edited to match. Nothing spent.
- Oct 3, 2026Ruling
John rules version 4's seven questions; one short run is owed before the registration review
Ruled as suggested, in the words 'Agreed on all'. The accuracy floor is on the transplanted piece only. The piece rule is applied after the layers are chosen. The registered fit is computed on the laptop's processor. The full-size measurement uses the toy's numbers of episodes. The other-agent control is described beside twenty random pieces. When an arm's three seeds disagree, two of three decide. The ordinary competing solver is run through the measurement as now registered before the review opens. Version 4 carries all seven; the record of the ruling and version 4 are each owed a check. Nothing spent.
- Oct 3, 2026Process
Version 4 of the proposal is drafted; three things stand between it and the registration review
Version 3 with every ruling of 2026-10-03 written in, as a new file; version 3 is left unedited. The accuracy floor sits on the piece that is transplanted; the other-agent control has no pass line and is said never to have run at toy scale; the too-early-position control holds and is described as a test of the pairing and the code; the fifth outcome is in. Toy figures come from the controls re-run. The 325 site sets for the full-size model are printed. Not ready for the review: the check of the short run and the evening ruling is filed but not merged, version 4 needs its own check, and seven questions go to John. Nothing spent.
- Oct 3, 2026Process
The short run stated in advance, and the record of the evening ruling, are checked and hold
A session that wrote neither checked both from committed files. The run's method and code were committed and pushed before any output and the script is unchanged; run again, the output files are byte-identical; separately written code agrees, and shows the redefined control is not an empty test. Three notes for version 4 of the proposal, none changing a figure: the record cannot show the script was never run before the method was committed; the outputs compared are the model's scores at its two action positions; the figure on the average over a site is good to an episode or two. The ruling record matches John's words except that it set down how the two figures are computed, which his words accepted only by implication; asked, he confirmed it that evening. Laptop only, $0.
- Oct 3, 2026Process
The day's work is checked by a session that wrote none of it, and the short run stated in advance is done
The controls re-run reproduces exactly. The diagnosis behind the redefined too-early-position control holds, so that ruling stands. The short run: the redefined control holds on all twelve toy models, by construction; a transplanted piece often does not carry the label at the other positions of its site, on the mixed model too; the other-agent control's code ran end to end as a test of the code. Neither ruling packet has a wrong number. Two small things went to John, who ruled both that evening: the fifth outcome is satisfactory and weaker than a reading, and the new figure is printed both ways. Nothing blocks version 4 of the proposal. Laptop only, $0.
- Oct 3, 2026Ruling
Three rulings after the controls re-run: one control loses its pass line, one is redefined and can veto again, one new figure is reported
The other-agent control is kept as a description with no pre-stated pass line, since no toy model can exercise it. The too-early-position control is redefined on both twins of a pair and becomes a control that holds, reversing the morning's ruling; this rests on an after-the-fact diagnostic not yet checked by a second session. The accuracy of a piece transplanted at several positions is reported at each. One short pre-stated run and three checks are owed before the registration text. Nothing spent.
- Oct 3, 2026Result
The controls re-run: the ruled change holds on the toy, one control cannot be exercised, and another turns out to be defined wrongly
Method committed before output; laptop only; $0. Readings under the registered rules: separable 0.0000, entangled 1.0051, 0.9926 and 0.9974, mixed 0.4886, 0.4860 and 0.5449, free no verdict on every seed. The other-agent control has no figure on any model because the named agent's read misses the floor on the one model that learned the condition. The too-early-position control came back above nothing on six of twelve models; an after-the-fact diagnostic shows it returns exactly nothing once its positions are taken before both twins' first own turns. Both go to John. A second session's check is owed.
- Oct 3, 2026Ruling
Proposal version 3 is merged, reviewed with nothing fatal, and ruled: the floor now applies to the piece that is transplanted, and a fifth outcome is registered
The first independent review found two serious findings and five minor ones (RT-230 to RT-236 in the red team ledger's sequence). John ruled all seven and the proposal's thirteen open decisions in one sitting. The controls re-run at $0 comes before the registration review. A free model that returns no verdict after the built models separate is reported as metric validated, degree not read. Nothing spent.
- Oct 3, 2026Process
The record catches up after a week's gap: Weekend 1 did its four jobs, proposal version 3 is written and unreviewed, and the registration slips one weekend
No work between 2026-09-26 and 2026-10-03. This entry is the skipped Sunday handoff and the skipped weekly re-plan, done late. Proposal version 3 and John's rulings on two of its decisions are open pull requests (71 and 72). The registration will not be committed on 2026-10-03/04 as the weekend roadmap planned; the realistic date is 2026-10-10/11, a week inside the first kill date, and the first reserve weekend absorbs the slip. About $228.15 of $450; nothing spent since 2026-09-25.
September 2026
- Sep 26, 2026Ruling
Proposal version 2's first independent review is fatal and John rules it the same day; at toy scale the measure reads the three built models and returns no verdict on the free one
One fatal finding and four serious ones (RT-212 to RT-229 in the red team ledger). Ruled: a four-fifths floor on the read that nominates the ownership subspace, the free model's toy reading withdrawn, the built models gated on the own-marker condition only, no first-layer sites where the acting channel is injected, the random-subspace comparison reported rather than used as a gate. The toy re-run under those rules reads the separable model at 0, the entangled one at 1.0 and the mixed one at about 0.5 on every seed. A search for another label on the free model tops out at 0.789, and John ruled the experiment proceeds with the first full-size free-model run as the stop point. The other-agent condition failed a fourth repair attempt. Thirty trained toy models committed. Version 3 written. $0.
- Sep 25, 2026Spend
The small rented test run, in two attempts: training speed measured, the registered runs re-priced at $161.90
The first attempt hung while starting the machine's own shutdown watcher and was stopped at about $0.50 with nothing measured; the launcher was fixed and checked. The second attempt ($0.0525) measured 13.08, 13.52 and 12.53 milliseconds per step on the separable, entangled and free models. The half of the shutdown test that runs on the rented machine was not exercised. The same day John ruled the nine-page queue of decisions blocking the registration, including the ceiling raised from $400 to $450.
- Sep 25, 2026Ruling
Amendment A3 is closed on the record: the closure block is appended to its registration file as the registration commit
Version 5 of the closure text, after two outside reviews, John's rulings, a tier 1 check by a session that did not write it and a re-check of the fixes, was appended verbatim to amendment-a3.md with the date filled in. Outcome as the block states it: not testable, because the registered comparison was undefined for every possible model; separately, removing the ownership-input channel reduced primary-battery accuracy on all three trained seeds. The same pull request lands the two annotations of 2026-09-25, ledger rows RT-204 to RT-211 with John's rulings, a note on what RT-172 to RT-203 refer to, and the compute ledger's record of the ceiling raised from $400 to $450. Nothing spent; the ledger's last row is still 2026-09-21.
- Sep 24, 2026Process
The record catches up: the stranded handoff is landed, the weekend roadmap is mirrored, and the launcher refuses to rent while the laptop can sleep
The 2026-09-22/23 handoff, drafted in a folder outside the repository, landed as pull request 29. A weekend-by-weekend schedule to 2026-12-21 was laid over the December-result roadmap and mirrored in TimeAssembler, where two steps were closed as already done. The launcher now refuses to rent a machine while the laptop could fall asleep (pull request 30). STATUS.md now says in one sentence how the programme roadmap and the December-result roadmap fit together. Nothing spent.
- Sep 24, 2026Process
Three days of checking and no money spent: the successor's reading gets a label and a limit on what may be quoted from it, and both rulings land
The two fatal findings from the review attached to the proposal were confirmed by measurement; the nomination step was found to have no definition at all; John ruled the label is which marker word is the model's own, and ruled a range and a direction only for the two arms whose training does not reproduce. Both rulings and their notes reached the main line on 2026-09-24 as pull request 28. Programme $227.6 of $400; Amendment A3 $46.2 of $100; nothing spent since 2026-09-21.
- Sep 21, 2026Null
The two follow-up localization reads are in and ruled: the exclusion confound is excluded, nothing is localized
Twenty-three findings ruled as drafted. One cell of 135 crossed the family-adjusted bar and is not carried as a clearance. The registered term for where the line stands is still not testable (localization). Work is measured in task time, not calendar time.
- Sep 20, 2026Ruling
December-result roadmap approved: the degree metric is the successor experiment, registered in 2026, one result or a named schedule failure by 2026-12-21
Three arms (separable by construction, entangled by construction, free), four registered outcomes, kill dates 2026-10-18 and 2026-11-01, successor cap $130. Standing rule: spending above a cap is proposed with the number, never worked around.
- Sep 20, 2026Process
Program roadmap approved: no release date; an open-ended, publicly logged research program with a hibernation deadline of 2027-01-04
Dated steps 5 to 9 and the paper-as-scheduled-output dropped. Lines of inquiry that each end in a ruling; the site is the public program log from its first deploy.
- Sep 20, 2026Ruling
Step 4 ruled early: Amendment A3 closes as not testable; the loss condition fired; claim scope narrowed; localization order set
On proposal v2 after a Gate C review (RT-94 to RT-117, three fatal). Option D is a successor experiment. 'Structural signature of ownership-specific learning' struck.
- Sep 20, 2026Null
Control battery converges on the name-blind solver even with its own loss term; fitted read finds nothing on the register index
Both entered STATUS.md through Gate B. The linear-read line is parked as not testable (localization).
- Sep 19, 2026Ruling
Eleven-position sweep finds nothing under a difference-of-averages read; outside review says that closes one read, not the linear read
First entry under the outside-review protocol. Rulings RT-33 to RT-51.
- Sep 19, 2026Result
Branch merged to main; the seeds replicate; the registered comparison was never computable
Primary battery collapses 7.5× to 8.6× past threshold on three seeds. Control never learned on any seed. Sensitivity rule ruled for future runs only.
- Sep 19, 2026Ruling
Amendment A4 ruled for in the morning, withdrawn the same day; outside-review protocol ruled in force
Independent red team found the control cannot fall far enough for the generic-binding cell to fire. $10 control-learnability pilot approved as the replacement spend.
- Sep 17, 2026Null
Ceiling measurement: the control battery's ownership-blind ceiling is 1.0, not 0.3227
The registered differential clause has been unsatisfiable since registration. Options B and E rejected and fenced.
- Sep 16, 2026Result
A3 pilot complete: the objective is learnable and genuinely about ownership; control never learned; $13.92
Threshold lock committed by John; seeds 1 and 2 launched; public path roadmap approved.
- Sep 15, 2026Process
Shortcut sweep finds a second fatal leak; red-team pass 3 finds the certified grammar did not test self-indexing; Gate 1 complete
Three grammar drafts map a real trade-off. Amendment A3 ratified by John, all fifteen decisions.
- Sep 14, 2026Process
Gate 0 complete: registered null calibration ran at $0; kill condition K0 does not fire
Project resumed after the stall; Claude Code agent picked up from Gate 0.
August 2026
- Aug 19, 2026Null
Register lesion: binding survives total register removal; construct problem is total
The installed register is numerically active and informationally inert. The acting channel is the one load-bearing authorship mechanism.
- Aug 16, 2026Result
30M learns: verdict H_scale; gate (iii) passes; Amendment A2 registered (scale 30M, cap $400)
Watchdog did the full fetch and deleted the pod itself; the 2026-08-12 failure mode closed.
- Aug 15, 2026Build
30M re-run prepped and launched with all three process fixes
Secure 5090 with network volume, watchdog armed, $75 top-up.
- Aug 12, 2026Spend
30M pilot lost: pod outran its backstop, balance drained to −$0.07, checkpoint never fetched
Terminate-after never fired. $97 on a single continuous run. RunPod ticket submitted.
- Aug 9, 2026Result
A1 pilot: 10M fails the ceiling; ladder climbs to 30M; gate (iii) passes on the A1 pipeline
Earlier the same day: gate run (iii) failed because rollouts carried an ownership fingerprint; the fix was red-teamed and the trilemma closed.
- Aug 8, 2026Result
Pilot result: 10M learns; the ladder stops at rung one
Then gate (iii) fails on the ownership fingerprint and the 5-seed spend is blocked.
- Aug 7, 2026Process
MVM-0a registered (v1.0 binding); corrigibility document committed; model and training loop built
All six blocking calls adjudicated. Compute cap $200.
- Aug 4, 2026Build
MVM-0a curriculum built and red-teamed; theory-to-instrument ledger written; uncertainty amendments run on experiments 1 and 3
Nature draft upgraded to v0.2 with confidence intervals and first figures.
- Aug 2, 2026Result
Stage 3 registered result: masked, not capitulated; W1 splits, W2 loses, W3 wins
Same day: fork adjudicated (MVM-0 build primary), flagship paper drafted, roadmap v2.
July 2026
- Jul 19, 2026Process
Stage 3 bank reconciled and construct-validity gate passed
Primary bank machine-verified and audited; reserve pool established.
- Jul 18, 2026Result
The registered removal test has run: router, not centre; report never subtracted; narrative not testable
Thresholds locked, spot-checks done, held-out test set authored earlier the same day.
- Jul 15, 2026Null
SAE-feature pilot: loss condition fires; both registered escalations exhausted
Pre-lock bench bundle complete; rank-k ablation exhausted.
- Jul 13, 2026Build
Substrate migration complete: all gates re-verify on Tulu-3-8B-SFT; cloud bench replaces the Mac mini
Dress rehearsal end to end the day before: rank-1 ablation too weak.
- Jul 12, 2026Result
Cross-patching: the two self-structures are functionally separable; narrative causally confirmed
RT-09 and RT-10 resolved on the sandbox; both deflations defeated.
- Jul 1, 2026Process
ROADMAP written; RT-09 (generic-speaker reflexivity control) registered; explainer written
Goal hierarchy and stage gates set down from a step-back review.
June 2026
- Jun 23, 2026Process
Design amended by red-team review RT-01 to RT-04; both self-structures localised
Stage 1 in progress; cross-patching and battery work next.