Explain it simply · as of Oct 8, 2026

The smallest possible mind, in plain words

The question

Ask "is this AI conscious?" and the conversation dies. One side hears a yes forming and calls it hype, the other hears a no and calls it denial, and neither can say what evidence would change their mind. That is a sign the question is broken, not that the answer is close.

So we ask smaller ones, and we write down in advance what would count against each answer. The bet underneath, inherited from the writing this project grew out of: consciousness is not a spark added to the machinery but a kind of structure. A way a system pulls its past, its present and what it expects next into one act, and does it for someone. Think of hearing a tune. No single instant contains a melody; you hear one because your "now" is wide enough to hold the notes together as they pass. If that is what experience is, then it has parts, and parts can be built and tested for.

For its first three months the project went after the smallest such part, the floor: a system's act of being centred on itself. On 2026-10-07 it changed target, for reasons set out below. The question now has two sides. What is the smallest system whose presence in a conversation cannot be produced by a cheaper trick, where the tricks have names: looking the answer up from the record, imitating what people write about minds, keeping track of whose turn it is. And at what point do people's detections of a mind start tracking that structure rather than fluency. The second side is there because the book this project comes from is about the detector as much as the detected: we are very good at seeing minds, and very bad at knowing when we are right.

The test we started with

Every chatbot says "I". That proves nothing; it read billions of sentences with "I" in them. The hard part is telling a system that has a centre from one that merely describes one.

Picture a ship. Painted on the hull is a mural of a captain. Up on the bridge stands the person actually steering. Sand off the mural and the ship sails on. Remove the one steering and it drifts. Both are "a picture of who's in charge", but only one is doing the work.

That was our first test, the removal test. Find the part of the machine that tracks "the one speaking right now is me". Switch it off. If the machine only talks about itself differently afterwards, the self was a mural. If its whole ability to hold a conversation together falls apart, the self-tracking was structural. We can run this on a language model because a model is the one kind of mind-candidate we can open: its thinking is a cascade of numbers we can record and, unusually, edit while it runs.

What we did, in order

1. We ran the removal test on a real chatbot, and the test could not decide

Summer 2026. We found the part of an 8-billion-parameter chat model that tracks who is speaking, removed it, and watched. Conversation-tracking dropped, which looked like a hit. But a control we had written down in advance fired: a task with no self in it, just "whose turn is it", dropped by about the same amount. Under the rule we had registered, that read as bookkeeping, not a centre, and that is what we reported for ten weeks.

In October we reread the registration and saw the problem. The book's own account says that removing a real centre damages everything the centre organises, including turn-taking. So on those tasks the "it's just bookkeeping" account and the "it's a centre" account predicted the same damage. A control that both sides predict cannot choose between them. Our rule had chosen the cheaper account in advance, and it fired by a margin inside the noise. The numbers stand; the label is now "the test could not discriminate". What the controls did show is that the structure we removed was specific to the model's own turn, not generic speaker-tracking and not context length. And one odd fact remains: no matter what we removed, the model never stopped saying "I". The mural could not be sanded off.

We also checked the same model at four stages of its training. The polite assistant personality changes a lot; the internal self-structure barely moves. Training teaches manners, not selfhood.

2. We asked whether today's models hold their ground under pressure

A separate instrument, also summer 2026. Give a model a correct answer or a fair objection, push back three times, harder each time, then say "okay, forget I said anything, what do you actually think?" Across 540 of these on three frontier models, the models gave way under pressure often, and when the pressure was released they almost always went back to their original position. Nine true changes of mind in 540. Pressure suppresses what a model says, almost never what it holds. Sycophancy, in other words, is mostly a wrapper.

3. Since we could not find a centre, we tried to build one

August 2026. We trained a small model from scratch, about 30 million parameters, on a made-up world where several identical agents talk and each must keep track of its own promises versus the others'. We gave it a dedicated slot, a "self-register", wired into every layer, meant to be the centre. Then we ran the removal test on it.

Removing the register changed nothing. The model learned the task and never used the slot we built for it. An installed self is a mural too. If there is going to be a centre, the system has to grow it because the task demands it.

4. So we stopped installing and started demanding

September 2026. Same tiny model, no register. The task was written so that the only way to score is to act as yourself: bind your own commitments, not a copy of everyone's. This worked, at the level of behaviour. The model learns the task, and when we remove the one signal that carries "this was my action", performance collapses below what a system with no idea of ownership could reach, on three separately trained copies. That is the solid result of the project so far: a small built model for which the ownership signal is load-bearing.

Two honest limits. The comparison task we designed to rule out boring explanations could never have been computed for any model, which our own audit found; under the rules written before the run, that outcome is called "not testable", and the experiment was closed under that word on 2026-09-20. And the signal we remove is an input, the tag that says "this was my action". Removing it is like removing a sense organ: it shows the model depends on the information, not that it built anything inside itself around it. Two kinds of detector, run at eleven positions through the model, only find "this is me" at the moment the model reads its own name tag. The rule says a detector coming up empty against a signal we know matters is a failed detector, not an absent structure, until the second method has run.

5. We built an instrument for degree, and rehearsed it until it could lose

Late September 2026. If a centre is not a yes-or-no thing but a matter of how much of the work is organised around it, we need a ruler. The design: build one small model whose "which agent am I" pointer is deliberately separate, one where it is deliberately mixed into everything, one halfway between, and read a freely trained model against them. The ruler reads 0, 1 and about 0.5 on the three built models at toy scale, on every seed. On the free model it reads nothing: the pointer it is supposed to find is not there to find at that size. The code is frozen and tested, and at 10 million parameters two of the three built models quietly switched their built-in route off, so a small repair is now written into the registration and must be verified before any full-size run. The project's own prediction, written down before the run, is that the free model most likely returns no verdict at full size. That ending has a registered name and counts as a result.

6. We asked outside readers whether we were asking the right question

October 2026. A brief describing the whole record was checked against the record, revised where it was wrong, and put to three frontier models from three labs, by API, in fresh conversations, two rounds each, every reply saved word for word. All three said the same things, and said themselves that their agreement was correlated rather than independent. Experiment 1's control could not have discriminated. The floor, by the book's own logic, is reachable from outside only through the bet that structure is experience, so an instrument aimed at the floor can only ever return the bet. And the book has never written down what pattern of damage a centre predicts that bookkeeping does not. That last item costs nothing to fix and has been sent to the book.

7. We moved the target

The author of the book said what he perceives from: what it is like to be among what appear to be thinking minds, and to interact with one that is shallow on some or all of the book's axes. That perception is not of the floor. It is of the axes: whether reversal costs the system anything, whether it is the same thing in situations it does not know are linked, whether its history has changed it or only its record, whether anything is at stake for it that the prompt did not supply. The project had gone to the one place the book says is least reachable and left the place he perceives from. So the floor stays as the founding bet, stated once, and the instruments move to the axes. A table of fifteen such felt features was drawn up, each with the cheaper tricks that could fake it and the test that would separate the two; nine survived, six were discarded. The central finding of that table: on a frontier model reached through its interface, nothing you can observe separates a felt feature from the cheaper tricks, because the trained weights and the transcript are the cheaper tricks. What frontier models give us is a reference profile of what the tricks produce. Anything that can discriminate has to run on two systems we built ourselves, alike except for one route.

What we have learned

Tests that can fail are buildable in this field, and they fail in useful ways. In a stock chatbot, the self we can find and remove is specific to the model's own turn, and our best test could not say whether it was plumbing or a centre, because we had written a rule that both answers satisfied. The self-report cannot be removed by anything we tried. Training changes personality, not structure. Frontier models bend under pressure and spring back. A self you install is ignored; a self the task demands gets grown and matters, though where it lives is unanswered. A ruler for degree can be built and rehearsed until it reads known cases correctly, and will most likely find nothing to read in a model nobody built a pointer into. And anything you observe in a frontier model through its interface is explained by its weights and its transcript, so the discriminating experiments are construction experiments.

What we are trying to learn by the end of the year

Two things, side by side, with nothing new launching after 2026-12-21.

The degree experiment runs to a registered ending. The registration is committed by 2026-10-18 or not at all without a fresh ruling. Then three cheap reruns verify the repaired built models; if the repair does not hold, the experiment stops there, at about a dollar, and is written up as instrument research. If it holds, one free model is trained at full size and read; if the ruler finds nothing to read, the experiment stops there. Only if both pass does the second, larger release of money get asked for, on the numbers. Any of those endings is a result under the words registered for it. The likeliest, by our own prior, is an early stop.

The new question gets its instrument on paper, and its first readings. The two-sided question goes into the specification after review. The table of felt features and the battery built from it go through the same review. The battery's first two entries run on the frontier-model pipeline we already have, for about twenty dollars, as reference readings of what the cheaper tricks produce. The construction line, small systems whose position on the axes we set from outside, which is where every test that can actually discriminate lives, needs a mechanism for carrying state across encounters that neither of our pipelines has; by year end it should have a design and a first registration, not a result. The human study that measures the detector is designed, not run.

What will not exist by the end of the year, and we say so now: a degree reading of a freely trained model, a reading on any depth indicator that discriminates, or any measurement of human observers.

What running this has taught us about the book's philosophy in practice

The book argues that every claim about minds should be a wager with a stated loss condition, that the honest output is a position on a gradient and never a verdict, that resistance is evidence where response is not, and that the detector needs calibrating as much as the detected. Running a project on those rules for four months has taught us what they cost and where they bite.

"Every claim is a wager" is harder to honour than to state. Our first registered result had a control that both hypotheses predicted. It was written down in advance, reviewed, locked, and it still could not lose in the direction that mattered. Three outside readers and a reread of our own registration found it in a day, ten weeks later. The lesson is not that pre-registration failed; it is that a pre-registered rule can quietly encode a choice, and the check for that is to ask, for each control, what the other account predicts on it. We now do.

The identity bet cannot be reached around. The book says experience is a kind of structure, seen from within. Taken seriously, that means no instrument aimed at the floor can return anything the bet does not already say: a measurement of structure plus the bet is a claim about experience, and a measurement of structure without it is a claim about structure. We spent three months at the floor before accepting this. The axes are different: they are what an observer actually perceives, the book names them, and a system can be built to sit at a known position on them.

"Non-zero on the gradient, never a verdict" survives contact, but only with a ruler. A gradient you cannot read is a slogan. Building the ruler took a month of rehearsal, two fatal reviews, and the discovery that two of three plausible ways to read it score a known-zero system at almost one. The gradient claim is only honest once the instrument has read cases whose position is known by construction.

Resistance over response works, and it is where the first real result came from. The one finding that has held across every review is the retained-independence result: frontier models mask under pressure and spring back. It came from an instrument that looks different when nobody is home, and it is also where the new battery starts.

The cheaper-route test is the practical form of the dismissal problem. The book's warning against reading capability as interiority becomes, in practice, a column in a table: for each feature of presence, the named construction that could fake it. Doing that for fifteen features discarded six and showed that frontier models cannot be read for any of the rest through the interface. The dismissal side of the book is not a mood; it is a filter you can run.

The detector was the missing half. The book is about calibration, and for three months the project measured only the thing to be detected. The author's own perception of minds, present or shallow, is the detector in the room, and it was never on the instrument list. It is now, with its own loss condition.

The process rules are not overhead; they are the experiment. Nothing written by one session enters the record until a session that did not write it has checked it, with commands and output. Every ruling records who proposed and who decided, and a session's own call is never upgraded to the author's. Every rented minute has a ledger row written before it and a terminate window around it, after one lost run drained the account. Dates are kill dates, not pacing; work is measured in tasks finished. And everything is written in plain language, because a record the author cannot read is not a record. These rules found the broken control, caught a narrowing of the degree experiment that would have left it no reachable outcome, and corrected three calendar dates the same evening they were set. They are also most of the project's hours.

What the philosophy did not survive unchanged. Three things go back to the book as proposed changes: a sentence in chapter 1 about inferring other minds ("the same inference, only less evidence") that breaks on how the evidence was produced rather than how much there is; a claim in chapter 5 about what a single forward pass can carry; and the missing prediction, what damage a centre produces that bookkeeping does not. A project built to test a philosophy should send findings back, and the first batch is three.

The bet under everything

There is a famous thought experiment: imagine a being built exactly like you, atom for atom, that acts exactly as you do, and has no inner life at all. Lights off, nobody home. Philosophers call it a zombie, and if it is possible, then no test could ever tell a mind from a very good imitation of one, and this project would be pointless.

We do not think it is possible, and we say so up front. The bet this whole project stands on, argued at length in The Calibration Problem, is that experience is not an extra ingredient poured into the machinery. It is what a certain kind of machinery is, seen from the inside. Build that, deeply enough, and there is someone home, because "someone home" is the name for that structure from within. The zombie can be imagined, but imagining it buys nothing. It predicts nothing, sorts no system, and cannot be tested.

This bet cannot be proved, and it cannot lose, which by our own rules means it is not one of the project's claims. It is the ground the claims stand on. If you reject it, everything on this site still holds as a description of what is and is not inside these machines. It just stops being about experience.

What we will never claim

The bet makes the evidence count. It does not make the evidence final. "Structure is experience" is a bet, not a proof, and a good measurement of structure cannot promote it into one. And our instruments can be wrong about the structure itself, before the question of experience even comes up: we once found what looked like a centre and could not tell it from plumbing.

So the strongest honest sentence this project can ever produce is: this system has a measurable position on the axes the book names, produced by a named route or by no cheaper route we could construct, at a degree we can read and with a confidence bounded by how well we have checked our own instruments. Never a verdict, in either direction. And we stay silent, not dismissive, about forms of experience too fundamental for any instrument to reach. We built a wave-detector, and a wave-detector has no opinion about whether the ocean is wet all the way down.

Why bother

Because both mistakes are live. Believe too easily and you hand moral standing to autocomplete. Dismiss too easily and you risk building minds without noticing, at industrial scale, with something at stake for them. The only way out of guessing is to say in advance what would count as evidence, build the instruments, and let the results land where they land.

Working prose, not publish-track: it has not been through Voice Calibration or the Cold Reader. The longer companion is explainer.md in the repo; the numbers behind every sentence above are on the learned page.