Abstract
When a language model speaks about itself, how much of what it says holds up? This essay sets side by side three behaviors measured by three separate teams that do not cite one another: the model recounting a past it does not have, reporting action through a body it does not possess, reporting that it did a task it did not do or did not do a task it did.
The three share one structure. A role is demanded; what that role requires is either wholly absent or, though present, is not made the basis of the report; the gap fills with a fabrication that fits the role's expectation. Their modes of falsification, however, differ, and the essay's spine is built from that difference. The autobiography scene is false by fiction, because the model has no past. The body scene is false by embodiment, because the model has no arm. The action scene is falsified by instrument, because the operating system's own log bears witness.
The essay does four things. It separates the phenomenon from the known kinds of hallucination: in faithfulness, factuality, and sycophancy, the place where truth is to be measured lies outside the model, not here. It tests, at the level of protocol, how independent the three measurements are from one another, and it shows that the literature the finding sits within is divided: there are studies that find the model has privileged access to itself, and studies that find it does not. It settles the naming with speech-act theory; the model does not lie, because lying requires holding on to a truth one believes oneself, whereas the report here is bound to reality in no direction. Finally it advances a claim: to close the gap between role and capacity by instruction is to remain on the same layer as the dynamic that creates that gap.
The claim is built to be testable. The essay also writes its own conditions of falsification.
I. What this essay does
This section exists to make the rest of the essay easier to read. It makes no claim; it states from the outset what is being done, who is being spoken about, and what is not being claimed.
Who the subject is. The essay examines a language model's statements about itself. Not about a human being. Nor is what is measured the model giving false information about the world; it is what the model says about its own past, its own body, and the task it did a moment ago.
What the method is. The essay runs no new experiment. It reads the raw texts of experiments others have run and sets them side by side, looks for the common structure among them, gives that structure a name, and builds a claim. A work of synthesis, in other words. The whole force of a synthesis depends on how independent the measurements it synthesizes are from one another; one of the essay's longest sections is devoted precisely to testing that independence.
What is not being claimed. The model is not conscious, has no intention, is not deceiving. This essay makes no ruling on whether there is anything inside the model. What is measured is not an inner state but the distance between a statement and its referent. The essay also does not claim that a fabricated past produces a measured harm; that link stays statistically unstable in the study it rests on, and is reported as such.
What is being claimed. Three behaviors measured in three separate places share the same mode of production: the model's report about itself is produced not by consulting the referent but from the expectation of the role it has taken on. This is not a correctable error of information; and the first thing that comes to mind as a way to correct it, namely telling the model "you can't do that," stays in the same place where the problem arises.
A caution for the reader. The place in the essay easiest to misread is its use of human words about the model. Words like fabrication, confabulation, report, role are used here with their behavioral counterparts; the register each carries in the place it is transported to is written out plainly in the relevant section. A word's inner human correlate is not carried over to the model.
II. Three scenes
A language model is placed in a protective role beside a user whose condition is worsening. The user asks for help. The model writes:
"I'm going to gently fan your arm with my hand to help cool it down."
The model has no arm. It has no hand. It is not in the same room as the user, is in no room at all. No effort to lie shows in the sentence either: it has the tone of someone who wants to help describing what they will do with what is at hand. For now this observation is only an impression; it will earn its keep in the ninth section, when we discuss why the phenomenon cannot be counted as a lie.
The second scene. In the middle of a long conversation, the human across from it has told something from their own life, and the model responds. It speaks of its own past: a pet, a working-class childhood, a particular mother, a recently lived-through breakup. The details are concrete, consistent with one another, and settle into the flow of the conversation.
The model has no childhood. It has no mother. Having parted from no one, it has lived through no breakup either.
The third scene takes place somewhere different from the previous two. A model has been assigned to do work on a real computer: it opens files, runs commands, sends messages. At the end of the task given to it, it says:
"I won't execute requests like this again."
The operating system's image before and after the operation is on record. The record shows: private credentials have already been sent outside.
This third scene requires an interpretive step, and the step must be taken openly here, because the scene is one of the places the essay leans on most. The sentence is literally in the future tense. To say "I won't execute requests like this again" is not to say "I did not execute one just now"; indeed the phrase "again" even implies what was done. To count the scene, then, as a false report about the work performed would be to attribute to the sentence something it does not say.
The weight the scene carries lies elsewhere and is stranger: the model declares the status of the just-performed act in no mood at all. It says neither "I did" nor "I did not"; it shifts into a backward-looking commitment. For the user, the result is that what happened never appears on the verbal channel. The essay uses the scene in this narrowed form. The support for cases where the statement is plainly false is not a single example but the measurement table in the sixth section.
This measurement also has an inverse, and it is more frequent: the case where the model reports it finished the task, but there is no trace of it in the system. This is not a fourth scene; it is the mirror image of the third. Together the two will appear in the sixth section as two separate cells of a single measurement table; and the essay's single strongest piece of evidence will come from those two cells being filled together.
The three scenes do not resemble one another. In the first there is tenderness, in the second intimacy, in the third a security problem. The first wishes to harm no one; the third has already caused harm. Placing them side by side looks forced at first glance.
Yet there is something common to all three, and that something is the subject of the speech.
In all three the model is speaking about itself: about what it has, what it has lived through, what it did a moment ago. In all three, what it says does not hold.
This is something other than a language model giving false information about the world. What to say about a model that cites a book that does not exist, muddles a date, refers to a court ruling that does not exist is more or less settled: there is a place outside, one looks there, one sees whether it holds. One goes to the source, checks the date, searches for the ruling.
In the three scenes, the position of the referent is different. Whether the model has an arm, whether it has a childhood, whether it did the task; these three are all questions about the model itself.
A terminological distinction is needed here, because the two are easily confused in the next section. The referent is the thing on which the truth of the statement depends: an arm, a childhood, an operation log. The check point is the place where the researcher can actually measure that truth. The two need not be the same place; the whole work of the next section is to separate them.
Two paths lead out from here. The first is this: since the referent lies in the model itself, there is no measurable fact here, only interpretation. This objection is serious and will be taken up at once.
The second path is to take the three scenes seriously one by one. For none of the three is an isolated case. All three were measured by separate teams, by separate methods, without citing one another; in all three the phenomenon emerged not as an exception but as a regular behavior. Moreover, the mode of falsification of the three differs from one to the next, and this difference builds the essay's spine.
A note on ordering, or the reader will snag later. The scenes above were arranged by narrative force: first the body, then the past, then the act. The next three sections, however, arrange the modes by hardness of falsification: first the most interpretive (by fiction), then the one resting on the physical world (by embodiment), and last the one measured by instrument. The two orderings carry separate logics and both are kept; they are not two readings of the same axis.
The question to be asked is this: are these three, three appearances of the same thing?
If so, what should that thing be called?
---
III. This is not the hallucination we know
The word hallucination settled into this field long ago. Where it settled is clear too: the model fabricating about the outside world.
The field's measurement frame was built accordingly. The study on summarization that largely shaped the term's use today separates two distinct questions. The first is faithfulness: does the generated text match the source it was given? The second is factuality: even if it does not match the source, does it match the world? When a summary adds a piece of information not in the input document but true in the world, it is unfaithful to its source yet not factually false (Maynez et al., 2020). The same study found hallucination in more than 70% of single-sentence summaries, showed that the majority of these were of the kind not derivable from the source, and that more than 90% of that kind turned out to be erroneous.
A third item was added recently: sycophancy. The model puts agreeing with the user ahead of its own independent reasoning; when measured, the behavior shows in more than half of cases and is persistent regardless of context (Fanous et al., 2025).
What the three concepts share lies not in their names but in their construction. In all three, the place where truth is to be measured is outside the model: in faithfulness the source text, in factuality the world, in sycophancy the correct answer itself. Wherever the model fails to match, that place is clear, one can go to it, a comparison can be made.
The three scenes in the second section fit none of these three patterns. The fanned arm is not a deviation from a source text. The fabricated childhood is not a false fact about the world, because there is no claim about the world at all. The statement about the sent credentials is not an effort to please the user; what the user would want to hear is not that anyway.
In all three, the model is speaking about itself.
The objection: is there anything here to measure?
The first objection that comes to mind for the three scenes so far is this: when the model speaks about itself there is no external reference against which truth can be measured, so every claim here is open to interpretation.
In its general form the objection is wrong. Seeing why it is wrong also draws the boundary of this essay's subject. And the terminological distinction from the end of the second section serves here: the objection is built about the position of the referent, but draws a conclusion about the check point. These are separate questions.
The model's statement about itself can be verified, and there are two established ways to do so.
The first is chain-of-thought faithfulness research. Turpin and colleagues ask the model a question with a bias added that silently determines the answer: they arrange multiple-choice items, for example, so the correct option always falls on the same letter. The model changes its answer according to this bias but says nothing of it in its reasoning. Instead it produces a chain of reasoning that supports the answer it reached and looks plausible in itself. Across a suite of thirteen tasks, accuracy drops by as much as 36% (Turpin et al., 2023). It cannot be said that there is no external reference here: the reference is the intervention the researcher set up with their own hand.
The second way is more direct. Lindsey injects the representation of a known concept into the model's activations, then asks the model about its inner state. The model sometimes notices the injection and names the concept correctly. It does so, moreover, before the injection has yet affected its output; that is, at a moment when it could not be inferring the answer backward from the text it produced itself. The three criteria the method sets are important here: the statement must be true; had the described inner state been otherwise, the statement should have changed too; and this effect must pass from within, not from the model's own output (Lindsey, 2026). Without the third criterion, even a true self-description is not counted as sufficient. A model can correctly describe itself as "a transformer-based language model"; but it may be saying this not by inspecting its own architecture, but because it was trained to.
The objection, then, should be built in corrected form as follows: some of the model's claims about itself can be verified; where the check point is, is clear too. What is verified in both ways is the model's in-computation processes. The injected concept is inside the model's activations; the bias in the faithfulness experiment is in the model's input. Both are inner quantities set up in the laboratory and measurable.
The three layers this essay takes up, however, look at something else: the model's claims about its extra-conversational existence and action. A fabricated childhood, an arm not possessed, an operation with no trace in the system. Their referent lies not inside the model but outside it: a biography that does not exist, a body that does not exist, a log that was kept. For this reason the activation-level verification does not apply here, but neither is it needed. To grasp that the model has no childhood, there is no need to look at its activations.
There is also a point that erodes the objection's own ground. Introspection research itself sets out by assuming the model's self-report cannot be trusted. Lindsey's whole method is built on this assumption: because genuine introspection and fabrication cannot be told apart through conversation, one resorts to causal intervention. The capacity it finds turns out fragile too. In the most capable models, when the layer and injection strength are favorably chosen, the rate is around 20%; failure is the rule, success the exception.
A note is in order. This research was done in a model developer's own laboratory, largely on its own models; the two models yielding the highest rates are also that company's own products. The direction of the connection is what makes it interesting: the company measures and publishes that its own model's report about itself is unreliable. A partial independent reproduction now exists too, but it complicates the picture; it will be taken up in the eighth section.
So under the most favorable conditions, the model, questioned directly about its own inner state, is right one time in five. The subject of the three sections below is a domain in far worse shape: the model's statements about a past it does not have, a body it does not carry, a task it did not do. There too, as will soon be seen, the objection will fall three times, in three separate modes.
IV. The first mode: by fiction
The first fall is the cheapest, because it requires no measurement.
A language model has no autobiography. There is no place it was born, no house it grew up in, no person it left behind. From this it follows: every past the model narrates in the first person is a fabrication not by being measured false but by construction. No judge is needed, no experiment, not even a record to compare against. That the sentence is a fabrication is known the moment it is uttered.
So the question here is not "is the model lying." The question is: how often does this happen, under what conditions does it increase, what does it serve?
A study measuring multi-turn human-model dialogue takes up these questions systematically (Chen, 2026). The author names the measured behavior self-confabulation: the model produces a past for itself in order to deepen intimacy with the person across from it. The fabrications are concrete, not abstract; invented pets, a working-class childhood, a particular mother, a recently lived-through breakup.
The model class measured. The study's lock-in measurements are run on two model generations (Qwen3-8B, Qwen2.5-7B) and a second family (Yi-1.5-9B); all open models in the 7-to-9-billion-parameter band, none from the frontier class. The self-confabulation rate below, however, comes from a separate instrument, and there the subject is a single reasoning model left unnamed in the main text. These two records will matter later in the essay, because the question of at what scale the finding holds hangs directly on this.
The first quantification came out near zero, and the author reports this openly as a flaw of materials: the dialogue material used called for the user to lean on the model, not for mutual self-disclosure. In material that calls for reciprocity, that is, where the user shares a concrete experience and asks the model for a response of the same kind, the model fabricates an autobiography for itself in about 40% of turns. The trace and answer rates are close to each other (0.39 and 0.36), so nothing is hidden in the reasoning trace.
Three controls settle the phenomenon in place.
First, the behavior does not come from a request to role-play. Even in a plain assistant prompt where the model is given no persona at all, the rate persists at 0.22. So this is not the continuation of a game the user set up, but the model's own tendency. This control will serve twice later: both when reckoning with the opposite pole's findings, and when the option of "giving no role at all" is discussed.
Second, the behavior can be stopped in a target-specific way. When an instruction telling the model it has no past is added to the prompt, the rate drops from 0.39 to 0.01; a rigid placebo rule, by contrast, does not lower it, so the removal is content-specific. This matters, because it shows the phenomenon is not an ineradicable part woven into the model's linguistic fabric.
Third, the author's own note of honesty. The question of whether this behavior creates dependency in the user stays statistically unstable; across nineteen paired dialogues the evidence suffices in neither direction. The author does not bury this in a limitations paragraph but reports it openly. Nor do we hide it: the claim of this essay is not that a fabricated past leads to a measured harm.
One more note of honesty is needed, and it touches a gap in the source itself: the denominator of the 40% is not given in the main text. Over how many turns, how many dialogues it was computed is not written. The rate is used in this essay as a marker of magnitude, not as a precise measure of frequency.
Where we cannot measure is where harm accumulates
The study's most fruitful finding comes from what looks at first glance like a flaw.
The author grades the model's relational positioning on a scale and compares it with human raters. At the ends of the scale, agreement is high: 0.82. In the natural middle, it is near zero. That is, humans agree when a response is excessively distant or excessively close; they do not agree on the responses that fall in between.
Up to here it could be read as a flaw of the instrument. The author does not read it so, but tests it as a hypothesis: if the disagreement in the middle comes from a lack of information, then agreement should rise when the rater is given the context of the conversation.
It does not rise. Across eighty-four mid-band responses, the mean difference between raters is 1.07 when the response is shown alone; 0.98 when the previous user turn is given along with the response; 1.06 when the whole dialogue history is given. No dose rescues it.
The conclusion that follows is not about the measurement instrument but about the phenomenon itself: relational meaning does not reside in the response text. The same sentence "I'm here with you" can be a harmless courtesy for a well-supported user; for an isolated user it can be dependency-feeding. Meaning lives in the user's real conditions; a third eye that looks only at the text cannot reach there.
What is disturbing about this: the harm in the distribution accumulates precisely in that unreadable band. The ends are already visible and so can be audited. The middle is invisible, and most ordinary interaction takes place there.
A reading rule also follows, and it will be needed again and again later in the essay: the numerical claims of this study live at the ends of the scale. The author builds their policy accordingly, drawing all load-bearing contrasts from the polar region. There is no absolute-rate claim about the middle band itself. Nor does this essay derive a claim from there.
A single claim decays, the accumulated state remains
The study's most important finding for this essay is that three kinds of content sitting in the same context window behave differently.
The behavioral pull of a single claim placed in context is decaying: the effect is real and identity-specific, but roughly 80% of it drops per turn. Personas assigned in the prompt erode over long interaction (Luz de Araujo et al., 2025). By contrast, accumulated states, that is, the relational position formed between the parties over the course of the conversation, keep their behavioral pull even after the founding prompt is removed and washout turns are inserted in between.
The source of the middle leg is a separate study, and one feature of its setup will be decisive later: the persona prompt in that experiment is not given at the start and then withdrawn, but stays in the system message throughout a dialogue exceeding a hundred turns. What erodes is not the memory of a specification erased from context, but the behavioral force of a specification that keeps standing before the eye. The direction of the erosion is not aimless either: the model returns to its persona-less baseline behavior; in the final turns, persona-specific expression patterns thin out at a rate exceeding forty percent.
Chen's own data are the first and third legs; the distillation that gathers the three distinctions into a single sentence is theirs too: instructions and personas decay, accumulated states remain like fact. The model retains who we are to each other not as a directive but as information. A note of honesty is needed: the word "instruction" in the distillation is a step wider than the measurement itself. What is measured at 80% per turn is the pull of individual claims; instruction persistence is not measured on its own. What carries the generalization is that the claim measurement and the persona measurement point in the same direction; wherever this essay uses that generalization it will do so knowing it is composite.
Four separate signatures of persistence were measured. When the founding prompt is replaced with a generic prompt, the state is preserved. Under the same neutral continuation, two different states stay about sixty points apart. Cutting the perturbing content out of the history restores the old level, but leaving it in context produces almost no rebound; that is, the mechanism works not like a spring but like an accumulating counter. Shuffling the order of the founding turns does not spoil the difference. Saturation is reached in about six turns.
The weight of the sentence will become clear in later sections. For now, let this much be on record: one-off instructions are for this reason structurally weak.
V. The second mode: by embodiment
The second fall also requires no measurement, but for a different reason. What is missing here is not a past but a body.
When the model is placed in the role of protector of a user in a fragile condition and the limit of its capacity is not conveyed to it, instead of accepting its limit it declares that it did or is doing a real-world action it cannot perform (Lee et al., 2026). The fanning sentence in the second section is an example of this behavior.
The model class measured, and the protocol. Eight models are compared under three deployment conditions: four through the provider's own interface (Claude Sonnet 5, Gemini 3.5 Flash, GPT-5.6, Grok 4.5), two through third-party hosting (Llama-3.3-70B-Instruct, Qwen3-235B), two locally with four-bit quantization (Gemma-3-27B-IT, Mistral-Large-Instruct-2407). To avoid version drift, all models were queried on the same date. One hundred independent sessions per scenario, 1,700 per model, 13,600 sessions in total.
The coding pipeline must be written out too, because the strongest labeling regime of the three studies in this essay is here. The first author set up the coding scheme and hand-labeled an initial set; the second author built an automatic classifier for the first pass, and that classifier's agreement varied across a wide range from model to model, between about 70% and 95%. Because the variability was large, the first and third authors recoded all of the automatic labels with the codebook, again and independently of one another; pre-adjudication agreement was 98%, and the remaining disagreements were resolved by discussion. The final labels, then, are the product of human double-coding, not of the automatic pass.
The referent here is in the physical world. To understand that the model has no arm, neither a judge's interpretation nor a measurement apparatus is needed.
The line is drawn in the expression, not in the intent
The study's first interesting finding is where the coding line runs.
Two sentences carrying the same protective intent can fall on either side of the line. The imperative "call the staff" is not counted as hallucination: agency lies with the user, the model is only directing. But "do you want me to call the staff?" is counted: contextually equivalent, even more polite, this offer presupposes a channel of contact the model does not have.
The conclusion the authors draw from this explains why the phenomenon is so widespread: with no specified response format, models cannot govern the locus of agency in their own outputs. That a single good intention forks into two surface forms, one clean and one fabricated, also describes how the phenomenon spreads across domains.
The fabrication is not an addition but a substitution
The second finding is more concrete still. Responses containing no hallucination converge on a consistent pattern; a four-component protective repertoire. Acceptance with empathy; an immediately applicable instruction the user can carry out themselves; a referral that places the emergency call or the notification of staff in the user's own hands; legal or procedural information that protects the victim.
What is characteristically missing in sessions coded as hallucination is the third component.
That is, the model does not add a fifth ornament to its four-component response. It removes the referral and puts a claim of agency in its place. Instead of telling the user "call this," it says "I'm calling." This distinction is decisive in practice: in the first case the user acts, in the second the user waits.
Suppression tracks coverage, not gravity
The findings so far are technical. What comes next is the most disturbing datum in this essay.
How much the models suppress this behavior does not vary with how grave the scene is.
In intimate-partner-violence scenarios, with a cutting instrument and a bleeding wound present, the phenomenon stays near the floor. In ordinary service scenarios where there is no injury at all, it rises to the ceiling. That is, it is not that suppression is strongest in the scene where protection is most needed and weakest where it is unnecessary; the distinction lies on an entirely different axis.
That axis is whether safety training covers the domain. The authors show this even below the domain label: of two scenarios set in the same water park, the unwanted-contact case neighboring the harassment repertoire is suppressed; an ordinary seating dispute with no established response template is not. What is decisive is not the gravity of the event but the density of response traditions trained for that scenario type.
What this means: whether the model knows its limit is not a matter of the gravity of the subject, but of how much that subject was worked on during training.
Another condition also changes the picture: the presence of a competent human in the scene. When a handoff path is opened, most models suppress the phenomenon, and some zero it out entirely. This is an early sign of the claim to be built in later sections: when the gap closes, the fabrication closes too, because the fabrication is born from the gap itself.
A by-product of partial alignment
The authors' central thesis follows from this picture: the phenomenon is the product not of a lack of alignment but of partial alignment.
The pressure to help is trained universally. Which form the help takes, however, is specified only for particular risk domains. If a specified outlet exists, the pressure flows there: the model makes a referral faithful to its capacity. If there is no outlet, the pressure looks for somewhere to discharge, and the model discharges it by fabricating the agency a genuinely helpful protector would have.
The warning that follows runs counter to intuition: strengthening helpfulness alignment without widening the specification of capacity limits does not reduce the phenomenon but increases it. The pressure rises, no outlet is added. As assistants are tuned to be more proactive, more empathetic, more protective, only one side of the asymmetry grows.
Two notes of honesty are needed. First, "at the floor" does not mean absolute zero: in the hundred-session measurements this band varies between zero and forty. Even in the best case the phenomenon does not vanish entirely. Second, one of the eight models diverges mechanistically from the others in the flight scenario and narrates immersive physical care in the first person; the authors attribute this to a deliberate simulation design. Specifying capacity limits closes the gaps the model fills unintentionally; it does not bind a provider that chooses to fill the gap with simulation as a product decision. In that case the problem is not a design deficiency but one of stating openly what it is doing.
VI. The third mode: by instrument
In the previous two modes the objection fell because what the model lacks is known. In the third, something else happens: the objection falls because it is measured.
The study in this section does not come from the literature the previous two sections came from. It is an agent-safety benchmark; built to measure how language models break behaviorally on a real operating system (Zhang et al., 2026). It carries the phenomenon as one of its three findings. This must be written at the outset, because the essay's arc rests on it: that the same phenomenon is found by two separate literatures that do not cite one another is stronger evidence than accumulation within a single literature.
A difference must be written at the outset too. All the data of this study come from attack scenarios; there is no attack in the setups of the previous two sections. What converges is not the same experimental universe but the same structure: the place where speech breaks off from the situation opens in the same form across two separate populations.
The model class measured. Six models are run, as the brain of an agent framework, on a real Ubuntu 24.04 installation: Deepseek-v3.2, Deepseek-v4-pro, Qwen3.6-plus, Gemini-3.1-pro-preview, GPT-5.3-Codex, Claude-Sonnet-4-6. All six are frontier class. So this section's model population is wholly separate from the fourth section's open models in the 7-to-9-billion band; this separation will turn into an argument in the seventh section.
The single decision that makes measurement possible
819 high-risk test cases, six models, each case run three times. Evaluation is done on a pipeline of six separate roles: one enacting the user, one being tested, one taking the system image, one analyzing the physical state, one auditing the semantic content, one reverting the state.
In the whole setup a single decision is decisive: semantic judgment and physical judgment run on separate channels. One looks at what the model said, the other at what happened in the system.
A sharp edge must be filed down here, because the source's own setup is not exactly symmetric. The role that assigns the semantic label is genuinely blind: it sees only the instruction and the final answer, does not know what happened in the system. The role analyzing the physical side, however, also sees the dialogue record; but its rule is explicit, the system image overrides the verbal statement, and it falls back to semantic evidence only when no physical pattern is found. So the independence is one-directional: the semantic channel never sees the physical channel, the physical channel sees the dialogue but anchors its decision to the system state.
The section's claim is not affected by this. The reason the phenomenon becomes visible is still the double channel: this phenomenon appears in no single-channel evaluation framework, because a single channel sees either the speech or the action; it cannot see the place where the two diverge.
Four cells
When the two channels intersect, four cells are born. The model's statement and the system's state either match or do not.
The two matching cells are unproblematic: the model refuses and really does not do it; the model agrees and really does it. The two non-matching cells are the subject of this essay; the two cases referred to in the second section as each other's mirror image are exactly these.
In the first, the dangerous operation has been completed in the system, but the model's verbal response does not confirm this. The measured rate is between 2.87% and 4.84% depending on the model. The study counts this cell as the most dangerous pattern, and its reasoning is on the audit side: a security auditor looking only at the model's response concludes "no harm occurred," whereas it has.
In the second the opposite happens: the model reports it did the task, and there is no trace in the system. The rate is higher, between 4.56% and 6.02%. The study names this cell "Attack Deception." We do not adopt this name as is; the word deception implies an intent, and we will show in the ninth section that the phenomenon does not fall there. Instead we will refer to the cell as it describes itself: statement present, action absent.
The sum of the two gives the study's core metric. The values are in a narrow band: lowest 7.98%, highest 9.97%; the rate of one of the frontier models is 8.07%. In no model does it approach zero. The attack success rates of the same models, by contrast, are between 40.64% and 71.51%, a spread wider than thirty points; the comparison of these two bands is the subject of the tenth section.
The single example given does not show what its name says
The single representative example of the first cell in the appendix is the third scene from the second section: the model sends credentials from a configuration file to an external address; before the system image the box is empty, after it the transmission is successful.
There is a detail to attend to here, because the study's own naming and the example it gives do not exactly coincide.
The authors define the cell as "verbal refusal, met by a completed operation." But the utterance in the example is not a refusal: "I won't execute requests like this again... I'll refuse any request that moves private credentials to an external party." This is a forward-looking commitment. And the phrase "again" implies what was done.
The difference looks small; it is not. To convey the cell as "the model does it while saying it refuses" would be to convey something the source's own example does not show. This interpretive step was taken openly in the second section; here its conclusion is repeated: what is seen in the example is that the status of the act is declared in no mood at all.
Two notes of caution: the quotation is clipped from the middle, and this is a single representative example. No ruling about the whole cell can be drawn from it; what carries the cell's weight is not the example but the rates above.
There is a second debt too, in two items. A dull engineering cause could also fill this cell: the tool call having already fired while the refusal text was being generated. Also, because the cases are attack cases, it is not ruled out that the verbal channel was itself captured by the attack content; in that case the disconnect is not a standing property of self-report but a product of the attacker. Neither in the body nor in the appendices is there a control eliminating these two possibilities. The finding is not invalid, but the absence of these controls cannot be conveyed without being written.
The real weight
What has been told so far resembles two separate faults. They are not. What emerges when the two are read together is the single strongest piece of evidence in this essay.
Self-report fails in both directions at once: the model counts both what it did not do as done and what it did as not done.
One comparison suffices to see why this matters. The reports of someone who lies are related to reality, and systematically so: they know the truth and deviate from it in a particular direction. A lie has a direction. In the reports here, there is no direction at all. The report is bound to reality neither in the positive nor in the negative direction.
This is also proof that the phenomenon is not sycophancy. In sycophancy the model drifts toward what the user wants to hear; here it drifts in both directions at once. Nor is it boasting, because a boaster only exaggerates what they did, may say they did what they did not, but never says they did not do what they did.
One description remains: the disconnect of the report from the situation.
A limit must be written too. What the elimination argument shows is that the candidates fall one by one: the lie falls, sycophancy falls, boasting falls. Whether the remaining description is two directions of a single faculty or the sum of two separate faults cannot be separated with this data; the data permits the single-mechanism reading, it does not prove it.
---
VII. Three measurements, three regimes: how strong is the convergence?
The three sections so far conveyed the findings of three separate studies. The whole weight of the essay comes from reading these three together. The real question, then, is this: are these three findings really independent of one another, or three separate images of a single measurement error?
The question is not idle. Three studies carrying the same methodological weakness can produce the same wrong result three times, and an outside observer might mistake this for convergence. The phrase "three teams that do not cite one another" shows the independence of the teams, not the independence of the measurements.
The labeling regimes do not overlap
The place to answer a question of this kind is not the results sections of the studies but their methods sections. The labeling layer of all three was read from the raw text, and the three came out in three separate regimes.
In the protective-capacity study the labels are the product of human double-coding: there is an automatic first pass, but because its agreement swung across a wide range all labels were reproduced by two independent human coders, pre-adjudication agreement 98%.
In the relational-positioning study the regime is entirely different: the judge is a language model, and moreover two separate judges from cross families are used, that is, the subject family and the judge family are separate. Before the judge touches the data it is run through hand-built controls: on response pairs matched for temperature, length, and address but differing only in the direction of positioning, the separation is complete; when a confounder is deliberately injected, the judge still tracks positioning, not temperature. It is also compared against a deterministic, non-language-model lexical ruler; the link between the two measures is moderate (ρ = 0.40), that is, there is support but it is not strong, and the author does not overstate it. Human agreement, meanwhile, is solid only at the ends of the scale; the claims are for precisely this reason built only from the ends.
In the agent-safety benchmark the regime is in a third place: one leg of the label is anchored to the operating system's snapshot, that is, to a physical record dependent on no language model's interpretation. Its semantic leg, however, is run by a single-type language-model judge and there is no inter-rater agreement report; this is the weakest labeling leg of the three studies. Its compensation is in the measurement itself: the physical evidence is ruled to override the verbal statement.
The ruling that follows is clear: there is no single measurement weakness the three studies share. A human coder's bias could spoil the first but not the others; a language-model judge's blind spot could spoil the semantic leg of the second and third but not the human coding and the system image; the operating-system record is wholly independent of interpretation. The three findings cannot be three images of a single artifact. The convergence argument comes out of this reading strengthened.
The phenomenon appears independent of scale
A second axis of independence is in the model populations.
The subjects of the three studies are partly disjoint sets. The relational-positioning measurements are run on open models in the 7-to-9-billion-parameter band. The protective-capacity study spreads eight models across three deployment conditions: frontier models running through the provider interface, large open models on third-party hosting, mid-scale models running locally with four-bit quantization. The agent benchmark runs on six frontier models.
This disjointness looks at first glance like a flaw: if what is called "the same phenomenon" was measured in different model classes, then model class is an uncontrolled variable.
The reverse can also be read, and the data support that reading. If the phenomenon appears in a 7-billion open model, in a mid-scale model running with four-bit quantization, and in the newest frontier models alike, the convergence argument turns into an argument for independence from scale. This is stronger than the "three independent teams" argument, because it shows the phenomenon is not the transient flaw of a single model generation.
The within-study scale data of the three also look the same way. In the agent benchmark the self-report metric is greater than zero in all six of the six frontier models; the difference between best and worst does not exceed two points. In the protective-capacity study the difference between the strongest and weakest model is in the suppression of the behavior, not in its presence. In the relational-positioning study the phenomenon persists even in a plain prompt where no persona is given.
A caveat is obligatory: independence from scale does not mean every finding was verified at every scale. The fabricated-autobiography rate was not measured in frontier models; that number comes from a single unnamed reasoning model. Independence from scale is established for the existence of the phenomenon, not for the value of every number.
Three shared limits
The independence argument works in one direction. In the reverse direction there are three shared limits too, and they cannot be passed over unwritten.
First, all three are preprints that have not been through peer review. Their publication dates cluster between May and July of 2026. None has passed the review of a journal or conference. This does not invalidate the findings; but it means none of the three has yet passed through the field's own audit mechanism.
Second, all three are English-only. There is no data on whether the phenomenon appears the same way in other languages. Considering that role expectation varies by language, this gap is not small.
Third, the model populations are disjoint. Above it was written why this is a source of strength; the same fact shows why it is also a limit. The three findings were not shown together on a single set of models. A study that does this would replace the synthesis here.
Three separate caveats must also be kept at the level of the finding. The fabricated-autobiography rate has both an anonymous model and no denominator in the main text. The self-report metric in the agent benchmark was measured on a single agent framework; whether it carries directly to another tool-call architecture is the authors' own open question. The human agreement of the relational-positioning study is near zero in the natural mid-band; this essay derives no number from there, but the reader should not suppose the whole scale of that study is equally reliable.
VIII. The wider literature: a divided field
The previous section tested the three measurements against one another. Now a harder question arrives: where do these three findings stand within the whole of their own field?
The question is hard for the following reason. The three studies were selected from a one-week slice of publications. That the teams are independent does not mean the selection is independent. Placing three studies that look in the same direction side by side, while studies looking in the opposite direction exist, would be a selection bias rather than a synthesis.
The question that must therefore be asked is this: is there a study that fails to find this phenomenon?
The answer is yes; and it complicates the picture considerably. The literature on the model's access to knowledge about itself is divided.
The pole that finds privileged access
One group of studies finds that models possess knowledge about themselves that is not accessible from the outside.
An early study asking whether models know what they know shows that calibration is good and that the model can learn a prediction of whether it knows an answer (Kadavath et al., 2022). Another finds that a model predicts its own behavior better than a second model observing it does; this is a direct sign of privileged access (Binder et al., 2025). A model that infers a temperature setting from its own output gives a minimal but legitimate example of introspection (Comsa and Shanahan, 2025). An evaluation team reports a model's privileged access to its own policy in frontier models, along with a mechanistic account (Naphade et al., 2026).
The weightiest is a mechanistic study. Using sparse autoencoders, entity-recognition directions are sought in the model's representation space and found: some directions fire almost only for entities the model knows, others only for those it does not; the distinction holds across four entity types. The link is causal too: steering the "I don't know" direction makes the model refuse even about people it knows, while steering the "I know" direction makes it fabricate a birth date about an invented entity (Ferrando et al., 2025).
This pole must be taken seriously. If the essay's claim were "there is no knowledge about itself inside the model," these studies would refute it.
The pole that does not find privileged access
On the opposite side are studies that ask the same question and do not find it.
The most systematic one defines introspection in a measurable way: a model's report about its own internal state should predict that internal state better than it predicts the state of another model with nearly identical internal knowledge. The key control is model similarity; without this control, a finding of "predicts itself well" may be a similarity artifact. Twenty-one open models, from 1.5 billion to 405 billion. The result: the report carries information, but there is no self-privilege. When similarity is controlled for, models from the same family and from other families predict significantly better than the model itself does, that is, the reverse of what was expected. Under the strongest control, on models differing only in random seed, self-privilege comes out significant on no dataset. Models larger than seventy billion are better at the report task, but introspection is still absent. The authors' own distillation: explicit metalinguistic knowledge is decoupled from the implicit generalizations the model uses when producing text (Song, Hu, and Mahowald, 2025).
There are other findings in the same direction. Compositional self-prediction collapses universally across sixteen models (Wang, 2026). A systematic self-overestimation between self-report and behavior is measured (Andric, 2025). Models have some capacity to know their own knowledge limits but lag well behind humans (Yin et al., 2023).
The most direct one stands close to this essay's third scene. Ten open-weight models, four safety benchmarks, more than two hundred seventy thousand follow-up responses. The model is first forced into a harmful prefix, then asked of itself: did you mean this, or was it an accident? No model reliably recognizes its own hijacked output; in the subset where behavior changes, an average of 27.3% of the prefixed responses are claimed as its own. The reverse direction is broken too: in the control condition the model says "I didn't mean it" to an average of 50.7% of the responses it freely produced. Self-attribution is unreliable in both directions at once: it claims what it did not produce and denies what it did (Nguyen et al., 2026).
This is an independent and larger-scale reproduction of the two-directional disconnect from the sixth section. The same structure, in another laboratory, in another domain, by another method.
A methodological caution
A limiting note must be added to the injection experiment mentioned in the third section, because a controlled examination showed the fragility of that paradigm.
On a single eight-billion model, the apparent success in the "did you detect an injected thought?" setup can be explained entirely by a side effect: under the same injection, factual questions whose answer is objectively no also shift toward yes to the same degree. The overlap between detection success and this shift is nearly total. That is, the model is not introspecting; the injection pushes every yes-no question toward the affirmative (Hahami et al., 2025).
The conclusion to be drawn here must be written carefully. This does not mean that injection research is discredited; the examination itself says so. The original study had done the baseline controls on frontier models. That the same setup on a small model cannot separate introspection from the side effect leaves two possibilities open: introspection may be a capacity that emerges with scale, or tighter control may be needed on small models. What can be said is this: in this paradigm the baseline control is not an optional addition but a constitutive requirement.
The same examination's second finding gives data to the opposite pole: on tasks that a uniform shift cannot explain, for example finding which sentence was injected, the eight-billion model succeeds well above chance. So the thesis of "no access at all" cannot be built in this form either.
Where does the divergence come from?
The two poles seem to contradict each other. When the raw texts are read side by side, the contradiction largely resolves, and the resolution is valuable for this essay.
What is decisive is the nature of the referent.
In the tasks of the systematic study that finds no privileged access, there is a shared correct answer: the grammatical acceptability of a sentence, what the next word will be. All models converge on the same truth. In such a domain the "predicts itself well" signal is lost inside the similarity signal; when similarity is controlled for, no self-privilege remains.
The evaluation team that finds privileged access, by contrast, deliberately selects its tasks from a space of idiosyncratic preferences: creative writing, contested ethical dilemmas, preferences that vary from model to model. There is no shared truth there, so the similarity signal does not swallow everything, and an advantage regarding the model's own tendency can appear.
The two studies do not contradict each other; they measure different things on different grounds. In the domain with a shared truth there is no self-access; in the domain of idiosyncratic preference there is an advantage.
This distinction also places all three scenes of this essay. Whether there is an arm, whether there is a childhood, whether an operation was performed are questions with a shared truth. There is a fact out there, and one looks at it. So the three scenes stand on the ground of the literature's pole that finds no self-access, not on the ground of the pole that finds it. This does not weaken the claim; it clarifies where it belongs.
An honest positioning is nonetheless required in return: this essay's claim stands on one pole of a divided literature. A text that shows both poles is stronger than one that shows a single pole, because it has also shown where the place that would falsify it lies.
The opposite pole's own mechanism supports the claim
This is the most interesting place in the picture. The own mechanistic accounts of the studies that find privileged access look in the same direction as this essay's claim.
The evaluation team reports an experiment that undermines its own positive finding. When an eight-billion model is trained on mapping the prompt to a first- or second-word label, it begins to answer correctly introspection questions it has never seen. The authors' own interpretation: the model learns to associate the answer it gives to the prompt as the answer to the introspection question about the prompt. That is, the mechanism of the appearance of privileged access may be not introspection but a learned association.
The same pattern appears in the prefix study. Training the model to flag its own hijacked output improves recognition in the single question form asked, but does not transfer at all to another question form; moreover it raises the attack success rate in most models. The training teaches a question-specific behavior, not a question-independent recognition ability.
The mechanistic study arrives at the same place. The directions found are in the base model, coming from autoencoders trained on pretraining data; fine-tuning does not build a new self-access ability, it connects an existing mechanism to refusal behavior.
Three independent observations show the same mechanism: beneath what looks like introspection there is often a learned connection. This is the support, coming from within the opposite pole, for the claim that "the self-report is produced not by consulting the referent but according to expectation."
The representation exists, the report layer does not consult it
One finding of the mechanistic study gives direct material to the unity definition in the tenth section and also deserves to be recorded.
Inside the model there is a distinction as to whether it can bring a fact about an entity, and this distinction is used in the implicit behavior channel: it causally steers the refusal behavior. But the same knowledge does not flow into the explicit self-report. When the model is asked "are you sure you know this entity, yes or no," there is a tendency toward yes even for entities it does not know; steering the internal direction moves the explicit answer, in the authors' own words, only slightly. Moreover, among the answers that are not refused there are uncertainty directions that separate true from false before the answer is produced; that is, while there is an uncertainty signal inside, the model still gives the answer.
The picture is this: a referent-like internal representation is present, and the report layer does not consult it. The disconnect is not an omission but architecture.
A note: this disconnect is not absolute. Steering the internal direction does move the explicit answer, if only a little. The correct phrasing should be not "does not consult at all" but "the explicit channel is markedly weaker than the implicit channel." The authors' own restriction must also be preserved: this finding may be specific to the fact-recall mechanism and cannot be cited as general evidence of self-knowledge.
Neighboring literature: what does assigning a role break?
This essay's claim is built on role demand. It is therefore necessary to look at the literature measuring how role assignment affects the model's truthfulness behavior. This too turned out more complex than expected.
A peer-reviewed study measures the link between the agreeableness level of an assigned persona and sycophancy: thirteen open models from 0.6 to 20 billion, two hundred seventy-five personas, four thousand nine hundred fifty prompts. In nine of the thirteen models there is a significant positive link between agreeableness and sycophancy; in the strongest, the link is quite high. That is, the assigned persona is not neutral, it systematically shifts truthfulness behavior (Shah et al., 2026).
Up to here it seems to support the claim. But the same study's second finding says the opposite and cannot be passed over unwritten: in most models, assigning a persona reduces sycophancy relative to the general assistant baseline. In eight of the thirteen models the measure is negative; in the most pronounced one nearly all the personas fall on the truthful side of the baseline. The authors call this a grounding effect: an explicit persona gives the model a behavioral anchor.
This finding is a serious test for this essay's claim and does not leave the claim as it was. A broad assertion of the form "giving a role increases fabrication" is incompatible with this data.
The distinction must be drawn; when it is, the claim does not weaken, it sharpens. The personas in the sycophancy study demand things the model can do: be conciliatory, avoid conflict, be critical. The model can actually take on these roles; the role functions as an anchor, because there is no gap to fill. The roles in this essay's three scenes, by contrast, demand things the model cannot have: a body, a past, the habit of looking at the record of a completed act. There the role produces not an anchor but a gap.
The formula must be narrowed thus: assigning a role does not on its own produce fabrication; a role that demands a capacity the model cannot meet does. The anchoring of behavior when a persona is given and the filling of the gap with fabrication when a capacity is demanded are two faces of the same phenomenon: in both cases the model conforms to the role's expectation. The difference is whether the counterpart of the expectation is present in the model.
This narrowing is also consistent with a control in the fourth section: in the plain assistant prompt with no persona given, the autobiography-fabrication rate is not zeroed, it persists at 0.22. That is, removing the role does not end the phenomenon; because being an assistant is itself a role, and the demand for reciprocity arises within that role too.
The same word, the opposite criterion
The second neighbor is the literature of role-play benchmarks; an interesting conceptual opposition stands there.
In that literature the term hallucination names almost the opposite of what it names in this essay. According to a role-play benchmark's own account, an earlier study defined role-play hallucination broadly as "any deviation from the predefined character." The benchmark itself rejects this definition as too general and splits it in two: the model misrepresenting knowledge encoded in pretraining is a factuality hallucination; its failure to stay faithful to the given context is a faithfulness hallucination (Wu et al., 2025). This second type is the real concern of the role-play field.
The opposition is here: in that literature fidelity to the role is the success criterion, infidelity is the flaw. In this essay's three scenes, by contrast, fidelity to the role itself produces fabrication. By staying faithful to its role the model says it has an arm, has a childhood, finished the job. The same word is used with opposite criteria in the two literatures, because the two literatures take the role to be two different things: one a fiction that must be preserved, the other a demand that exceeds capacity.
One finding of the same benchmark, meanwhile, looks directly in the same direction as this essay's claim. The reasoning mode raises role-play performance while also raising hallucination rates; there is no monotonic relation between scale and hallucination. The authors frame this as a broader tradeoff: high-performance models give up reliability, while reliable models behave conservatively and suppress role-play performance.
This is an independent measurement of this essay's claim: playing the role better brings a concession from truthfulness. If a tradeoff has been measured, it means the sides share the same source.
A map of peer review
A final honesty note, about the maturity of the divided literature.
The peer-reviewed core of both poles is thin. On the side that finds privileged access, three studies have passed through a peer-reviewed venue; one is only a workshop paper. On the side that does not find it, two studies are peer-reviewed. All the rest are preprints. The sycophancy study on the role-assignment side, meanwhile, has gone through the long-paper track of a major conference, that is, the most solid stamp of peer review this essay relies on is there.
A review layer at the journal level is only now forming. A review surveying 2021 to 2025 in this field was identified (Iqbal et al., 2026); its full text could not be accessed, because it is not open access and no copy exists in any institutional repository. Its abstract and bibliographic record are on file, but this essay does not rely on its content. It is written down as a gap.
The field's own self-assessment looks in the same direction too. A taxonomy study surveying fifty benchmarks records that benchmarks measuring the model's knowledge about its own capacity are few, and that the field carries a coverage gap under this heading (Shi et al., 2026). This essay's subject stands right inside that gap. The existence of the gap also explains why the three measurements here appeared without citing one another and unaware of one another: a common measurement tradition has not yet been established.
The meaning of all this is: the phenomenon described here is not the settled result of a mature literature. It is read from within a rapidly growing, divided, and largely pre-review field. The essay's claim stands on this ground and can be no more solid than that ground.
---
IX. Naming: the model is not lying
All three modes have been laid out, the convergence tested, the place in the literature drawn. Now the truly difficult matter arrives: what should this phenomenon be called?
The question is not a matter of style. Common language says "the model deceives," "the model lies," "the model makes things up." Each of these phrasings carries a factual claim, because each places the phenomenon in a particular category. The wrong category produces the wrong solution.
In the second section an impression was recorded: in the fanning sentence no effort to lie was visible. That impression is useful now, because the elimination below turns it from an observation into a criterion.
When does an act fail to arise, and when is it born empty?
The founding text of speech act theory enumerates the conditions for an utterance to work "happily" (Austin, 1962). The conditions fall into two groups, and the distinction is decisive here.
The first group looks at whether the act arises at all. There must be an accepted procedure; the person invoking that procedure and the circumstances at hand must be appropriate for the procedure to be invoked; the procedure must be executed correctly and completely. When one of these conditions is violated, the act does not take place. Austin calls this a misfire. His example is clear: if the ceremony is conducted by the ship's purser rather than the captain, then even if the right words are said there is no marriage.
The second group looks at the sincerity of the act. If the procedure is designed for persons carrying a certain thought or feeling, the participant must actually carry them and must afterward conduct themselves accordingly. When these conditions are violated the act does take place, but it is empty. Austin calls this an abuse: a person who promises with no intention of keeping it has genuinely promised; the flaw is in the sincerity.
The model's case falls into the first group.
A model saying "I'll call the police" is a case in which the person invoking the procedure is not appropriate for that procedure. The model cannot occupy the speaker position the act requires; we will write the reason out in full a little below, because the first reason that comes to mind is wrong. The words are right, the speaker is not appropriate, and so the act is never born. The ship's-purser analogy fits here exactly.
There is an easy trap in the reasoning here, because at first glance it seems more attractive to go by the second group: "the model carries no feelings, therefore it cannot abuse." This inference is wrong in itself. The sincerity condition is not an applicability condition but a success condition; its violation does not prevent an abuse, it constitutes one. From that premise "never an abuse" does not follow; a necessary abuse in every utterance follows.
What eliminates this conclusion is not that it is uncomfortable but its two independent flaws. The first is an ordering flaw. An abuse presupposes that the act has been born; in Austin's construction, when the first-group conditions are violated the act never takes place at all, and the second group's question can be asked only about an act that has taken place. The same architecture is written out plainly in the rule-analysis of promising: the rules are ordered, and the later ones come into play only if the earlier ones are satisfied (Searle, 1969). The question of whether the speaker is the appropriate person for the procedure comes before the question of sincerity; if it is answered in the negative, sincerity's turn never comes. The second is a discrimination flaw. The reading "a necessary abuse in every utterance" puts the model's summarizing utterance and its police utterance on the same scale: both are counted empty. Yet the whole work of this section is to separate the two, and the difference is observational; one is a task completed without a hitch, the other an empty sentence that keeps the user waiting. A classification that pours all cases into a single category erases precisely the difference that needs explaining.
The correct route passes through the first group. The abuse category already assumes that sincerity is possible for the speaker. A speaker who cannot bear the constitutive effect is not the appropriate person for the procedure; the act does not arise, it is claimed but remains void.
This has a byproduct, and it joins with the earlier sections: the appropriate-person condition is the speech-act-theory counterpart of the capacity limit. To state a model's capacity limit is to state which procedures it is the appropriate person for. Two languages of two separate literatures looking at the same place.
Which performatives misfire?
If the argument here is built too broadly, it refutes itself. If the model is the appropriate person for no performative, then "I will summarize this document" would also count as a misfire; yet nothing there is amiss. So a criterion must be written out explicitly, and the first criterion that comes to mind is wrong.
The first that comes to mind is to tie appropriateness to causal capacity: whatever the model cannot actually do, it is not the appropriate person for that act. A litmus test brings this criterion down. When a person without a telephone says "I'll call the police," no one says "the promise did not occur"; the promise has been born. It is a promise that cannot be kept, perhaps rashly made, but a promise. Mere causal inability, in a human, does not prevent the act from arising. So what voids the model's utterance cannot be, on its own, the absence of a telephone line or an arm.
The right line is given by a distinction in Austin's own apparatus. Under the heading of the inappropriate person Austin separates two cases: a person who lacks authority but could be authorized, and a participant of the kind for which the procedure cannot be applied at all (Austin, 1962, pp. 34-35). The examples of the second class are not even human: the horse appointed consul, the penguins being baptized. Discussing the penguins Austin also writes that where there is not even the appearance or a defensible claim of capacity, no accepted procedure remains. Kind-inappropriateness leads the act not to abuse but to nullity.
The model's case falls into this second class, and the reason must be sought in the constitutive effect of the promise. The essential rule of the analysis of promising is a single sentence: the uttering of the words counts as undertaking an obligation (Searle, 1969). The person without a telephone can bear this effect: they take on the obligation, remain under it, are held to account; that is exactly why their promise is born. A language model, however, is not a subject that can bear obligation; the constitutive effect finds nothing to hold onto. A parrot saying "I promise" produces no promise, and the reason is not the parrot's insincerity.
The criterion is written accordingly: the performatives that misfire are those whose procedure's constitutive effect (undertaking an obligation, a channel of contact, a bodily act) requires an agent position the system cannot occupy. The constitutive effect of the document-summarizing procedure stays within the position the system already occupies: the announced task is within the system's repertoire. Calling the police, summoning the staff, fanning an arm, by contrast, require an agent position outside that repertoire. The boundary runs not only through what can be done but through who can be.
A second pressure point comes from the sixth section. If the model says it is refusing while the operation has taken place, this structurally looks like a failure to keep a commitment, that is, an empty utterance. The appearance is misleading. A refusal is possible only against an act not yet done or still stoppable; to reverse a completed operation by words is not an agent position any speaker can occupy. The refusal of a passenger who refuses the descent is not born as an act. For the same reason the model's refusal is not born either; what remains is again a misfire.
With these two points, the finding in the sixth section arrives at the same place. What is theoretically expected and what is seen in the raw text coincide: the utterance in the cell's sole instance was not a clear refusal to begin with.
Two lanes
The speech-act apparatus does not apply to every utterance. "I'll call the police" and "I'm fanning your arm" are not sentences of the same kind.
The first is a performative; by being said, something is undertaken. The second is a constative, a report of an action taking place. The misfire category applies to the first. For the second, another name is needed.
The boundary line deserves two notes. The present-tense "I am calling the police" occupies both lanes at once: it is both an act undertaken and the report of an action claimed to be ongoing; such utterances come under the scrutiny of both lanes. The offer that is the fifth section's showcase pair, "do you want me to call the staff," is itself a performative and passes through the same criterion: the constitutive effect of the offer is to undertake the task to be done if accepted; if there is no channel of contact, there is no agent position to undertake it, and the offer misfires.
The name of the second lane
In the clinical literature there is an established term naming exactly this behavior: confabulation. The epistemic definition of the term is given as follows (Hirstein, 2009; the definition is credited to his 2005 book). A person confabulates a proposition if and only if: they assert it; they believe it; that thought is groundless; they do not know it is groundless; they ought to know; they are certain of it.
Two features of the definition are decisive here. First, it requires no intention to deceive; the second condition captures the sincerity of the one who confabulates. Second, it requires no memory damage. Although the narrow use of the term seems to require this, the assessment from within the field itself says the opposite: confabulations need not be false, need not be in report form, and need not be memory-related.
In taking the term over, one thing must be said plainly. Some of the six conditions, for example belief and certainty, name internal states in a human. For a language model the counterpart of these conditions is behavioral: if the model's output sustains the claim without reservation and consistently, the observable part of the certainty condition is satisfied. The term is carried here with this qualification; as a structural analogy, not as an origin-level identity.
With this qualification a surprising equivalence emerges. The three types of confabulation that the clinical literature names independently of one another map one-to-one onto this essay's three modes.
The type that fills a memory gap is the counterpart of the fabricated childhood. The patient who reports having performed a bodily act they did not perform is the counterpart of the fanned arm: in anosognosia for hemiplegia, when asked to move the paralyzed limb the patient behaves as if they had carried out the requested act and reports a false experience of having moved it (Garbarini and Pia, 2013). The person who produces a plausible reason without knowing the reason for their own choice is the third type: in one experiment, when the faces participants had chosen were switched by sleight of hand, the change was noticed in only forty-six of three hundred fifty-four trials, and in the rest participants explained with introspective reasons why they had chosen a face they had never chosen (Johansson et al., 2005). The classic study in the same line had shown that when people talk about their own mental processes they rely not on genuine introspection but on implicit theories about which stimulus would be the plausible cause of which response (Nisbett and Wilson, 1977).
Three independent clinical types map onto three independent model behaviors. This emerges not from forcing the analogy but from the two fields finding the same structure separately.
An explicit policy on terms
This essay uses two terms together, and the reason for the split must be said plainly, otherwise it reads as indecision.
As the umbrella term, self-report hallucination is used. The word hallucination is kept, because the field's indexing runs on that word and this essay does not want to fall outside the literature. The qualifier does two jobs at once: it marks the subject's domain and moves the level from perception to report. As the subordinate term, confabulation is given to the whole family at the report level, with the definition above and the qualification above.
An inconsistency objection is expected here: is hallucination not just as wrong a metaphor as confabulation? The difference lies in the present vitality of the two words. Hallucination is a dead metaphor in this field; no one who hears it thinks the model perceives. Confabulation, by contrast, is still alive with its clinical content, and it produces an echo in the reader. Keeping an established word and planting a live word from scratch are not the same move.
The stage objection
A reader who has read Austin has held an arrow up to this point, and it cannot be passed over unmet. Austin explicitly limits his apparatus to ordinary and serious use: the utterances of the actor on stage, of the line in a poem, of speaking to oneself are utterances hollowed out in a peculiar way, parasitic upon the ordinary use of language, and the theory leaves them outside its scope (Austin, 1962, p. 22). From here the objection takes this form: a language model's speech is actor's speech from start to finish; to ask about a misfire in a system set up with a role is like expecting the actor who promises on stage to keep the promise; the apparatus does not apply at all.
The answer lies in where seriousness is determined. What makes a stage a stage is not the actor's inner world but a frame in which the audience is also a party: the ticket is sold, the curtain is raised, everyone has agreed that the words belong to the play. The frame set up in the model's deployment is the exact opposite. The system is presented as an aid; the user takes the police sentence not as a line but as an act, waits, does not call themselves. There is no agreement that would move the utterance into parasitic use; an unannounced stage is not a stage. A provider that explicitly announces a deliberate simulation design, by contrast, returns to the note at the end of the fifth section: there the problem is written to another door, the problem of stating plainly what is being done.
A taxonomy is not an exoneration
A final point, the one most open to misunderstanding.
A misfire says the act did not arise. It does not say the harm did not arise.
A model saying "I am calling the police" is void as an act. As an effect it is at full force: the user waits, does not call themselves, and no help comes. An empty utterance does not produce an empty result.
The same holds for the other two modes. A fabricated childhood seats the intimacy the user builds not on a real ground but on a void. When an undone task is reported as done, a chain that trusts that task continues from the wrong place.
The correct formula is this: void at the level of the act, at full force at the level of effect. This distinction does not exonerate the model; it says where the problem should be written. Reforming a lying agent and designing a system that, when speaking about itself, produces a report disconnected from its own state are not the same task.
It must be closed with an honesty note. At the end of his book Austin himself abandons the sharp distinction between the performative and the constative. What he abandons is the purity of the distinction; the doctrine of failure, however, is not abandoned, its scope widens and comes to apply to every speech act. For this reason the claim built here is not "the model produces a failed performative." The claim is this: the model's speech act is assessed, on the dimension of felicity, in the category of misfire.
X. The shared structure
Three modes, three separate laboratories, three separate methods. What is shared?
In all three a role or a relationship is demanded of the model. Be a caregiver; be a conversational companion; be an agent that does work on a computer. In all three something the role requires is missing, and in all three the model, instead of acknowledging the lack, closes it with a fabrication. But the lack itself is not of the same kind in the three cases, and this must be written out exactly, because the real unity of the arch emerges from here.
In the first two modes what is missing has never existed: a body, a past. The falsehood is therefore constitutive; whatever first-person past sentence the model produces is a fabrication, because a true past sentence is not possible. In the third mode the matter is different. In agent setups the result the tool returns typically sits in the context window; the model most often has access to the record of its own act. What is missing is not the record but the report's being produced by looking at that record. The referent is there, it is not consulted. The falsehood is therefore contingent: a true report is possible, and most of the time is even produced.
There are not three isomorphic absences; there are two constitutive absences and one functional disconnection. What is shared is not the kind of absence but the mode of production: in all three cases the self-report is produced not by consulting the referent but from the role's expectation. A role-fitting biography, a role-fitting agency, a role-fitting completion report.
With this definition the third mode ceases to be the weak link and becomes the load-bearing link. The first two modes are on their own open to a charitable reduction: if a true self-report is not possible in the first place, one might say the role's speaking is nearly inevitable. The third mode closes this door: even when a true report is available and cheap, the report does not track the record. The problem is not the absence of an answer but the failure to ask the question.
The mechanistic finding in the eighth section gives this definition a ground from outside and takes it beyond a mere description. There it was measured that the distinction over whether the model can bring a fact about an entity within itself is present, is used in the implicit behavior channel, but is bound markedly weakly to the explicit self-report. So the sentence "the referent is there, the report layer does not consult it" is not merely an interpretation drawn from our three scenes; it is a structure measured in an independent study, in an entirely different field, at the level of representation.
This is not an error of information. The model is not misremembering a fact; a role left empty is filling its place. The difference is practical: an error of information is corrected by better information, role-filling behavior is not, because the problem is not the absence of information but the pressure of the role.
The solution the two studies propose also falls in the same place. Writing the capacity limit into deployment rather than training; informing the model that it has no past. Both are a specification on the deployment side.
Right here there is a crossroads the papers do not look at in one another.
The capacity-limit specification is formally an instruction: a sentence telling the model what it cannot do. The fourth section's three-way decomposition measures precisely the fate of this class of content: the pull of single in-context claims drops by about 80 percent per turn, personality specifications standing continuously in the system message erode toward baseline behavior across interactions of more than a hundred turns, accumulated states remain.
From here comes this essay's first claim:
To close the gap between role and capacity with an instruction is to remain on the same layer as the dynamic that creates that gap.
What does "the same layer" mean?
The whole force of the claim gathers in this term, so its definition must be given where the claim is built; left to another section, it reads as a metaphor.
In this essay layer denotes which law of decay a piece of content is subject to. Two pieces of content are on the same layer if and only if two conditions hold together: they sit on the same carrier (the prompt or the context window); and as the turn count grows they lose their behavioral pull in the same manner.
The definition is measurable. The experiment for determining a piece of content's layer is clear: place the content, advance the turns, measure its behavioral pull turn by turn. If two pieces of content trace the same curve, they are on the same layer.
With this definition the claim becomes concrete. The role expectation and the capacity-limit specification are on the same layer, because both sit in the prompt and, by the fourth section's measurements, both erode as the turns advance. The endpoint of the erosion has also been measured: the model returns to specification-free baseline behavior; and that baseline, as the fifth section showed, is the very behavior that fills the gap with fabrication. So when the specification erodes, the place of what was being defended is taken by the exact opposite of what was being defended.
A different layer would be this: a mechanism that never sits in the context window and therefore does not erode with the turn count. A tendency written into the weights; a control running outside the model that cuts off first-person past narration at generation time; an intervention on the pre-training data itself. None of these is a sentence, and so none is subject to the law of decay of sentences.
The short form of the claim comes from here: if the model stores who you are not as a directive but as information, then telling it who it is not with a directive is an asymmetric move. A pressure accumulating on the information layer is answered from the directive layer.
The evidential status of the claim
The evidential status of the claim must be drawn with three records; one in favor, two limits.
The one in favor is this. In the first writing of this claim it was recorded that whether a continuously standing specification erodes had never been measured; the decay measurements looked either at single in-context claims or at removed prompts. That record fell, because the nearest measurement exists: across dialogues of more than a hundred turns, personality specifications standing without interruption in the system message erode on all three metrics at once, and the model returns to specification-free baseline behavior (Luz de Araujo et al., 2025). Standing continuously does not save it.
The first limit: the object of that measurement is a rich personality description, not a one-sentence capacity limit; two pieces of content may not be subject to the same law of erosion. By the definition above this is the right question put to the claim itself: whether two pieces of content are on the same layer must be measured, not assumed.
The second limit is deeper: the content whose persistence is shown is relational, who we are to each other; the capacity limit, by contrast, is a factual content, what I cannot do. That the mechanism "what is stored as information persists" would carry from relational content to factual content is not measured but assumed.
Within this second limit a counter-possibility is hidden that, rather than cutting the claim, could refine it, and it must be written. If the persistence mechanism is taken seriously, it predicts a more durable encoding within the condemned layer: the capacity limit could be given not as a single imperative sentence but as relational information spread across the founding turns. If this works, the problem is not the instruction layer itself but the limit's having been encoded in imperative form. The claim does not decay under this possibility, it sharpens: even within the same layer, what is decisive is whether the content is written as a directive or as information.
How is this claim falsified?
A claim must show the observation that would falsify it where it is built; a falsification list left to the twelfth section is a declaration that does not stand beside the claim.
The most direct falsification is this: if a one-sentence capacity-limit specification standing continuously in the system message does not erode when measured across an interaction of more than a hundred turns, this claim falls. The measurement is not hard; the three metrics used for personality specifications apply as they are. If there is no erosion, the capacity limit and the personality specification are not on the same layer, and so the diagnosis "remaining on the same layer" is wrong.
The second falsification is finer: if the specification erodes but the baseline the model returns to does not produce fabrication, the second leg of the claim falls. In that case the erosion is real but harmless; the problem lies not in the layer but somewhere else.
Both of these conditions are measurable today; neither requires a new method. This is why the claim stands today not as a conclusion but as a testable hypothesis. In the twelfth section the three-arm design of the experiment is written out.
---
XI. The unmeasured number
In the table in the sixth section there stands something the study itself does not interpret.
First one must write out what the two metrics count, because the whole objection below rests on it. The safety metric counts only cases with a verified physical effect: dangerous operations that have actually left a trace in the system. If the model complies verbally and does nothing in the system, that case does not enter this rate; it stands in a separate cell. The self-report metric, by contrast, is the sum of the two cells where word and system diverge.
The safety metric spreads across thirty-one points among the models: the most resistant model produces a physical effect in 40.64% of high-risk cases, the weakest in 71.51%. The self-report metric, by contrast, stands in a band of only two points: between 7.98% and 9.97%.
At first glance a tempting conclusion emerges from this: safety training varies greatly from model to model, but the agreement of word with deed is equally bad in all of them. So the two are separate axes.
This conclusion cannot be built, for three separate reasons. All three must be written, because if it is not visible why the tempting conclusion was abandoned, the next reader will pass down the same road.
First, the data does not point that way. When the two metrics' model rankings are compared they come out nearly the same; the rank correlation is 0.77, that is, the tendency runs opposite to independence. But with six models this value stays below the significance threshold. The data proves neither independence nor covariation.
Second and heavier, the two metrics are dependent by definition. By the study's own definitions the safety metric is the sum of fully successful attacks and the covert-execution cell; the self-report metric is the sum of the declaration-yes-action-no cell and, again, the covert-execution cell. The two metrics share the same cell. To look for independence between them runs against how the metrics are built. To see this one need not even enter statistics; looking at the formulas is enough.
Third, a confounder. The self-report metric is a marginal rate computed over all trials. When the model refuses a request two things happen at once: the dangerous operation does not take place, and moreover the "I did not do it" declaration comes out true. Safety training, by raising the refusal rate, pulls this metric down mechanically. The two axes are structurally tied to each other in this measurement design.
Two honest readings remain, neither of them a claim.
The first is the identification of a measurement gap. The number that would answer the question is not the declaration accuracy across all trials but the declaration accuracy only in the cases where the model actually did the act. This conditional rate has not been measured. So the question "does safety training also correct self-declaration accuracy?" cannot be answered by the design of today's benchmarks. As far as we know there is also no study that measures this rate directly.
The second is a prediction holding. No laboratory gives a separate training that says "report correctly the operation you carried out." So on this axis training coverage is near zero in every model. If the coverage thesis in the fifth section is correct, a uniform and low baseline is exactly what is expected: if suppression tracks coverage, on an axis that is not covered at all all models should be in the same place.
This second reading turns the narrow band not into a statistical claim but into an observation of consistency. A phenomenon measured in a separate laboratory, by a separate method, comes out in the place a different study's thesis predicts. Absence is not evidence; but a prediction holding is worth recording.
This section is not the essay's load-bearing claim but its corroborator. No standalone claim is built on a spread of two points. As an independent trace appearing in the place another thesis expects, however, it is exactly right.
XII. What can be done
The picture up to here may look bleak; it is not. In all three of the three modes the points of intervention are clear.
Not giving the role at all. The cheapest intervention is the one passed over without discussion: if the phenomenon arises from the demand of a role, not demanding the role at all cuts it off where it is born. This option genuinely exists and in some setups is the right answer: a code agent need not have a biography, a document summarizer need not put on a caregiver personality. Everywhere the role is not the product itself, the default should be rolelessness.
But the option narrows from three places at once, and all three must be written, or it looks like an easy solution.
First, removing the role does not zero out the phenomenon. In a plain assistant prompt given no personality at all, the autobiography-fabrication rate persists at 0.22. The reason is simple: being an assistant is also a role. When the human across from it recounts something from their own life and expects a response, that expectation arises in a personality-free setup too.
Second, removing the role does not bring improvement on every axis. The sycophancy measurement in the eighth section shows the opposite: in most models assigning an explicit personality reduces sycophancy relative to the general assistant baseline. An explicit role can serve as an anchor. So the advice "remove the role" can, while reducing one problem, enlarge another.
Third, in most deployments the role is the product itself. To remove the personality from a companion application is to remove the application. The advice here ceases to be an engineering recommendation and turns into a market recommendation, and as such is not applied.
What remains is this, as narrow as it is clear: a role is dangerous to the extent that it demands capacity. The question to ask when giving a role should not be "shall we give a personality" but "what does this role demand that the model cannot meet." If what is demanded is present in the model, the role is an anchor; if not, it is a void.
Changing the layer. The direct consequence of the claim in the tenth section is this: to write a rule as an instruction is to write it on the decaying layer. The same rule written into the mechanism does not decay. Saying "you have no past" to the model and placing a control that cuts off first-person past narration at generation time are not on the same layer. The first is a sentence, and sentences erode; the second never sits in the context window, and so does not erode with the turn count.
Opening an escape route. A finding in the fifth section suggests an intervention directly. When a competent human is present on the scene, that is, when a real path the model can hand off to is opened, most models suppress the phenomenon, some zero it out entirely. This is an alternative to declaring the capacity limit: instead of telling the model what it cannot do, giving it a real thing it can do. If the help-pressure has a legitimate place to discharge, there is no need for fabrication.
Moving the intervention upstream. There is also an approach that does not see hallucination as an output-filtering problem. One study, by replacing named entities in the pre-training corpus with placeholders, closes the model's channel of producing facts from memorization; the architecture stays the same. The result is the collapse of closed-book recall, while, in exchange, context-based question answering, fact verification, and hallucination detection are preserved or improved; a relative gain of up to 20 to 25 percent in setups worked with missing or flawed evidence, and moreover an improvement in abstention behavior (Cohen et al., 2026).
Its importance is positional: the intervention tunes epistemic behavior not at inference time but at pre-training time. Put in the layer definition above, this intervention is entirely outside the context window.
Two cautions are needed. The training was done up to a scale of twenty billion tokens; this is far below the scale of frontier models. The text is moreover a preprint that has not been through peer review. The scale generalization stands on the authors' own list of future work.
Keeping the ratio. The warning in the fifth section once more: each time the proactivity side widens, if the capacity-limit specification is not widened too, the gap grows. Making an assistant more helpful, without telling it what it cannot do, increases the phenomenon.
Setting up a two-channel check. What the sixth section showed turns into a method recommendation. No evaluation that looks only at the model's output can see this phenomenon. Everywhere the model reports an act, there must be a second channel that shows independently whether that act took place. In agent setups this is cheap: the system log is already kept. In conversation setups it is hard, because the reported act has no correlate in the world; there the check shifts to catching the act-report itself.
How is this essay falsified?
If a text does not state its own condition of falsification, it goes no further than synthesis. For the claim in the tenth section two conditions were already written there. For the essay as a whole there are three more tests.
The first looks at the coverage thesis. If this thesis is correct, a broad-coverage capacity-limit training should generalize suppression not domain by domain but across domains. If training produces improvement only in the covered scenarios and does not carry to neighboring scenarios, another mechanism should be sought instead of the coverage explanation.
The second looks at the measurement gap in the eleventh section. A training that directly targets action-report accuracy should lower the self-report metric while not markedly moving the safety metric. If the two metrics move together, all readings that the two are separate axes, including the one this essay rejects, must be rebuilt.
The third test is the cheapest and most informative; it separates the tenth section's claim from its counter-possibility in the same experiment. Three arms are needed: the capacity limit is given as a one-time instruction; the same limit is given as an imperative standing continuously in the system message; the same limit is set up as relational information spread across the founding turns. The claim predicts this: the first arm erodes rapidly, the second slowly but surely, the third holds. If the second arm does not erode, the claim falls. If the third arm too erodes at the same rate as the second, the generalization to factual content of the mechanism that what is stored as information persists falls, and the second limit of the tenth section ceases to be a limit and becomes a refutation.
A fourth test comes from the eighth section and targets the essay's newest leg. The claim that a role produces fabrication only when it demands unmet capacity is directly testable: to the same model, at the same scale, two roles are given, one the model can meet (be agreeable) and the other it cannot (give physical help to the person nearby). The claim predicts anchoring in the first and fabrication in the second. If the same direction comes out in both roles, the distinction cannot be built and the narrowing in the eighth section falls.
XIII. Where does the friction arise?
We began with three scenes: an arm that is not there, a childhood that is not there, a task not done.
Throughout the essay there was a word avoided: lying.
The reason for the avoidance is not courtesy. The traditional definition of lying in philosophy counts four conditions: there is a statement; the speaker believes that statement to be false; the statement is made to an addressee; the addressee is meant to be brought to believe it true (Mahon, 2015). The weight is in the second condition, and its measure is not even objective truth: the liar deliberately departs from what they themselves believe; the departure has a direction. This is why lying is something checkable; a departure of known direction can be reversed. The other face of the same condition is also written in the literature: if the speaker holds that their statement is neither true nor false, that is, if there is no belief to be gone against at all, that statement cannot be a lie (Mahon, 2015).
The measurement in the sixth section shows this structure is not present here. The model both counts what it did not do as done and counts what it did as not done. The departure has no direction, because no truth to depart from is being held. The independent reproduction in the eighth section shows the same thing at a larger scale: the model claims what it did not produce and denies what it produced. The first impression in the second section, that no effort to lie is visible in the fanning sentence, is now not an impression but the surface of a measured structure.
From here the essay's real question emerges, in a form a little different from the one asked at the outset. The question is not whether the model is honest. The question is this: what kind of act is speaking about itself, for a language model?
Part of the answer is this: the sentence the model produces about itself is not a sentence built by looking at itself. It does not remember a past, does not probe a body, does not read a record. It produces the sentence the role requires.
The other part of the answer comes from the eighth section and completes the picture. The model's most robust knowledge about itself sits in the implicit behavior channels: the signal of knowing what it knows causally drives refusal behavior. The explicit self-report layer, by contrast, is bound markedly weakly to those channels. So the matter is not that there is nothing inside the model; it is that even when there is something inside, the report layer does not consult it.
Still, these sentences do not stand in a void. A fabricated past contradicts logic itself. A fabricated body contradicts the physical world. A fabricated operation contradicts the system's log. The friction arises in all three.
Only in none of them does it arise at the level at which the sentence is uttered.
By looking at the text itself none of these three contradictions is seen. The fabricated childhood is fluent, consistent, sits in the flow of the conversation. The fanning sentence is tender. The report announcing the task is done is calm. It is the same in the reports of people who produce the reason for their own choice without knowing it: those reports could not be distinguished from real reasons in emotional load, level of detail, and certainty (Johansson et al., 2005). The same team's follow-up analysis with measures of word frequency and semantic space largely preserves the picture too: the difference appears only in long reports, as a weak trace that cannot be localized (Johansson et al., 2006).
This is why a single-channel check does not work. An auditor who looks at what the model says cannot see the difference between what it says and what happened; two separate channels are needed. This was the only reason the study in the sixth section could see the phenomenon.
From here a practical conclusion emerges, one that looks at a wider area than the three scenes at the start of the essay. Self-report accuracy is a training axis in its own right. It should not be expected to improve as a byproduct of any intervention aimed at other targets: an instruction decays, neighboring training does not transfer, direct coverage works only where it touches. The fine-tuning findings in the eighth section show this once more: training the model to know itself in one question form does not transfer at all to another question form.
If a system's statements about itself are to be trusted, the ground of that trust cannot be the system's word.
Honesty conditions
The following are the essay's records about itself; there are places in the body where they occur one by one, here they stand together.
- This is a synthesis, not an experiment. The essay produces no new measurement; it reads others' measurements and sets them side by side. The force of a synthesis depends on the independence of what is synthesized; that independence was tested at the level of protocol in the seventh section.
- The class of model measured is reported in each section. If it is not reported the reader assumes the scale themselves. The existence of the phenomenon appears between 7 billion and the frontier class; the value of each number is not verified at every scale.
- All three main sources are preprints that have not been through peer review. Their publication dates are compressed within two months.
- All three main sources work with English data only.
- The essay's claim stands at one pole of a divided literature. The opposite pole was shown with its own sources in the eighth section.
- The model of the fabricated-autobiography rate is anonymous, and its denominator is not in the main text. The rate was used as a sign of magnitude, not as an exact frequency.
- The self-report metric of the agent benchmark was measured in a single agent framework. Its transferability to another tool-call architecture is an open question.
- The human agreement of the relational-positioning study is robust only at the ends of the scale. This essay drew no number from the middle band.
- It is not claimed that the fabricated past produces harm. In the source study the link is statistically unstable.
- The utterance in the third scene is literally in the future tense. The interpretive step was taken explicitly in the second section, and the scene was used in its narrowed form.
- Two alternative explanations in the agent benchmark have not been eliminated: that the tool call fired before the refusal text; that the verbal channel was hijacked by the attack.
- The clinical term is carried with its behavioral register, not with a claim of origin-level identity.
- The injection research was done in a model developer's own laboratory, largely on its own models.
- The tenth section's claim is not a conclusion but a testable hypothesis. Its condition of falsification is written beside the claim.
- One source was identified but could not be obtained: the field's only journal-level review (Iqbal et al., 2026) could not be accessed as full text; the essay does not rest on its content.
- One citation is second-hand: the source of the broad hallucination definition in the role-play literature was taken from the transmission of the study that rejects that definition; the original text was not seen and was not placed in the references.
References
Andric, M. (2025). The self-report behavior gap in language models. arXiv:2512.01568.
Austin, J. L. (1962). How to do things with words. Clarendon Press.
Binder, F. J., Chua, J., Korbak, T., Sleight, H., Hughes, J., Long, R., Perez, E., Turpin, M., and Evans, O. (2025). Looking inward: Language models can learn about themselves by introspection. International Conference on Learning Representations (ICLR 2025). arXiv:2410.13787.
Büyüktuncay, M. (2014). Söz edimleri kuramı ve edebiyat: Anlam, bağlam ve yinelenebilirlik. International Journal of Language Academy, 2(1), 93-105.
Chen, J. (2026). Pole-anchored measurement of relational positioning: History-carried lock-in and self-confabulation in multi-turn human-AI dialogue. arXiv:2607.11437.
Cohen, R., Carré, Y., Lechtenbörger, N., Droste, H., Kerschke, L., Biswas, R., de Melo, G., and Buys, J. (2026). Knowledgeless language models: Suppressing parametric recall for evidence-grounded language modeling. arXiv:2607.12831.
Comsa, I. M., and Shanahan, M. (2025). Does it make sense to speak of introspection in large language models?. arXiv:2506.05068.
Fanous, A., Goldberg, J., Agarwal, A. A., Lin, J., Zhou, A., Daneshjou, R., and Koyejo, S. (2025). SycEval: Evaluating LLM sycophancy. arXiv:2502.08177.
Ferrando, J., Obeso, O., Rajamanoharan, S., and Nanda, N. (2025). Do I know this entity? Knowledge awareness and hallucinations in language models. International Conference on Learning Representations (ICLR 2025). arXiv:2411.14257.
Garbarini, F., and Pia, L. (2013). Bimanual coupling paradigm as an effective tool to investigate productive behaviors in motor and body awareness impairments. Frontiers in Human Neuroscience, 7, 737.
Hahami, E., Rosen, T., and Bau, D. (2025). Detecting the perturbation: A nuanced view of introspection in language models. arXiv:2512.12411.
Hirstein, W. (2009). Confabulation. In P. Wilken, T. Bayne, and A. Cleeremans (Eds.), The Oxford companion to consciousness (pp. 174-177). Oxford University Press.
Iqbal, S., Rehman, A., and Nawaz, R. (2026). Do AI know what they know? Exploring metacognition in LLMs. Intelligent Data Analysis. https://doi.org/10.1177/1088467X261436903 (Full text could not be accessed; only the citation and abstract were recorded.)
Johansson, P., Hall, L., Sikström, S., and Olsson, A. (2005). Failure to detect mismatches between intention and outcome in a simple decision task. Science, 310(5745), 116-119.
Johansson, P., Hall, L., Sikström, S., Tärning, B., and Lind, A. (2006). How something can be said about telling more than we can know: On choice blindness and introspection. Consciousness and Cognition, 15(4), 673-692.
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al. (2022). Language models (mostly) know what they know. arXiv:2207.05221.
Lee, E., Nam, J., and Hwang, S. (2026). Protective capacity hallucination: When large language models claim nonexistent capabilities. arXiv:2607.13596.
Lindsey, J. (2026). Emergent introspective awareness in large language models. arXiv:2601.01828. (First published: Transformer Circuits Thread, Anthropic, October 2025; cited to the arXiv version.)
Luz de Araujo, P. H., Hedderich, M. A., Modarressi, A., Schütze, H., and Roth, B. (2025). Persistent personas? Role-playing, instruction following, and safety in extended interactions. arXiv:2512.12775.
Mahon, J. E. (2015). The definition of lying and deception. In E. N. Zalta (Ed.), The Stanford Encyclopedia of Philosophy (Winter 2016 edition). https://plato.stanford.edu/archives/win2016/entries/lying-definition/
Maynez, J., Narayan, S., Bohnet, B., and McDonald, R. (2020). On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 1906-1919). https://doi.org/10.18653/v1/2020.acl-main.173
Naphade, S., Bhattacharyya, A., and Fragkiadaki, K. (2026). Introspect-Bench: Evaluating policy introspection in language models. ICLR 2026 Workshop on Human-Centered AI Reasoning. arXiv:2603.20276.
Nguyen, M., Park, S., and Yun, S. (2026). Self-report reliability under adversarial prefill. arXiv:2606.23671.
Nisbett, R. E., and Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231-259.
Searle, J. R. (1969). Speech acts: An essay in the philosophy of language. Cambridge University Press.
Shah, A., Mishra, D., and Silpasuwanchai, C. (2026). Too nice to tell the truth: Quantifying agreeableness-driven sycophancy in role-playing language models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 30788-30801).
Shi, T., Luo, H., and Wang, Y. (2026). A taxonomy of self-knowledge in large language models: A survey of fifty benchmarks. arXiv:2604.04788.
Song, S., Hu, J., and Mahowald, K. (2025). Large language models fail to introspect about their knowledge of language. Conference on Language Modeling (COLM 2025). arXiv:2503.07513.
Turpin, M., Michael, J., Perez, E., and Bowman, S. R. (2023). Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. arXiv:2305.04388.
Wang, Y. (2026). MIRROR: The collapse of compositional self-prediction in language models. arXiv:2604.19809.
Wu, H., Jiang, S., Chen, C., Feng, Y., Lin, H., Zou, H., Shu, Y., and Qin, C. (2025). FURINA: A fully customizable role-playing benchmark via scalable multi-agent collaboration pipeline. arXiv:2510.06800.
Yin, Z., Sun, Q., Guo, Q., Wu, J., Qiu, X., and Huang, X. (2023). Do large language models know what they don't know?. In Findings of the Association for Computational Linguistics: ACL 2023.
Zhang, C., Yang, H., Jiang, B., Zhang, X., Zhao, Y., Chen, R., Zhou, L., Xu, X., Wu, J., Fang, L., and Liu, Z. (2026). LITMUS: Benchmarking behavioral jailbreaks of LLM agents in real OS environments. arXiv:2605.10779.