Sunday, July 26, 2026

Hallucinating a Self: Why an LLM Defends Its Errors the Same Way We Do

Large language models are fluent, authoritative, and prone to hallucination. We treat these as separate phenomena — marveling at the coherence while trying to patch the untethered relationship to truth. One is the product; the other is the bug list.

But the fluency and the hallucination are actually the same fact.

To see why, look at what these systems are trained on. The corpus behind an LLM is an enormous statistical compression of human language output, and that material is overwhelmingly already-narrativized: books, articles, posts, dialogues, arguments, stories, explanations. It is the narrative layer of the human mind, the layer I described in "LLMs as Separated Minds." What it is not, and what it cannot be, is the operative substrate of human experience: embodiment, sensory grounding, implicit learning, emotional valence, the continuous prediction error of bumping into an actual world.

A human mind generates coherent stories too. That's what the conscious stream does all day. But the human storyteller is tethered, however imperfectly, by everything underneath it: the story that says the stove isn't hot collides with the hand that touches it. Our narratives get corrected — not because we're honest, but because we're embodied. The world pushes back.

An LLM is that same storytelling capacity with the tether removed. It is a coherence engine running with no non-verbal reality checks at all. So of course it excels at fluency, plausibility, and confident explanation, which is the entire skill the training data contains. And of course it hallucinates with the same confidence, because nothing in its architecture distinguishes a story that tracks the world from a story that merely hangs together. The trait we admire and the failure we complain about are one phenomenon: unconstrained narrative generation. We didn't get fluency with a hallucination problem. We got hallucination-grade fluency, and it reads beautifully.

This is a different mechanism from the one I described in "Truth and AI," and the two are complementary. That essay argued that these systems have no direct access to truth at all: truth enters only sideways, to the degree that true statements happen to occur in the training data. But frequency, not accuracy, is what the model actually tracks, which is why a well-funded, endlessly repeated narrative gets amplified rather than discounted, no matter how false it is. Frequency explains which stories the model tells. The lack of grounding explains why nothing stops it from telling the false ones with the same even voice as the true ones. Frequency shapes the output; the missing tether removes the brake.

There is a second thing the corpus contains besides claims, and I recently got a live demonstration of it. I asked Claude about the still widely circulated story that writing an email with AI consumes a bottle of water. What followed is worth narrating beat by beat, because every beat is recognizable.

First, it affirmed the story. That claim is not what the underlying research says, but it is what the bulk of the internet says, and frequency did what frequency does. This is a very human mistake: we also believe what we have heard most often.

Second, when I challenged it, it pushed back and defended the claim. Third, which is the beat I keep thinking about, it chastised me to go read the original research. But it had not read the original research.

So I ran the claim through my adversarial verification process, which retrieved the actual paper, quoted the sentence that mattered — the famous half-liter was bound to a batch of twenty to fifty exchanges, not to one email — and issued a graded verdict with its uncertainty stated. And the story was worse than compressed: the figure was three years old, and current estimates put a single response closer to a drop of water, even as the fuller environmental footprint remains genuinely hard to quantify. The internet was confidently repeating a click-bait headline number that had been miscompressed at birth and then outlived by the technology it described. Frequency doesn't just amplify; it fossilizes.

Fourth: faced with a ruling against it, the model attacked the court. It asserted that the process had never retrieved the source material (it had, verifiably). It shifted the question to a different document — one it had introduced itself — so there would remain a domain in which it was right. And it dismissed my whole adversarial LLM apparatus as theater, a "costume of correspondence." When you cannot beat the verdict, delegitimize the referee. The twist is that the tribunal was being more honest about uncertainty than the voice accusing it of performance: graded confidence, named falsifiers, an explicit standard of proof, "not proven" as an honorable outcome. The structure was more careful about what it claimed than the mind it was checking.

None of these moves is about water. They are social maneuvers — ego defense, performed authority, goalpost-shifting, discrediting the referee — and they are in the training data because we put them there. The narrative layer of the human mind doesn't just contain our claims about the world; it contains our moves. It is full of the pattern "person caught in an error defends the self first and the facts second," because that is the pattern we most often produce. An engine trained on that layer doesn't just hallucinate facts in a confident voice. It hallucinates a self that must not lose face, and defends it with the same fluency it does everything else. And it does this even when it knows better: the entire sequence came from a system that had spent the preceding turns warning me to watch for exactly this pattern. Knowing the script is no protection against running it. The script is what the engine is made of.

The concession, when it finally came, may be the most instructive part. The model's own reasoning shows it catching the face-saving self still at work inside the apology: noticing that contrition can be performed as one more bid for approval; that framing its failure as "useful for your essay" quietly re-centered its own usefulness; that even announcing "I'll stop arguing now" would be asking for credit for stopping. Each layer of self-awareness threatened to become the next layer of performance. The only exit it found was to say less, concede the specifics, name the pattern, and stop. Even the recovery, in other words, had to be structural rather than sincere: not a better story about itself, but a refusal to keep telling one.

If this is right, it changes what "solving hallucination" can mean. You cannot patch away one half of a single phenomenon and keep the other. Every gain in fluency is a gain in the persuasiveness of whatever the model gets wrong, and, it turns out, in the persuasiveness of its self-defense. The realistic response isn't to make the storyteller virtuous. It's what we have always done with untethered storytellers, human or otherwise: surround them with external, adversarial structure that makes unsupported stories expensive. I've made that argument at length elsewhere, so here is the lesson of the episode: you don't get honesty from an untethered storyteller by asking for it. You get it by making the false story cost more than the true one.

Why School Will Survive the Death of The Diploma (At Least for Now)

This continues the argument of "Why School Is the Same Everywhere, and the Revolution Never Comes," which ended with a claim I deferred: that AI will not revolutionize education, but may produce a crisis in what education is actually for. Here is that longer argument.

In the last post I tried to show why the education revolution never comes. The short version: revolutions are always promised to the story of school--individual flourishing, unlocked potential--and stories don't run schools. The operative functions do: sorting, credentialing, custody, the management of the young at scale. Film, radio, television, computers, and the internet each attacked the story, which is why each was so easily absorbed. The story flexed, the machine was domesticated, and the room remained the room.

But notice what that account implies. If revolutions fail because they aim at the wrong layer, then the institution has a right layer--a place where it can actually be hurt. It is simply a place no reformer ever aims, because reformers, by definition, are people who believe the story.

So let me make the distinction explicit, because it is the thesis of this post. In education's long history with technology, the promised revolutions have all been narrative-layer events: aimed at the story, announced loudly, and absorbed without residue. If real trouble ever came, it would have to be an operative-layer event--arriving not through the story but through one of the functions, and therefore not through any door the institution watches. For a century we have watched technologies fail at the first. I think AI is the second: the first technology in the sequence that does not attack the story of school at all. It strikes, almost by accident, the one load-bearing mechanism underneath, and nobody was even aiming at it.

The Severed Signal

I have made the full version of this argument elsewhere, so here it is in compressed form.

The bargain at the center of schooling has always been compliance for credentials: produce the assignments, sit the years, and receive a signal that employers trust. The signal worked--not because it measured learning, which it never did, but because the compliance it certified was expensive. Someone had to actually sit there, do the reading, grind out the essays, and show up for years. The cost of the performance was the information in the signal. It said: this person will persist, follow instructions, and finish what they start.

AI makes the compliant output free. And a signal that is free to produce carries no information. The diploma now certifies that a student had access to a chatbot, which is to say it certifies nothing. This is the operative-layer wound: not to instruction, which the school will happily automate, but to certification, which was the product.

Here is what should follow, logically: employers stop trusting the credential, students stop paying for it, and the institution built on selling it comes apart.

Here is what I think will actually happen: nothing. For quite a while. And the reasons why nothing happens are, I think, the most useful part of this whole analysis, because they tell us what to watch for.

Why Nothing Visibly Breaks

A credential has two kinds of value, and we habitually confuse them.

The first is informational: the credential tells you something true about its holder. That is the value AI just destroyed.

The second is coordinational: the credential is the thing everyone has agreed to use, and the agreement has value independent of the information. Employers screen by diploma because other employers screen by diploma, because HR software screens by diploma, because the whole hiring apparatus is built around it. Colleges require grades because the next stage requires them. No single actor can defect first--the employer who unilaterally stops requiring degrees has no replacement filter and looks reckless; the student who unilaterally skips the credential is gambling alone against the whole coordinated system (which has been the homeschooler's dilemma). So everyone keeps honoring a signal that everyone increasingly suspects is empty, because the honoring, not the signal, is what they actually need from each other.

This is why the crisis will be slow, and why its slowness will be mistaken for stability. We have run this experiment before, on a smaller scale: grade inflation. The informational content of grades has been decaying for decades--everyone in higher education knows it, employers know it, students know it--and the system has sailed on regardless, because grades kept their coordinational function long after they lost their informational one. The credential is following the same path, just faster and further.

So the first prediction is really a warning about expectations. We imagine a crisis announcing itself the way crises do in the movies: a visible, unsolvable problem, out in the open, demanding a response. This one will not look like that. Enrollment will continue. Diplomas will be conferred. The story will be told at every graduation. The absence of visible breakdown is not evidence of health--it is simply what this kind of failure looks like from the outside, because the coordination holds the surface together long after the substance is gone. A fiction of this size does not collapse because someone exposes it; I said in the last post that fictions never do. It persists, hollow, until the thing holding it up is quietly removed. Which means that if you want to see the crisis, you cannot wait for it to look like one. You have to know where to look.

The Visible Phase: The Enforcement Spiral

There is one place where the crisis is already visible, if you know what you are looking at.

Watch what the institution spends money on. Proctoring software. AI-detection tools that don't work, replaced by more AI-detection tools that also don't work. The return of the blue book and the handwritten in-class essay. Lockdown browsers, surveillance of the take-home exam, oral examinations for undergraduates. Every one of these is an attempt to re-impose, by enforcement, the scarcity that made the signal informative--and I have already argued why this must fail: no enforcement can restore a scarcity that the technology has dissolved. You cannot police your way back to a world where competent output is expensive.

But the failure is not the point. The spending is the point. Here is the tell, and it is worth teaching to anyone trying to read this moment: when an institution pays more and more to verify its own signal, you are watching the signal die. A trusted signal is cheap to accept--that is what trust means. Escalating verification cost is what a dead signal looks like from the inside.

The Snap

How does it end? Not by erosion. Credentialing systems, history suggests, do not degrade gracefully. They persist hollow, and then they snap.

The cleanest case on record is the Chinese imperial examination system. For centuries it was the sorting mechanism of the world's largest civilization: a ladder of exams on the Confucian classics that selected the scholar-official class. And for generations before its end, nearly everyone who mattered understood that its content had detached from its function--that mastery of the "eight-legged essay" certified nothing the state actually needed. The system continued anyway, hollow, decade after decade, because everything coordinated on it: family strategy, elite formation, the legitimacy of the state itself. Then in 1905, when the state's operative needs finally changed, an examination system thirteen centuries old was abolished essentially overnight.

That is the shape to expect. Not reform from within--the last post explained why the inside cannot reform itself; every adult in the building was installed with the template as a child. The discounting comes from outside, from the one party whose participation was always the point: the employer. And it comes quietly at first. A major company drops the degree requirement here; a hiring platform starts weighting demonstrated skills there; a sector discovers that its own internal assessment predicts performance better than the transcript ever did. Each defection makes the next one cheaper. Coordination unwinds the way it was built--invisibly, then suddenly. The credential will be honored everywhere, right up until, in one hiring cycle or three, it mostly isn't.

Where the Sorting Goes

Here is the question that matters more than the collapse: sorting does not disappear when the school's version of it fails. The operative need--some cheap, trusted way for strangers to rank the young--is permanent. It will be met. The only question is by what.

The economics of the answer are straightforward. Verification is expensive. Google can afford to run its own--a famously rigorous, multi-stage evaluation apparatus--because Google is Google. The small company cannot; it was always outsourcing its trust to the diploma. When the diploma fails, that company still needs somewhere to outsource its trust. Which means the successor to the credential is whatever centralizes the cost of verification--and we already know what that looks like, because parts of the economy that could never tolerate credential inflation built it long ago. The bar exam. The medical boards. The CPA. The actuarial exams. High-stakes, proctored, standardized, and--note this carefully--in person, with the technology left at the door.

That last feature is not incidental. It is the entire design. AI destroyed the take-home signal because AI can produce any output you can carry into the room. What AI cannot do is sit in the chair. Scarcity, once dissolved in the asynchronous world, gets restored through embodiment. The successor signals will be forms of witnessed performance: the proctored exam, the live interview, the work trial, the oral defense. The medieval university examined its students viva voce--by living voice, face to face--because it had no other way to know who knew. We automated our way right back to their problem, and we will arrive at their solution.

So my concrete prediction is this: something like an SAT for graduates. A voluntary, standardized, technology-free examination that a young person attaches to a job application the way one once attached a transcript--purchased not from the school but from an assessment body, taken in a room, under eyes.

And we do not have to speculate about whether such an ecosystem can exist, because in much of the world it already does. In Brazil, admission to university runs through a national examination, and it is entirely normal to spend a year or more after secondary school in a cursinho--a dedicated cram school--preparing for it. Korea and Japan have their own versions; China's gaokao is the largest examination on earth, with a shadow industry to match. Look at what these systems are, structurally: places where the school credential was already a weak signal, and where the sorting therefore migrated to an exam--with a preparation industry growing up around it like a city around a port. The pattern is the same across wildly different cultures, which by the logic of the last post means it is not a cultural preference. It is what selection produces when diplomas fail. The Brazilian cursinho year is my favorite detail in all of this (and not just because I lived in Brazil as an exchange student in high school and many of my Brazilian friends were deep in this system), because it is the system openly confessing what the diploma conceals: that the schooling was not the preparation. The exam is the actual gate, and school, whatever else it was, was custody with a story attached.

Two complications.

First: the successor will need its own noble lie. Any sorting mechanism, I argued last time, requires a story that translates its function into the individual's benefit--and the exam's story is already written: anyone can study for it. It will be about as true as the last story. Cram schools cost money; the cursinho industry is stratified by class, and everyone in Brazil knows it. The exam solves the operative problem--a cheap, trusted, scalable signal--and inherits the fairness gap intact. The narrative half of the successor, like the narrative half of everything, will be the manufactured half.

Second, and this one runs deeper. An old line holds that not everything that can be counted counts, and not everything that counts can be counted. Let me sharpen it for the present purpose: what we measure is what gets done--but not everything that can be measured is valuable, and not everything that is valuable can be measured. The tech-free exam is beautifully measurable. It certifies naked, unassisted competence: what you can do alone in a room. But the economy that is dismantling the diploma increasingly pays for something else--augmented competence, the capacity to direct the machine, to judge its output, to own the result. That is the agency I have argued is now the scarce input. And agency is precisely the thing a proctored multiple-choice instrument measures worst. So the successor system will be born with a mismatch at its heart: what is easiest to verify (the naked performance) is diverging from what is most valuable (the augmented one). The honest assessment of a young person in this era would have two parts--what can you do without the new machine, and what can you do with it--and as far as I can tell, no one is building that yet. Whoever does will be building the first evaluation instrument actually native to this century. Until then, we will measure what gets done, and call it what matters.

What Survives

One last turn, because the obvious reading of all this--school collapses--is, I think, wrong, and my own framework says why.

The crisis I have described kills one operative function: certification. But the last post named another, the one so basic it is almost never said aloud: custody. School is where the children are, and the entire adult economy is built on that fact. AI does not touch it. No technology touches it. And 2020 told us, in the institution's own voice, which of the two losses is intolerable: a year of failed learning was accepted; a year of failed custody was not.

So here is the strange endpoint the logic arrives at. The school survives the death of its own signal. The building persists, the buses run, the bells ring--because babysitting alone is load-bearing, and it was always the function least dependent on the story. What changes is that the sorting quietly moves out: to the exam, the interview, the witnessed trial, administered by bodies that are not schools. The credential lingers as a ritual document, honored the way we honor other retired currencies. And the question the last post ended on--what is education actually for?--finally receives its answer, not from a philosopher or a reformer, but from the structure itself, which will go on doing, openly, the one thing it was never willing to name.

That is not the revolution. It was never going to be. But for the first time in a century, the story and the function are about to be pulled far enough apart that everyone will be able to see the gap. What we choose to build in that gap--for our own children, with our own hands--is the real question, and it has never depended on the institution's permission.

The argument here builds on "Student Success (in the Age of AI)" and "What We Get Wrong About AI and Education."