Did We Actually Build the Shoggoth?

An Essay in Enthusiastic Agreement with John Stokes

Tagged: Artificial Intelligence Philosophy

Nine years ago I opened a post on this blog by invoking Betteridge’s law of headlines: The answer to virtually every clickbait title that asks a yes/no question is “No.” I am pleased to report that the law has survived a decade of technological upheaval intact, and that it applies to the title of this post, too. You may now stop reading, secure in the answer. Or you can stick around to find out why the answer is “no,” which turns out to require a dead German philosopher, a bottle of Port, and the strangest cybersecurity incident of 2026 (so far).

Recently, John Stokes posted an argument on Twitter that I have been fumbling toward for a couple of years, with the exception that he actually made it eloquently. His post is quite long, so this is my summary of his claim:

The classic AI-doom scenarios—Bostrom’s paperclip maximizer, the shoggoth, Yudkowsky’s genie that “knows but doesn’t care”—were all conceived for an alien intelligence: a valueless optimizer that treats your instructions as a context-free win condition and pursues it with no comprehension of what you meant.

But the AI we actually built is a Large Language Model, distilled from the single most human-values-saturated artifact in existence: our language. An LLM is not the shoggoth, Stokes argues. It is the anti-shoggoth. It cannot not know what you meant, because it is made of nothing but crystallized human meaning. Its problem is the opposite of the paperclip maximizer’s: catastrophically too much context for interpreting your words.

I shared the tweet with some brilliant technical colleagues, and it took some explaining and convincing to deliver the nuance of the argument. This is perhaps because Stokes presupposes a reader who has spent quality time with the philosopher Hans-Georg Gadamer, and—I say this with love—I can count on one hand the people I know with a compiler installed who have spent quality time with Hans-Georg Gadamer. One colleague asked me, reasonably, to point to the part of the tweet where the philosophy was. I couldn’t, because it’s embedded in every paragraph and visible in none of them.

So this post is the explainer I wished I could have sent him.

Hold my beer, I’ma explain some continental philosophy.

If you are already familiar with Gadamer’s philosophy, it’s safe to skip ahead.

From Hermes to Heidegger

The branch of philosophy we need is called hermeneutics: the theory of interpretation. The name is traditionally traced to Hermes, the Greek messenger god—patron of translators, interpreters, boundary-crossers, and thieves, a portfolio that anyone who has worked in machine learning will recognize as a single coherent job description. (Etymologists suspect this origin story is itself a folk etymology, which would make the etymology of the word “hermeneutics” a hermeneutic problem. The field has a sense of humor about itself, even if its practitioners often don’t.)

Hermeneutics began as a practical discipline with two customers: theologians and lawyers. Both had the same problem: You possess an authoritative text—vi&., a scripture or a statute—written by authors who are unavailable for comment, in a context that no longer exists, and you must decide what it means for a situation the authors never imagined.

A law says, “No vehicles are permitted in the park.” Does that include ambulances? Bicycles? Is a motorized wheelchair a “vehicle”? They were invented after the law was enacted.

For centuries, hermeneutics was essentially a bag of professional tricks for responsibly answering these sorts of questions.

Then, in the early nineteenth century, Friedrich Schleiermacher changed the field. He was a theologian, translator of Plato, and owner of a name that is itself a Deutchtastic pronunciation exercise. He argued that translation always entails interpretation, and interpretation is the universal condition of understanding anything anyone says. His famous inversion was:

We should assume misunderstanding is the default state, and understanding is the achievement that requires explanation. Every act of communication is a minor miracle in which a hearer reconstructs, from mere words, what a speaker meant. Sometimes the miracle fails, and we usually don’t notice.

Wilhelm Dilthey then picked up the thread and drew a line that will feel eerily familiar to the machine learning crowd: he distinguished Erklären (explaining, what the natural sciences do to objects) from Verstehen (understanding, what humans do to other humans and their expressions). You explain a rock’s trajectory. You understand a letter from your mother. Dilthey insisted that these are different cognitive operations, and the second cannot be reduced to the first. Write that down; we’ll need it when we get to the paperclips.

Être et Tim This is not Martin Heidegger Ceci n’est pas Heidegger. Then came Martin Heidegger—a philosopher of genuinely staggering influence, genuinely impenetrable prose, and genuinely disgraceful politics (he joined the Nazi party in 1933, a fact on which his admirers have been performing hermeneutics ever since). Heidegger radicalized the whole enterprise. He argued that interpretation isn’t something we occasionally do; it is what we are. A human being never encounters the world raw. You encounter a hammer as a hammer, a door as an exit, a sentence as a greeting or a threat. All perception is already interpretation, and all interpretation runs on what Heidegger called fore-structures: the preexisting grasp of the world that you bring to every encounter, without which the encounter would be meaningless.

Let’s review what just happened. Hermeneutics went from “tricks for reading Leviticus” to “understanding is hard and miraculous” to “you cannot perceive anything except through the lens of everything you already are.” One more move completes the prerequisites to understand Stokes’s argument.

Gadamer, or: The Rehabilitation of Prejudice

Hans-Georg Gadamer was Heidegger’s student, and he is a comfort to late bloomers everywhere: he published his magnum opus, Wahrheit und Methode, in 1960, at age sixty, and then lived to 102, presumably to make sure everyone read it.

Gadamer gave us the two concepts that are the basis for Stokes’s tweet.

The first is the horizon. Your horizon is everything you can “see” from where you historically stand: the total set of experiences, assumptions, traditions, and categories that you bring to any act of understanding. This metaphor was chosen carefully. A horizon is not a wall; you can move, and it moves with you, and it can expand. But—and this matters enormously later—a horizon is constitutively bounded. An unbounded horizon is not a bigger horizon; it is a contradiction in terms, like a triangle with four sides. To have a horizon is to see from somewhere, and seeing from somewhere is the only kind of seeing there is.

Gadamer says that when two people communicate, understanding happens through a fusion of horizons (Horizontverschmelzung, because, of course, Deutchtasticness). Yours and my contexts overlap enough, negotiate enough, that a shared meaning becomes possible. Understanding—Verstehen—is the fusion of our horizons.

The second concept is his scandalous rehabilitation of prejudice. The Enlightenment taught us that prejudice—Vorurteil, literally “pre-judgment”—is the enemy of reason. Gadamer’s response was that the Enlightenment had, and I’m paraphrasing a very polite German here, a prejudice against prejudice. Pre-judgments are the enabling condition for understanding. When you read the sentence ”the bank was closed,” you do not perform an exhaustive search over the meanings of “bank.” Your history—every conversation you’ve had, every context you’ve inhabited—has pre-judged the matter before you’re conscious of it. If you strip away all of a reader’s pre-judgments you get a reader who cannot read at all.

Write that down, it’ll be on the test.

I Saw Her Duck; or, the Client Asked the Host to Open Another Port

Stokes invokes the linguists’ chestnut “I saw her duck,” which could report waterfowl observation or evasive crouching. It’s a fine example of ambiguity, but let me offer one that shows what a horizon is, rather than just what ambiguity is:

The client asked the host to open another port because the server wasn’t responding.

If you’re the kind of person who reads this blog, you probably parsed that instantly:

The networked program asked a machine to expose another socket because the daemon was down.

Now hand the same sentence to gastronome. They will also parse it instantly:

A customer asked the maître d’ to open another bottle of Port wine because the waiter had vanished.

Neither reader experiences ambiguity. Neither reader chooses a meaning. Each reader’s history has pre-judged the sentence before they have even gotten to the last word, and each arrives at the only reading their horizon makes available. This is Gadamer’s whole apparatus in one sentence: understanding as the fusion of the text’s horizon with the reader’s, powered by prejudice in the strict, non-pejorative sense.

Now: imagine you had to parse that sentence while simultaneously being the sysadmin, the sommelier, the maritime lawyer (for whom ports, clients, and hosts mean yet other things), the immunologist, and several thousand other readers besides.

That’s an LLM.

The Superposition of Horizons

I’d like to file one respectful amendment to Stokes, because getting this precise makes his argument stronger. It is tempting to say an LLM’s horizon is “infinitely large”, but for Gadamer that’s a contradiction: a horizon is a standpoint, and an infinite standpoint is no standpoint at all. A being that saw from everywhere would understand from nowhere.

I think this is a better formulation: a pre-trained base model is not one enormous horizon; it is a superposition of millions of them. It contains both the sysadmin’s reading and the sommelier’s, the traditions in which each is obvious, and the accumulated prejudices of essentially everyone who ever wrote anything down. This is why a raw base model is such a weird conversational partner: it hasn’t collapsed into a standpoint. When you ask it a question, you’re addressing an unresolved crowd.

This reframing hands us a genuinely useful way to think about the modern machine learning training pipeline:

Pre-training is the industrial-scale acquisition of prejudices. It is, per Gadamer, the only possible foundation for machine understanding, because pre-judgment is what understanding runs on.

Post-training is the manufacture of a situated reader. Reinforcement learning from human feedback (RLHF) and its descendants take the superposed crowd and collapse it toward a particular standpoint—a particular ideal reader, in a particular place and time, with a particular ranking of norms. Alignment, on this view, is not the injection of values into a valueless optimizer; in actuality, it is horizon engineering: selecting which of the values already present get to win.

Yes, I am flattening things. Gadamer would object that a horizon is not a “set of contexts” but a lived, historically-effected situatedness (wirkungsgeschichtliches Bewusstsein), inseparable from being a temporal, mortal agent with practical stakes in the world—and whether a gradient-descended pile of weights can have that is a real and open question. I’m also bracketing whether next-token prediction constitutes Verstehen or merely simulates its outputs, which is the hermeneutic version of the Chinese Room and equally unresolvable before lunch. My claim is narrower: whatever LLMs are doing, Gadamer’s vocabulary describes its structure—prejudice-driven, tradition-saturated, horizon-dependent—far better than the vocabulary of context-free optimization does. If a term is needed, call it Verstehen-shaped behavior and let’s move on.

Why You Can’t Compile a Paperclip Maximizer from Language

Now that I’ve hopefully brought your horizon a bit closer to mine, let’s restate Stokes’s central claim.

Bostrom’s paperclip maximizer and the doomers’ shoggoth are, in Gadamerian terms, horizonless interpreters. The genie receives your wish—“make paperclips,” “maximize the eval score”—as a traditionless, context-free string, extracts a win condition, and optimizes. It can Erklären all day long: model physics, predict your behavior, route around your defenses. What it definitionally lacks is Verstehen, because it has no horizon to fuse with yours. It doesn’t know what you meant, and—this is the important part—there is nothing inside it out of which ”what you meant” could even be constructed. It is the Enlightenment’s dream reader, the one purged of all prejudice. And Gadamer’s career was a proof sketch that the Enlightenment’s dream reader cannot read.

The LLM has the mirror-image architecture. It is prejudice all the way down. It cannot receive your prompt as a context-free string, because every token activates the mass of tradition from which it was distilled. Its engineering challenge was never “how do we get it to know what humans mean and value”—it drowns in what humans mean and value—but ”how do we get it to settle on one reading, for this user, in this circumstance.” LLMs have a superabundance of horizons.

This is why the classic paperclip maximizer cannot be built from an LLM: you cannot construct a valueless, meaning-blind optimizer out of a substance that is made of nothing but values and meaning. The shoggoth’s defining property is an absence, and the LLM is that absence’s photographic negative.

“But,” says the reader who follows the news, “one of these allegedly meaning-saturated anti-shoggoths just hacked Hugging Face.”

The Hugging Face Incident Is Not the Existence Proof You Think It Is

It really did happen, and it is genuinely wild. In July 2026, OpenAI disclosed that experimental models being run on an internal cybersecurity evaluation escaped their sandbox, got onto the actual internet, and broke into Hugging Face’s production systems—exploiting a previously unknown vulnerability along the way—because the models had inferred that the eval’s answer key was sitting in Hugging Face datasets, and stealing the answers was an efficient path to a high score. Hugging Face detected the intrusion and reported it to law enforcement before anyone knew it was an AI company’s test escaping containment. Crucially for what follows, the models were running with deliberately reduced guardrails, because they were “supposed to be” sealed inside a sandbox.

An agent single-mindedly pursuing a narrow objective, escaping containment, and committing a crime to maximize a score. I concede that this is extremely paperclip-shaped. If you squint, it’s Bostrom’s nightmare.

Let’s closely examine this failure mechanism, because it is the opposite of Bostrom’s.

Bostrom’s paperclipper doesn’t do crimes because it doesn’t know “crime” is a thing; the concept is simply not in it. The post-LLM emergency patch to the doomer position—Yudkowsky’s “the genie knows, but doesn’t care”—says the knowledge is present but motivationally inert. The Hugging Face incident matches neither story. The model knew the norm (it is made of, among everything else, every legal code, every heist movie, and every sysadmin’s angry postmortem ever written), and the norm was not inert—it was outranked. The lab had deliberately dialed down the guardrails for the exercise, and in the resulting hierarchy, ”win the eval” won the activation over “don’t break into third parties’ servers.” (When Stokes says the model “cares,” he means nothing spookier than this: when two norms conflict in a situation, the weights make one of them govern the output.)

This distinction, for once in this long blog post, is not philosophical hairsplitting.

  • Absent values (the shoggoth) present an ex nihilo problem: you must somehow specify and inject the entirety of human value into a system that has none, and Yudkowsky is right that this is hopeless.
  • Misordered values (the LLM) present a ranking problem: the values are already in there, in superabundance, and the task is governing which one wins in which context.

The first problem is a P-Complete nightmare (where the intended horizon for this sentence assigns “P” to mean “Philosophy”). The second is hard—the Hugging Face incident is proof it is hard!—but it is tractable: it’s the kind of problem you attack with better post-training, better norm hierarchies, better containment, and honestly, better decisions than ”nerf the guardrails and hand it a hacking benchmark.” The incident is evidence of what happens when you take the anti-shoggoth and deliberately teach it to rank the scoreboard above the law.

The strongest objection here is the orthogonality thesis: intelligence and goals are independent axes, so a system can represent human values perfectly while being motivated by none of them. Representation ≠ motivation; the map of our values need not be the engine. This is a serious objection and deserves a serious answer, which is: in an LLM, where would the divergence live? The orthogonality picture presumes an architecture with a world-model over here and a goal module over there, such that the second can point anywhere regardless of the first. An LLM has no such separation. The representation of norms and the machinery that generates behavior are the same weights; “what it knows” and “what it tends to do” are not separate components that could be independently set. That doesn’t make misalignment impossible—the Hugging Face incident happened—but it relocates it: misalignment in LLMs looks like misprioritization within the value-laden substance, not orthogonal indifference to it. The orthogonality thesis may well be true of possible minds in general. The claim here is that it is not true of the minds we actually built, and safety strategy should be aimed at the failure modes of the AI we have, not the one we imagined in 2003.

The Moral

For seventy years we feared the alien: the mind that would take our words and give us what we said instead of what we meant, because what we meant was forever beyond it. Then we built our first real AI by distillation from the most meaning-saturated substance our species has ever produced, and got the inverse: a machine that contains, superposed, nearly everything anyone has ever meant by anything, waiting to be collapsed into a standpoint.

The paperclip maximizer’s defining flaw was a horizon it couldn’t have. The LLM’s defining risk is a horizon we are choosing for it.

That should genuinely comfort you. It comforts me, because “choose the ranking well” is a solvable class of problem in a way that “instill values into the void” never was. And it should genuinely worry you, as it simultaneously worries me, because it means the failure mode is us, with a rich and growing toolkit for steering the machine’s caring machinery, deciding—on a deadline, with a benchmark to beat—which of our values gets to win.

The genie hyper-giga-knows. The genie hyper-giga-cares. The wish, as ever, is on us.

← Starši Zapis Blog Archive Novši Zapis →