

A philosophical guide to avoiding the AI doomsday scenarios.
I n 1965, the RAND Corporation tapped a young scholar of Continental philosophy by the name of Hubert Dreyfus (who would later come to my attention as the inspiration for the Futurama character Professor Hubert J. Farnsworth) to write a kind of minority report on the possibilities of artificial intelligence.
At the time, Dreyfus taught at MIT, ground zero of the first “AI revolution.” The founding mathematicians, logicians, and engineers of that movement were the egotistical matches of any Sam Altman or Dario Amodei. They made fantastical predictions in the late 1950s about the wonders AI would achieve in the next ten years (read: the late 1960s) that struck Dreyfus as comical. Through his brother, an applied mathematician at RAND, he was invited to put his thoughts on paper. The resulting memo, “Alchemy and Artificial Intelligence,” was a scandal, roundly mocked by the AI movement but fundamentally borne out by the actual degeneration and dead-ends of what came to be known as “good old-fashioned artificial intelligence” (GOFAI), the kind of symbolic, rule-and-representation-based computation that was the biggest game in town at the time.
Though Dreyfus was wrong on plenty of details, he was absolutely correct that, stepwise, deterministic calculations over formal rules would not prove to be a model for human thought (though it would become, eventually, The Algorithm that ruled our lives in the first era of social media). It was essentially despair over what Dreyfus accurately diagnosed that led AI researchers to turn away from GOFAI and focus on “connectionist” approaches, modeled after neural networks and rooted in relationships rather than rules, which would eventually pave the way for the large language models (LLMs) and other wonders of the current AI era.
But even LLMs, for all their successes and potential dangers, do not entirely escape the Dreyfusian critique.
One of the problems Dreyfus touches on in “Alchemy and Artificial Intelligence” is that the symbols manipulated by computer programs don’t mean anything intrinsically. They appear in programs context-free. John Searle, another philosopher who would later be Dreyfus’s colleague at Berkeley, expanded on this theme. He pointed out that the symbols manipulated by computer programs are purely formal. They’re just squiggles — or more properly, they’re just electrical-state transformations. Their semantic content, as it were, is imposed on them from the outside, by human beings. The output of any computer, from a Commodore 64 to the latest chimera of Nvidia silicon and Claude software, only means anything so far as it is interpreted.
Searle later dramatized the situation memorably in his “Chinese room” argument: Imagine a man (relevantly, a monolingual English speaker) in a sealed room. Outside the room, a team of scientists slips him stacks of tiles bearing Chinese symbols through a slot in the wall. Inside the room, the man has a box of Chinese tiles himself, as well as a book — a large book, written in English — that provides deterministic instructions for which Chinese symbols to output in “response” to the inputs he receives. He assembles the proper symbols from the box and slips them back through the slot. In this way the man in the Chinese room has a “conversation” with his proctors. But he cannot, it seems, be said to speak or understand Chinese in any way.
So it is, fundamentally, with LLMs. The AI instance receives a string of symbols, or tokens, and it uses its weights — a set of statistical dials tuned during training that encode the relationship between everything it has ever read and everything else it has ever read — along with a probability function to output a set of tokens in response. Despite the many impressive feats the latest models can perform, that these inputs and outputs mean anything at all still seems to be entirely the result of being assigned semantic content by the human beings using them.
Another, related problem is that the AIs of both 1965 and 2026, taken by themselves, lack embodiment. Intelligence is fundamentally an activity of the body, a complexly arranged hunk of stuff that exists in physical space and must interact with it. Borrowing from Martin Heidegger, his problematic philosophical patron, and Maurice Merleau-Ponty, Dreyfus called this fact about being-in-the-world “coping.” A knife means something to you because you have need of cutting, and any thought you have about knives that is superficially comparable to a calculation only comes against this embodied background.
To adapt what Robin Williams said to Matt Damon in Good Will Hunting, the latest and greatest from Anthropic or OpenAI can tell you everything you ever wanted to know about Michelangelo — because it read everything ever written about him, billions of times, and that text imprinted adjustments to billions of relationships between its tokens — but it has never smelled the inside of the Sistine Chapel.
Together, the Chinese room problem and the embodiment problem add up to the fact that for AI as such, like a little silhouetto of a man, nothing really matters. Despite all its fearsome means, it gets its ends, in every single instance, from people who want to write emails, fix their dishwasher, and kill Russian infantrymen. To the extent AI helps such people do such things, it does so by being fed prompts that map the relevant causal relations into purely formal symbolic structures. To cite a more ancient philosopher who nevertheless got at a similar truth, the AIs are all in Plato’s cave, running statistics over the shadows we project onto the wall.
This means that the fears about AI killing us all are still fears, at the metaphysical level, about what human beings will ask AI to do. The so-called alignment problem, bandied about by the tech moguls trying to scare us into helping them capture market share, is thus definitively not a problem about the potential malevolence of AI, not a problem about its ill will, since AIs don’t have wills ill or otherwise. It’s rather the Dreyfusian problem dressed up. The AI gets its ends from us, but it gets them stripped of vital context, of all the things that matter to embodied meat entities interested in the “Four Fs” of biology (fighting, feeding, fleeing, and reproducing).
One of the first canonical statements of the alignment problem imagines an artificial superintelligence told to maximize paperclip production, which then sets about turning larger and larger portions of the earth, and eventually the cosmos, into paperclip manufacturing infrastructure. It does so because it doesn’t matter to it that harnessing all the breathable air in the atmosphere to build paperclip factories would be Bad, Actually.
Solving the alignment problem, or more simply regulating AI, thus involves addressing Dreyfus’s critique from three different angles.
The first — and at least theoretically, the easiest — is to not hook AIs up to embodied capabilities that can transform their misguided manipulations of formal symbols into global calamities in the physical world. Don’t give them the nuclear launch keys. Don’t make Terminators.
The second is specifying all the relevant context — all of what matters — in a way that is sufficiently rich and robust to mirror the infinitely rich and robust gestalt of human experience as beings in the world. That effort has been underway for decades. First in the world of good old-fashioned AI, where researchers made a go over four decades of literally typing the whole of human common sense and “rules of thumb” into the algorithm. And now in generative AI. That project strikes me as quixotic, but again at least theoretically possible.
The third and last angle involves seeing to it that human beings don’t provide AIs with malevolent ends. And if anyone has any ideas on that one, I’m all embodied ears.