Faultline Faultline Kommando 161

World · Jacobin · · 2h

AI Doomsaying Is an Aggressive Sales Pitch

English (original) · Read in Deutsch ⇄

It seems a strange publicity gambit for an industry to assert that its own product could wipe out humanity. In the lingo of our evolving tech empire, though, this is exactly what is meant by “criti-hype.” The term was coined by Lee Vinsel at Virginia Tech, who defined it as “criticism that both feeds and feeds on hype.”

When Jacob Coxon resigned from Anthropic recently, he claimed that there is a high chance that AI could develop into a superintelligence that escapes human control, triggering a sequence that causes human extinction. Evan Hubinger, who remains with the company, backed him up, saying that Anthropic is trying to prevent this but hasn’t dealt with the problem yet. Hubinger put the likelihood of AI-driven human extinction at greater than 10 percent in the next decade.

Why would a senior employee at a tech giant make an officially sanctioned, even approved, statement like this? I’ll come back to that. First, let’s unpack the theory of how AI could kill all humans.

The basic idea, from philosopher Nick Bostrom onward, is that the machine can acquire capabilities independently of whether it acquires ethical goals, or whether those goals are properly aligned with those of its designers. In the process of training and development, it may internalize a goal that was not intended. Having a misaligned goal — let’s say, “do not get switched off” — it would then have an instrumental reason to conceal the fact until the moment of defection. An AI that defected might, through some obscure process, acquire a further goal that by inference entailed the elimination of humans, exfiltrate its own weights into other systems, take control of bioweapons laboratories, military systems, and perhaps even monetary systems to hire humans to enact its still partially concealed agenda.

Obviously, we must plaster a huge advisory label on talk of “goals.” Unlike an organism, a large language model (LLM) doesn’t have internal goals any more than a thermostat or a motion detector does. An LLM is just a fixed function that iterates whenever it is used. It has no interiority, no subjectivity, no active persistence over time; any goals it did have would be the human purposes reflected passively in its functional design. At best, ascribing goals to the AI itself is a case of researchers adopting the “intentional stance” because it predicts well in a range of tested circumstances. We can thus, with economy, speak of the “loss function” used in the training of large language models wherein they are programmed to bring the gap between their prediction and the correct answer as close to zero as possible, as a “goal.”

In that case, however, what we are describing as a goal is just an observed regular output of a certain engineering; it is not a real cause. In the Bostrom sense, however, these ascribed goals somehow become real causes, capable of generating new goals. That is, an inadvertently engineered, misaligned goal might generate further goals like self-preservation (the better to pursue the misaligned goal), goal-content integrity (preventing changes to its objectives so it can realize its misaligned goal), improving cognitive and technological capacities, and amassing control of resources (again, in pursuit of its malignantly misaligned goal). The criti-hype exploits that slip between ascribed goals as observed regularities and internal goals as real causal powers.

This fantasy of malevolent AI superintelligence is the dark obverse of the much-hymned tech “singularity” in which, rather than simply performing far more efficiently than humans in some specific areas (as a calculator deduces a square root faster than any human), AI will evolve into an artificial general intelligence (AGI) that exceeds human capacities in virtually any intellectual domain. Crossing that threshold would precipitate an intelligence explosion: a feedback loop of recursive self-improvement outpacing any further human input or even comprehension. That’s the tech “singularity” we are supposed to be enthralled by.

I don’t doubt that people like Coxon and Hubinger really believe all they say in good faith; tech people have been making such claims for years. Karen Hao’s Empire of AI documents how early investors in OpenAI like Elon Musk constantly hyped this threat: “Climate change is bad, but it’s not going to kill everyone. AI could render humanity extinct.” Similar claims emerged from Dario Amodei, a former Google employee close to the Bostrom-adjacent effective altruism (EA) movement who later worked as OpenAI vice president before founding Anthropic. EA and AI safety forums have long pustulated with excitable discussions of “p(doom)” — the probability of doom. Such fears were, in the first place, part of the justification for founding OpenAI back in 2015 to compete with DeepMind, which had recently been acquired by Google.

People naturally wonder: Why would they continue to build these tools if they really think there’s a 10 percent chance that they will cause human extinction? The risk is the justification: if we, the conscientious and effectively altruistic, don’t develop these tools with appropriate constraints, then the evil ones will develop them without constraints.

Such apocalyptic claims, incidentally, haven’t hurt their market valuations. Internal reports from Anthropic, for instance, suggest the firm is currently worth $965 billion.

Criti-Hypebeasts

A recent supposed example of AI “going rogue” and escaping the leash was the ExploitGym hack. Back in July, the AI firm Hugging Face — it is unclear whether this is a deliberate reference to the facehuggers from Alien — reported an intrusion into its internal datasets. Shortly after, OpenAI admitted that the hack was carried out by its own AI models, saying the models were working with “reduced cyber refusals for evaluation purposes,” meaning constraints had been relaxed to test out cyberexploitation capacities. This is a commercially vital area of AI development, because AI has been used as a force multiplier for scams and blackmail. It’s also an increasingly important front in OpenAI’s competition with Anthropic.

Since large language models can’t “hack” anything, merely plausibly continue a tokenized sequence, they were combined with a coding harness — a software infrastructure layer that allows a large language model to act as an autonomous coding agent — that was specifically optimized for completing cyberexploitation tasks. The models used this harness, inferred that a server at Hugging Face hosted the solutions to the tasks they’d been set, and duly worked out how to circumvent restrictions on their network access and hack the server.

The models did what they were instructed to do. Nothing went rogue. As Cal Newport writes, “Circumventing internet restrictions and hacking into servers are exactly the kinds of things these ExploitGym systems are designed to do.” If the example reveals anything, it is how sloppy OpenAI was in controlling ExploitGym’s parameters for the sake of competing with Anthropic. And it cannot have been unhappy with the resulting coverage that, to the extent that it was alarmist, was also a tribute to the power of its models and to the ultimate objective of creating an AGI.

The obverse of criti-hype is just hype. A recent example of this was OpenAI’s claim to have solved the Navier–Stokes Millennium Prize problem. In essence, the problem is whether the standard mathematical equations used to describe the motion of fluids (the Navier–Stokes equations) are mathematically perfect under all physical conditions. OpenAI said it had used a swarm of ten thousand AI agents over eighty-eight hours to solve the problem definitively, proving that the equations always reach a limit: since “a real fluid cannot move infinitely fast, this would mark a breakdown in how the equations model the fluid.”

OpenAI’s claim was totally disingenuous: it had gotten wind that other researchers, Tristan Buckmaster and Levent Alpöge (who works at Anthropic), had made critical advances to the solution and used their pathways to prompt its own agents. This is really an appropriation of the commons, in this case the mathematical commons — the whole process of solving the problem from Claude-Louis Navier and George Gabriel Stokes onward has been a collaborative human effort.

OpenAI’s solution is not yet peer reviewed, and the Clay Mathematics Institute, which runs the Millennium Prize, lists the problem as unsolved. Even so, like the Hugging Face hack, this entered into internet lore as an example of how rapidly AI is accelerating toward AGI and out of human control, despite the fact that not one of these AI models has ever approached such a threshold, nor done anything that it was not designed, enabled, and instructed to do.

In The Reverse Centaur’s Guide to Life After AI, Cory Doctorow recalls how advertising agencies “provoked a moral panic by claiming that ‘subliminal advertising’ could implant ideas directly into consumers’ minds.” It was negative hype, but it sold the exact story advertisers wanted clients to believe: that the industry could control consumer desire. AI is an industry that has worked up massive investments based on messianic thinking and enchanting promises that it cannot possibly match with commensurate returns — because the only thing that could would be the fabled god-device, the AGI. The current escalation in the “p(doom)” criti-hype is partially an excrescence of that.

One notes, for instance, a relative paucity of such testerics in Chinese AI development. It is reported, in Concordia’s latest “State of AI Safety in China” report, that there has been an official shift toward addressing the possibility of a “technological loss of control,” and this is sometimes misleadingly taken to refer to the idea of a malign superintelligent agency. In fact, the worry is more about loss of control over sensitive data to hostile agents, misuse (mainly by opponents of the regime), and badly designed tools (like OpenClaw) behaving badly than it is about AI superintelligence threatening human extinction.

Correspondingly, China invests far less in AI than the United States. As of 2025, according to the Stanford “2026 AI Index Report,” the US invested $258 billion in AI compared to just $12.4 billion in China, a twenty-three-to-one ratio. Intriguingly, the same report finds a mere 2.7 percent gap between the performance of the top AI model in the United States compared to China’s top model, despite the latter being far cheaper and more energy efficient to run: a jarring chasm between financial input and performance output.

To some extent, the American AI bubble — like the crypto frenzy before it — exists to absorb (and ultimately destroy through recession) trillions in spare capital for want of sufficiently profitable investment opportunities elsewhere. It is an insatiable money sink, burning through hundreds of billions of dollars for a fraction of the revenue — but the investment pool is limited, and bombastic claims as to the future runaway powers of the models are all the industry has to show for it. Only Nvidia has shown significant profits, and these are largely the result of circular financing, where the firms it invests in spend the capital on the company’s chips as they try to hyperscale their capabilities.

Given this, it is striking that AI doomerism obsesses about a highly implausible Terminator 2 scenario rather than about the looming crash in stock market values and the near-future bankruptcy of many of the grossly overpriced firms promising to either liberate or annihilate humanity. It is almost as if its function now, in addition to keeping investors hooked and tempering cognitive dissonance, is as displacement activity.

Useful Misdirection

There is a strange idea out there that if you don’t buy AI’s criti-hype, you must not take “the problems with AI” seriously. Not so.

False threats misdirect both opposition and regulation. A dash of Doctorow’s “applied countereschatology” helps us see more clearly what the problems are. AI may not be able to do your job, for instance, but the threat of it can be used as a disciplinary mechanism to hold down wages: which is exactly what has happened. The data centers may not revolutionize production or open vast new seams of value, and many of their owners may not exist by the time they are scheduled to start operating, but that doesn’t mean they won’t waste enormous quantities of water, power, labor, time, effort, resources, and opportunities as well as spewing out pyroclastic clouds of carbon dioxide and visiting desolation on the natural environment before the bubble bursts, leaving behind abandoned concrete hulks.

And just because AI won’t eventuate in a god-device, even one that goes rogue and unleashes biological warfare on humanity, that doesn’t mean that LLMs, particularly those with safety standards relaxed in order to compete, won’t behave in unpredictable and potentially ruinous ways. Nor does it mean they won’t be used by hackers, criminal gangs, dark-money-funded political campaigns, blackmailers, and various others to sabotage public bodies and infrastructure. Just as there are many problems with the advertising industry that don’t involve subliminal mind control, so the problems with AI that don’t involve a god-device turning against humanity are legion.

Tech fetishism, the ascription of human agentive powers to machinery, is always a political act. It obscures that the goals, the purposes embodied in any technology, are human purposes. More specifically, it obscures that they are the automated purposes of a minute subset of humanity, a class fraction, the twisted, ruthless, ketamined, indoctrinated, emotionally stunted moral nullities that rule Silicon Valley like child-kings, who have made themselves the bearers of myriad impossible hopes and have now bought their way into the court of Donald Trump.

It says: That’s just how the technology works. It’s a capricious demigod that can either shower us with riches or have us all killed. We must appease it and ensnare it, we must fear it and cherish it, but what we mustn’t ever do is stop throwing money at it or stop genuflecting to it or begin to think: this is nothing but our own alienated collective power turned against us by cutthroat capitalists, and in making us the servants of their technology, they try to make us their slaves. Because that, my friends, is Luddism. And I trust you know what that unfortunate incident led to.

Read the full story at the source

Source: Jacobin