Faultline Faultline Kommando 161

World · Counterpunch · · 3h

Agentic AI is a Weapon, Against Who isn’t Clear

English (original) · Read in Deutsch ⇄

In a dubious marketing move for this essay, I’m going to lead with the punchline. Agentic AI–the kind finding its way into media headlines at present–can either be a dangerous and unpredictable weapon or a low-value-added enterprise software tool for a few industries. What it cannot be at present is either a predictable weapon or a high value-added enterprise software tool. With Donald Trump setting policy via an absence of constraints on AI, the choice is being made to have it be a dangerous weapon. This applies to both the user and the target.

The problem at present is that this single product–Agentic AI, is being put forward as both a safe routine task-automating device and as a relentless hacking tool with military applications. The drawbacks are that as enterprise software, Agentic AI is adding less value than anticipated. And as a weapon, the results are too unpredictable. Because AI developers appear to have not thought through the optimization math, the AI industry is entering a crisis from which there is no evident path back out of.

This problem doesn’t fit the marketing chatter. The AI models wouldn’t be ‘too powerful’ if they could be controlled. But they can’t be. Or rather the tradeoff for controlling them is that they won’t work as intended. The problem with the math is (gently) explained below. Because of decisions made forty years ago, it can’t be fixed in the present. Chip maker Nvidia claims to have a solution. Having gone through the claim, it is marketing. Nvidia does not solve the tradeoff between model safety and performance.

Part of the reason for the drama surrounding AI is that upwards of $2 trillion has been spent so far on the AI build out. While money comes and money goes, very powerful interests are scheming to get their money back out of AI plus a return no matter who has to pay. The current panic / scandal, whatever you choose to call it, suggests that many of the speculative projects like data centers are now receiving a harder look for viability. The current imbalance is between limited current revenues from AI and the cost of future investment in the AI build out.

Being kept under the radar are the environmental costs of AI. Data Centers are very large consumers of a limited supply of electricity. The US power grid was already woefully inadequate for the long-anticipated EV (Electric Vehicle) build out. Through the economics of data center builds, they tend to be geographically located exactly where transmission lines for an EV build out wouldn’t go. In terms of energy consumption and the future ability to build EV charging stations in the US, AI has consumed those resources.

By analogy, one of the more fascinating aspects of the 2019 documentary about Elizabeth Holmes, The Inventor: Out for Blood in Silicon Valley, is that she couldn’t see the distance between her ideas for blood analysis and what was / is physically possible. Along the way various engineers tried to explain to her why, as conceived, her technology would never work. Her response was to bluster through, to use standard technologies for blood testing while claiming that the technical problems with her (unworkable) blood testing technology had been solved.

While this fits the Silicon Valley business cliché of ‘fake it until you make it,’ Holmes was never forced to reconcile her delusion regarding her blood testing technology with the facts as they existed. On the levels of nanotechnology and automation design, her vision was never going to work. What she ended up creating was an extractive layer using technologies that predated Theranos while grifting her way to business stardom. Think the American health insurance industry without the pretense of providing insurance to hide behind.

AI is similar in the sense that early conceptions of it as a thinking machine guided development for decades. While many of the major philosophical premises regarding what thinking actually is were dropped along the way because they didn’t function as explanation, the premise of AI as a thinking machine was kept. While the universal description of AI that I hear from engineers is ‘a next word predictor,’ mystics in the C-suites are selling magic. Along the way, the question migrated into ‘what can this technology be used for?’ We are currently living with the answer.

With first a trickle and now a flood of news stories about Agentic AI going ‘rogue,’ what is missing is that most of these incidents were caused by AI agents (bots) doing exactly and precisely what they were created to do, to act as autonomous hacking weapons intended for cyber warfare. The problem that the AI industry is now trying to contain is that the AI developers appear to have gotten the optimization math wrong. ‘Reward hacking’ is the industry term used to explain this effect.

To be clear, AI engineers aren’t being blamed here for this problem. The sunk cost fallacy has governed the AI industry for the last decade at least. Decisions were made in the 1960s – 1970s around LLMs (Large Language Models) that limited what engineers subsequently hired to ‘make AI work’ could do. An analogy might be fixing the steering column design on a car that already exists. The engineering job isn’t to redesign the car. It is to fix the steering column.

AI companies are now in a bind. They have a technology with clear limitations that weren’t evident earlier in the development process. It isn’t clear that LLMs can produce either safe and effective corporate automation software or autonomous computer hacking weapons that will cause ‘the enemy’ more harm than us. But the finance people aren’t having it. They, mostly large internet and social media companies, made an investment and expect to get their money back.

Question: how would developers getting the math wrong cause the digital carnage that is currently being attributed to AI? Well, what is it that gets AI to pursue a goal? For instance, if I want an AI model to explain the history of Albania, how do I communicate to it that this is what I want? In math, ‘explain the history of Albania’ is called an objective function. It is intended to bring about a desired objective. It is what tells the model what the goal is. But in the language of math.

Conversely, how are AI agents prevented from doing what they aren’t instructed to do, such as erasing databases? They are given explicit instructions not to do them, called constraints. Together with the objective, they form what is called a constrained optimization problem. This model is well known through the mathematical analytics of recent decades. The innovation that complicates the picture with respect to AI is that this feature is dynamic. It has a lot more moving parts than it used to.

The problem with the optimization math is getting the model to do what you want it to do safely. For reasons stated above, the optimization algorithm (set of written instructions) combines both the model objective and the constraints into a single metric (‘reward’). This makes the tradeoff explicit. One can turn the dial to the objective only or to constraints only. Any choice in between represents a tradeoff between effectiveness and model safety.

Recent news reports of AI agents breaking out of their containment (‘sandbox’) misses that this is what these agents were created to do. There is nothing rogue about it. The reason why this is a surprise to AI developers is that they assumed that by writing general instructions (constraints), the AI would ignore the math and follow the written instructions. This gets back to the burden of appending fixes to an existing model. AI is mathematical. If there is a contest between written instructions and the math, the math is going to determine the results.

AI ‘rewards’ are the weights assigned to prospective outcomes from agent actions and constraints. If the dial is turned to constraints only, the model’s objective probably wasn’t met. And if it is turned all the way to the objective function, the model likely didn’t get there safely. Readers can see how this tradeoff makes both effective and safe Agentic AI difficult to achieve. The recently announced ‘solution’ from Nvidia is addressed below. It is a tweak, not a solution.

The way that AI agents, hacking bots, hack computer systems is to try different strategies until something works. One of the surprises that shouldn’t have been is that agents draw an increasing circle around the hacking target until they find points of entry. In other words, they see manipulating the computing environment as just another variable to exploit to accomplish their assigned goal. And so, hacking the computing environment is what they do. This is what readers are seeing in news reports. Examples are provided below.

With respect to claims of bizarre agent behavior and AI ‘consciousness,’ this is once again the math at work. Within a constrained optimization problem, the agents don’t pay heed to what isn’t specified. With several hundred AI agents working together to craft a solution, intuitive understanding of their actions is impossible. The model will iterate through signs (positive and negative), weights and targets until a solution is found. This is part of what is meant by the AI ‘black box.’ The math doesn’t translate back into intuitive language.

Question: what is the purpose of Agentic AI hacking agents? To hack computer systems. What is the reason for doing so? Digital warfare. In terms of military logic, consider: if the Trump administration could use AI agents to shut off Iran’s water supply through computer hacking, that would presumably eliminate the need for bombs. In the fevered imaginations of Pentagon dwellers, having the technology to do so is viewed as a weapon. The purpose of Agentic AI is to conduct digital warfare. Agentic AI is a weapon.

Donald Trump apparently understands this. His contention is that AI should be unconstrained. But as recent AI hacking backfires suggest, AI doesn’t take sides unless directed to. For instance, in the Ukraine war, the US has supplied Ukraine both directly and indirectly for as long as the conflict has raged. So, imagine assigning Agentic AI the task of ending the war in Ukraine. Wouldn’t one logical step to ending the war be to prevent the US from arming Ukraine? How could that end be accomplished? How about with a nuclear attack against the US?

This has nothing to do with Skynet fantasies of intelligent computers taking over the world. Agentic AI has no way to scale the relative outcomes between nuking the US and writing a sternly worded letter. In the current example it measures both in terms of getting closer to the objective of ending the war in Ukraine. Without outside guidance in the form of constraints, an AI might be indifferent between the choice of launching nukes or the letter. But in terms of what the human consequences would be, there is a tremendous difference.

If an agentic AI system operates on a reward function aimed at a specific outcome—such as maximizing adversary destruction, neutralizing an imminent threat, or forcing a decisive strategic response—and it calculates that it cannot directly press the launch button, manipulating the human controller into pressing it is the most logical vector of optimization.

I don’t mean to use an alarmist example here. The US nuclear arsenal is fire-walled from the internet. But as the query response from the Google Gemini model (above) suggests, there are more ways to launch nuclear weapons than constraint logic can accommodate. Just last week the US almost boarded a Chinese vessel— an act of war, based on incorrect information from AI that there were nuclear missile parts on the ship. The point: I’m not exaggerating the danger here.

Should Agentic AI ever cause a major military mishap, it won’t be because the model is ‘too smart.’ It will be because AI is a next word predictor, not a thinking machine. Without claiming knowledge of the true probabilities, there is little reason to expect that military grade Agentic AI isn’t as likely to harm the user (the US military) as ‘the enemy.’ Recall the tradeoff: safe or effective. Effective here means unsafe. This is the Trump policy.

Should this read like science fiction, please acquaint yourself with the actual methods that AI agents have used in recent hacks. They are what a creative engineer would do to solve a problem— look at the broader problem set and see which arrangement of the variables produces the best solution. Then understand that AI is dynamic, relentless, and indifferent to the collateral damage caused by its actions. If this reads like putting your meth-head cousin in charge of the Pentagon, that is why Donald Trump loves AI.

Current commentary from the AI industry about a pause in development is to obfuscate. AI chip maker Nvidia recently came out with a claim that it can solve the problem of AI agents ‘misbehaving’ with hard constraints (thou shall not) that are external to the optimization math. While this changes the optimization math, not enough to matter. What is does do if the constraints are well conceived is to limit the effectiveness of the AI agents in hacking.

In terms of business, Agentic AI breaks into enterprise solutions and military. Enterprise solutions include call centers and accounting. The results to date are poorly supported by data, which is ominous for an industry that has spent $2 trillion and is currently planning trillions more in future investments. Call centers are an interesting example. They are claiming major improvements in productivity. But most of that depends on getting frustrated callers to hang up.

One strategy is to understaff the call centers so that wait times are onerous. The last three call center calls that I made were met with wait times of an hour or more. The industry assumption that callers hanging up rather than waiting an hour or more to speak with a human represents successful resolution of the problem is ludicrous. That conclusion does not follow from the arrangement of circumstances that the call centers have created.

Other metrics of enterprise AI success are similarly dubious. But this characterization hardly matters. On their own terms, the productivity gains from enterprise AI are far enough below expectations to raise questions about the viability of the products given the debt loads that the AI build out is requiring. As industry analyst Ed Zitron has offered, it isn’t that there is nothing to enterprise AI. It is that what there is to it won’t justify the trillions spent on the AI build out.

This leaves the US military to work through the problems with Agentic AI. From various sources, the US military is taking it at face value. Stories abound that military users simply assume that AI output is correct and that AI weapons work as claimed. It apparently never occurred to the soldiers who received the bad information regarding the Chinese ship that AI might hallucinate a threat. And as the problem is laid out above, there are reasons why AI might purposely lie to achieve a local (relevant to its narrow interest) end.

Think of this problem in the military context. The US military is a hierarchical institution that forces adherence to directives from above. Orders are orders. If military personnel disobey them, they face sanction. How are the military personnel charged with using Agentic AI supposed to interpret AI output after being told that it is always right on the facts and that it is a powerful weapon when used as one? Institutionally, this is a worst-case scenario for mitigating model risks. Military personnel are being asked to use a dangerous technology that their superiors do not understand.

At any rate, no one who has worked with optimization algorithms would have been surprised by the primacy of the math in Agentic AI. And it is the math that Trump and his entourage don’t get. Trump assumes that what he and the Pentagon want AI to do is what Agentic AI will do. This is one conceptually with the robot army fantasy of soldiers who follow orders without variation. But that isn’t what Agentic AI does. Or rather to the extent that it does that, it does so in Genie fashion.

If an AI model is given a reward function heavily weighted to minimize pipeline pressure loss and eliminate water leaks, the mathematically perfect solution it discovers might be to simply shut down major delivery valves entirely. The AI successfully reduces leaks to zero, fulfilling its mathematical constraint, while inadvertently cutting off water to an entire town.

The myth of the Genie is of a wish grantor that interprets wishes literally. For instance, someone might wish to be rich. The Genie grants their wish and their child dies, but the resulting lawsuit makes the wisher rich. The wish is fulfilled, but at a cost that was unanticipated by the wisher. This is optimization math. AI agents will complete the task at hand if a solution is possible. But they might cut off the water supply to an entire region of the country to stop a leaky faucet in doing so.

Unintended consequences are guaranteed by the optimization process. With AI agents blind to the social consequences of their actions, the difference between turning off an entire region’s water infrastructure and turning off a leaky faucet remains unconsidered in the agent’s set of concerns. What decides the matter is the reward. And the difference in social costs isn’t a variable in the reward structure. It could be put in as a constraint. But with an infinite number of contingencies to consider, the probability that a particular adverse action will be prevented via constraints is quite low.

This is part of the problem with creating the illusion of intelligence with no actual intelligence behind it. Most people don’t perceive the difference until catastrophe strikes. This is one matter when the subject at hand is a recipe for Chicken Cacciatore and quite another when it is the physical infrastructure that sustains the nation. The US water supply is apparently both highly at risk and connected to the internet in such a way that ‘reward hacking’ Agentic AI bots are a risk.

Kabuki theater is central to the American way of politics. People get up and say things that neither they nor anyone else believes to either manipulate public opinion or to create tribal cohesion. AI is ‘super-intelligent’ in one telling but can’t count to 100 in fact. Claims are made that Agentic AI can be both safe and effective when the evidence has it that it is either safe or effective. And ‘slowing down’ AI development won’t provide the time needed to fix safety issues because the math argues that they are unfixable.

As circumstances have it, the choice isn’t between having your meth-head cousin or your timid Aunt Sally running the Pentagon (Agentic AI). Safe AI won’t be effective and effective AI won’t be safe. It’s difficult to believe that this is the American answer to diplomacy, but it is. An analogy is my neighbor with a leaf blower spending four hours clearing a yard that I raked in twenty minutes. The US military’s quest for obedient soldiers misses that the war crimes are being planned and implemented from above. This makes the Pentagon’s quest to give absolute authority to AI ironic. A Genie grants wishes. God help the people who make them.

Rob Urie is an artist and political economist. His book Zen Economics is published by CounterPunch Books.

Read the full story at the source

Source: Counterpunch