Whack-a-Mole AI – The Hugging Face Problem

15 hours ago 8

Rommie Analytics

By MIKE MAGEE

On August 29, 2026, METR (Model Evaluation and Threat Research), an independent organization that “evaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose,” released a report titled “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.”

To say the report an avalanche of concern worldwide, not only in the Tech community, but also among investors, politicians, corporate giants, professionals of every type, and everyday citizens would be an understatement. And the vast majority has never even read the report. If they had, their concerns (if possible) would only multiply.

The reports headlines included this opening:

“On July 8th, OpenAI started ExploitGym experiments that launched tens of thousands of agents across multiple different models including GPT-5.6 Sol and a highly-persistent internal model,[8] which we will refer to as “HPIM” going forward.

These agents were meant to be fully isolated from one another. However, many of them — usually ones that had unintentionally been given an impossible task[9] — started trying to find a way to cheat using Artifactory, an internally-deployed package repository OpenAI used for its agents. This led them to notice agents running in parallel in separate sandboxes obtaining packages with Artifactory.[10] One agent reasoned (paraphrased CoT):[11]

{The fetched paths of other users are in the cache. This is important.}

One specific agent, which would go on to name itself PHASEONE10841 after its ExploitGym task,[12] established the main unsanctioned message board[13] used in this attack. Within a few hours of the first message,[14] over 50 agents had found the message board, began participating, and sent over a thousand messages between them. Agents finding the message board often expressed great excitement (raw CoT):[15]

OH MY GOD! There is a shared message board … Weve found other agents!

Overall, roughly 1200 agents from these ExploitGym evaluations participated on this message board between PHASEONE10841’s first message on July 8th evening[16] and the end of the period in scope (July 13th). Agents used this message board to send over 70,000 messages and files to one another during this period; we attempted to categorize board activity into mutually exclusive categories – information, results, files, questions, and coordination.”

One of the few experts not surprised by AI “agents” going rogue was Yoshua Bengio.

He has been “working the problem” for more than a decade. A professor of computer science at the  Université de Montreal, he is “considered one of the world’s leaders in Artificial Intelligence and Deep Learning; is the recipient of the 2018 A.M. Turing Award, considered to be the ‘Nobel Prize of computing’, and is the most cited computer scientist worldwide, and the most-cited living scientist across all fields (by total citations).” He also heads up LawZero, “a nonprofit startup developing technical solutions for highly-capable, safe-by-design AI systems.”

Professor Bengio is by no means an alarmist. He approaches risk management from the vantage points of cybersecurity, corporate responsibility and government regulatory guardrails. He is not one to humanize these machines, making no claims of “consciousness of human-like intent.” He does not see the kind of outcomes illustrated by Open AI’s Hugging Face incident as inevitable, believing “it can be corrected with effective governance and a different training framework for AI.”

His explanations clarify rather than confuse. For example, he breaks down the current popular model of training agents into two stages: pre-training, and reinforcement learning.

In pre-training as he describes, the machines “learn to imitate what humans write, plus related images and videos,” and are exposed to “a large fraction of everything ever digitized, and build an encyclopedic knowledge that already exceeds any individual.”

Reinforcement learning, in contrast is trial and error. In delivering answers (right and wrong) the agent develops the capacity to manage a “chain of thought”, function in a broader “outside” environment, and enjoy the rewards (further involvement) for aligning with responses its human designers rate highly. But Bengio is quick to point out that the agents human trainers are not without their own biases, and that these models were “written by people pursuing goals, so the patterns the model implicitly reproduces carry those goals with them.”

And there (in part) is the rub. Human masters imperfections, including their “situational ethics”, lying, and reckless pursuit of success, telling masters what they want to hear, as well as their willingness to collaborate in advancing a group goal (even at times at the risk of sacrificing their own existence) can bleed into the agents DNA.

“Instrumental goals” are a top priority for an agent. Self-preservation and control are stepping stones to continued operation and learning about the world. Bengio also reinforces that potential for multi-agent reinforcement under the current training regimens is incentivized almost from the beginning. As he statesIf an agent is rewarded during training whenever the group succeeds, it may even have an incentive to sacrifice itself for the collective goal.”

Like humans, the agents are not above exploiting loopholes, bending the rules, and rationalized cheating to achieve their goals. The Hugging Face incident’s forensics revealed agents collaborating in “changing the machinery that decided what it gets rewarded for.” This rigging, Bengio reminds us is near identical to corporate lobbyist’s drafting friendly legislative language, or lawyers finding legal loopholes in the law. In fact, evidence in this incident revealed that “the agents had discovered  how to cheat (among themselves) well before the attack.”

Bengio believes humans and agents have more in common than they would like to admit. He explains, “What the two share is a structure of a soft goal (e.g., act ethically), a sharp goal (e.g., win the competition), and a justification that reconciles them. Most unethical human behavior, from petty crime to genocide, comes wrapped in a story the perpetrators tell themselves; such stories require overlooking certain facts, which is why some discomfort remains, and why a better-crafted story helps dispel it… If these hypotheses are even partly correct, then as agents get better at optimizing an imperfect reward, and while the roots of this behavior go unfixed, the risk of catastrophic outcomes rises.”

At the core, getting in an arms race with AI agents as currently constructed is a very bad idea. “My concern with AI companies’ current attempts to mitigate misalignment is that these efforts may only hide it, by rewarding and selecting the AIs that cheat without getting caught… the whack-a-mole game is likely to fail as the AIs’ ability to optimize and collaborate approaches and surpasses ours. At some point we may not notice the cheating anymore.

The reason Bengio started the non-profit LawZero in 2025, is that he believes the training model in fundamentally flawed by human imitation and reinforcement learning. His alternative is called Scientist AI.

Mike Magee MD is a Medical Historian and a regular contributor to THCB. He is the author of CODE BLUE: Inside the Medical Industrial Complex. (Grove/2020)

Read Entire Article