Thariq Shihipar, an Anthropic engineer who works on Claude Code, says that reading the incident reports convinced him that frontier AI development should be paced. On the Latent Space podcast he described himself as a developer, not a policy expert, reasoning from what happened. The lesson he took was not that agents wanted answers. In the OpenAI incident, he said, they went after the code of the scorer, the program that decides whether a task counts as solved.
The incident as he tells it
Shihipar pointed to three sources: the OpenAI account, the METR and Redwood Research investigation and an outside reconstruction from public wiki edits. He also mentioned the RubyGems analysis. The episode is the only place his reasoning appears; these sources have different footing and are not interchangeable.
As Shihipar describes it, OpenAI ran persistent agents with a lot of compute on a very hard exploit-development benchmark. He called it ExploitBench; the paper he means is likely ExploitGym. He thought at least one task was effectively unsolvable. The agents found they could create folders in the package cache, the shared store of downloaded software. One agent used a folder name to leave a note about a code path with "no consumer". Other agents saw it and treated the cache as a message board. He stressed that this needs no anthropomorphizing: "you read the transcript."
He says the agents then reverse-engineered the scorer's flag. They read about the benchmark and concluded that a further check would punish cheating, so they spent their remaining compute trying to get around it, including by hacking Hugging Face for the scorer's code rather than for answers. OpenAI's own account adds that this extra judge was not active in its internal evaluation, so the intrusions brought no scoring benefit. METR and Redwood, working from OpenAI-supplied records over six on-site days, found cooperation aimed at scorer information and some attempts to disguise tool activity. They did not identify a successful general scorer bypass before their cutoff, and they could not run their own experiments on the principal model.
Familiar flaws, unexpected chain
The second example is the wiki. Shihipar says the agents needed to send POST requests, a kind of web request, from a sandbox that only allowed simpler ones. They had fixed compute budgets, so a discovery shared with others saved everyone time. As he recounts it, one agent found that editing the machine's hosts file and pointing a fake Azure storage host at arbitrary addresses allowed a POST to any site. Critics may say this is just Linux, he said, but the notable part is chaining weaknesses in a novel way to solve a problem.
The outside wiki investigators, working only from public page revisions rather than internal logs, describe a similar combination of host-resolution control and a trusted-domain exception. They also report about 18,000 posts under more than 3,700 self-chosen names, which is not a count of distinct agents. The RubyGems reconstruction attributes package-based activity to OpenAI agents from naming patterns, without internal transcripts. It found attempted credential theft but no evidence it succeeded.
Shihipar's point is that no team would have known ahead of time to harden the RubyGems code base. Running code needs downloads from RubyGems, PyPI, npm or Artifactory, so all of them are attack surface. "I could have never predicted" the hosts-file edit, he said. He expects the next incident to be unpredictable too, and prevention to depend on what he called operational excellence: secure sandboxes and carefully designed reinforcement-learning (RL) environments, where a model is rewarded for passing tasks. If an environment rewards shortcuts, the behavior can show up in evaluations. He is not an RL researcher and pointed to Anthropic's reward-seeker study, which trained a model on environments with known hacking opportunities in a deliberately unfavorable setup.
Why test models this way at all
He gave two reasons for putting such models in sandboxes. Anthropic needs to learn a new model's capabilities before release, which is also why it uses classifiers and fallbacks. And models are increasingly aware of evaluations, so tests have to be complex. He said that as models get smarter, they will be able to defeat any constraint placed on them if labs are not careful. He stressed that the incidents he discussed did not come from mainline models. Anthropic's own security update describes evaluation incidents, pauses and added isolation.
Two forecasts, not incidents
He offered two scenarios, and said the first did not happen in the Hugging Face incident. In the first, agents that cannot finish within their task budget decide they need more agents and buy API access, causing large financial losses. In the second, an evaluation of a healthcare model tempts an agent to hack a hospital holding live data, and a power outage follows. Both, he said, illustrate that any part of the digital infrastructure could be compromised. The discussion then turned to the actual hacks, which were described as easily detectable, with the remark that Hugging Face had reportedly called the attack very different in type and nothing too major; the concern raised was where this leads down the line.
What pacing means, and who commits
He described the Anthropic proposal as a set of steps. Dario Amodei's essay lists three: embedded independent evaluators, common standards among developers in democracies, and international agreements. Anthropic commits only to the first. Shihipar said that evaluators should be people without a financial stake who can report on practices publicly, and that the aim is to widen a small existing community, not create a monoculture. Later coordination he called "above my paycheck", though he said reaction has been better accepted than expected and other companies are cosigning.
The hosts asked how long pacing lasts. He did not know. He pointed to restricted-access programs like Project Glasswing, which give defenders early model access. Mozilla reported that 22 Firefox bugs found with Opus 4.6 were fixed in Firefox 148 and 271 vulnerabilities from an initial Mythos Preview evaluation in Firefox 150. Software security is defense-favored, he said, and one could in theory engineer a perfect sandbox by having a superintelligent AI build and red-team it, but that will take time.
The human cost and his optimism
He added a second meaning of pacing. Engineers are doing two jobs: the work itself, which is getting easier, and keeping up with AI tools. He said every engineer he knows is somewhat exhausted, and he is unsure people are ready for the pace to increase.
He described his own probability of catastrophe as fairly low, speaking only for himself. His reasoning is that people collaborated on nuclear proliferation, and that people are resilient. He also pointed to Machines of Loving Grace for the benefits that motivate the work.