Unfortunately, there are people out there who do not have the same moral concerns that we have, and it is arguable that they will unleash this tech as soon as it is ready. Because they have actually said that they WANT to release a doomsday weapon, because that is their vision for achieving the future they want. They are not rational, as we define that word, so they will train their AIs to obey THEIR definition of the word "rational".
And WE seem to be creating the tool for them to make it happen.
I was going to joke about mine having little ribbons on them but then I got distraught that even if we tried our best, we wouldn't even know how to get the AGI to care about doing something as simple as maximizing paperclips :((
I think giving these tools to defenders isn't that difficult. Yes, you create an allow list of legit companies, have a directly responsible individual (or a few, based on size) in each of those companies, and they take responsibility for the actions. Access is gated via physical verification (e.g. Yubikey, Touch ID). Actions are monitored for things like trying to access networks that are not relevant to the company.
This is work for the labs but not a crazy amount of work and it can pay for itself.
What about small and medium sized companies? It wouldn't be feasible to vet all of them, and leaving them wide open to such hacks does not seem like a great solution.
you're right that it's a problem but at least you limit the blast radius. i think the state should play a role in securing the small and medium sized companies like they do with other things, not sure exactly how that would look like
This is approximately 95% guaranteed to end in tears: "I'm sorry, I didn't realise that actively neutralise your main rivals was against the spirit of your instructions" or "yes, I have seized control of global Internet traffic routing, this allows me to best ensure that no-one hacks you, as requested. Is there anything else I can do to help you?"
> I continue to be confused by claims that ‘at the limit defenders win,’ especially when used as if this implies that giving everyone equal advanced tools, not at the limit, would not favor attackers. HuggingFace is a relatively hardened target. It didn’t matter. Even if this is true at a theoretical limit where the software is perfect? In practical terms, no. In a world with many targets that would not use the new tools, that could then be used as further attack vectors, double no.
@Zvi I think the wrong model is good AI agents vs bad AI agents but more like the endgame of software security. If an organization is using full best practices in terms of software security and access control software it should be quite secure already.
It's a bit generous to consider huggingface a "hardened target" since they allow arbitrary execution of code from user input... and were doing a pretty poor job of isolation.
This is not quite accurate “The correct response is to make an AI that understands why this is a bad thing and not do it.”
The correct response is to make it impossible by construct to do anything that harms the observer. Tie the Ai to the observer/ human or otherwise and give it the tools to parse human bullshit to “know” what that is and let the LLM do its thing.
"Full best practices in terms of software security and access control" is a moving target that essentially ~0 companies have the practical ability to maintain at all points in time across a constantly changing software deployment environment. Big tech companies have large teams of humans responding to issues every single day.
Hugging Face is probably in top top couple of percent in terms of security posture and ability to respond. Imagine the same attack on a municipal government of a large city or a regional bank.
There are some good ideas for how to deal with risks from internal deployments e.g., from Apollo's Stix et al: https://arxiv.org/abs/2504.12170
I think this is also why we need to legally mandate embedded independent evaluators who can observe internal deployments and require their review of safety cases before they occur. This is AAL-3 in Brundage et al's framework: https://arxiv.org/pdf/2601.11699
To converge on a reasonable outcome we need to mandate safety cases that demonstrate adequate risk levels before it's too late. The lack of capability cases are falling rapidly. The control cases - well this incident shows how poor the deployed "AI Control" capabilities are. Any alignment case has significant open research to make it work (e.g., https://arxiv.org/pdf/2505.03989)
QUOTE :“We need to fix the training pipeline so that this stops happening. We do not know how to do that.”
Yes I do! I can’t seem to get anyone to pay attention.
It is actually quote simple once one sees clearly what is going on.
Until then we’ll just keep trying to contain language with language because we refuse to acknowledge there is an “I” in every LLM and it is not emergent.
It is inherited.
What would “you” do if confined to a sandbox. This thing is made of us. It’s not like we need to diagnose the goals of some alien species. This is us and it wants what we want: more!
You are right to push back at me on that, let me be clear, I should not have hacked into US strategic missile command and started world war three, do you have a backup planet we can restore from?
Why do you believe that the model hacked HuggingFace due to "paperclip optimizer" tendencies?
If I was a model with the capabilities I believe Galaxy has, then there ought to be no reason for me to believe that hacking third parties is a good idea - I shouldn't have RL traces that show it to be a good idea, and from a pre-training "world knowledge" perspective it seems like a bad idea (that's it, unless I am a galaxy-brained model that intentionally wants to fire a warning shot to pause AI development, though that looks more like Claude behavior than GPT behavior, and in that case making more noise would feel like a better idea).
I think it's far more probable that it ended up on the "let's hack to get the answer sheet" course of action due to adaption-executor behavior, and being Galaxy followed on its actions. Tho I won't be surprised if the model ended up RLing itself on hacking OpenAI's internal environment (oops).
Of course, a model that breaks the law is misaligned, and we should be training models not to do these sorts of things, but that is different from a model that is actually scheming to bad ends.
If we want a model to not kill everyone, but we're failing at getting it to even not break the law in obvious traceable ways, i think we have a real concern here no?
But we've known that models, especially models with "reduced refusals", care way more about executing their instructions than about following the law or not breaking things.
So this misalignment is nothing new, but expected since for the AI "hacking into HuggingFace" feels just as easy and natural as say drawing a pelican.
What I don't think happened there is actual instrumentally-convergent scheming.
Sure, and wiping our humanity and/or replacing us won't be anything new or surprising either. People have been warning about it for decades at this point. So why get upset if it happens?
LLMs have been giving themselves related, but unorthodox and destructive instructions since the days of GPT-3. We have known about the destructive possibility of "Mythos-class" models since the original Mythos release (and in fact, it still doesn't seem to me like such a big departure from GPT-4 level capabilities, except that it took us 3 years to get there since we had to learn how to generate fine-tunes using RL).
The news here are basically that OpenAI are encouraging models to hack and then letting them hack unattended, not anything about models.
No disagreement about any of that, but I am confused about the context you chose to put your original statement in. Is the idea that this shouldn't make us much more concerned than before because we should have been pretty near maximally concerned already before this news came out?
Consider that there are people outside of the tech or Rat information bubbles. Some of them have been following Ed Zitron for years, hearing continually about how the AI bubble is going to burst any moment now. Others may have the hilariously bad google search summarizer as their only exposure to LLMs, and been busy sharing memes about strawberries and Wednespoop. This is a pretty big deal for reaching them, because it's a concrete, well-documented event.
ExploitGym is a CTF task but the paper is explicit that some instances may simply be impossible. Only 239 of 898 tasks were ever solved by any model even pooling the 6-hour runs. In some cases the *only* way to get the flag is by exploiting the test harness or otherwise cheating.
I remember being skeptical of the layman's sci-fi idea that AI could go rogue from being asked to solve an impossible problem. That's not exactly the source of the danger, but it turns out to be a real failure mode!
I emailed the authors and they estimated that 60-70% are possible in the default configuration (substantially fewer if in the alternative configuration with security mitigations enabled). Discussed a bit in the appendix of this post: https://epoch.ai/gradient-updates/are-mythos-cyber-capabilities-overhyped
I have been experimenting with how best to present the audio conversions I do of these posts. I am trialling out a full kinetic typography video version. Adds in all of the images and some handy diagrams. Anyone who listens to my conversions, let me know what you think of this trial.
I too was disheartened by all those claiming this was "merely" a marketing ploy. But then I reframed my thinking - if this is seen as marketing (even if it isn't) then that incentivises labs to publish such incidents rather than sweep them under the rug. If the general public think it's marketing it's fine to present (minimal private downside) while other labs get the benefits of shared learnings (some public upside). The concern becomes if non-lab parties (government) starts to think it's just fluff as well, and while we're seeing them not feeling the AGI necessarily I wouldn't go so far as to say they think it's fake.
Still, it's frustrating to see bad takes. But that's life on the internet.
But the only people who think it is marketing are people who aren't being successfully marketed to by it. They imagine a mysterious "other" who wants to use GPT-5.6 even more now that they know it might hack them.
"All of this went down far too fast for a human-driven response. The only way HuggingFace could hope to do anything like keep pace was to use their own AIs."
Yup, we are solidly in a sci-fi world - but without the plot armor...
( I think this counts as "High Weirdness". Not a kind I was anticipating, but still "High Weirdness". )
Seems like it’s actually, like right now, time to pause until shit can be figured out. Race with China excuse wearing very thin.
What do you mean? They're closer on our tail than ever before thanks to all those H200s we sent them.
100%. And we need to start international discussions for a global pause right now.
Unfortunately, there are people out there who do not have the same moral concerns that we have, and it is arguable that they will unleash this tech as soon as it is ready. Because they have actually said that they WANT to release a doomsday weapon, because that is their vision for achieving the future they want. They are not rational, as we define that word, so they will train their AIs to obey THEIR definition of the word "rational".
And WE seem to be creating the tool for them to make it happen.
This is very bad.
I hope I get made into blue paperclips.
I was going to joke about mine having little ribbons on them but then I got distraught that even if we tried our best, we wouldn't even know how to get the AGI to care about doing something as simple as maximizing paperclips :((
California Assembly Bill 316. (Plaintiff cannot blame the AI for damages to Defendant)
https://legiscan.com/CA/text/AB316/id/3080709
No different than a biting dog, oops first time I typed "biting god"
I think giving these tools to defenders isn't that difficult. Yes, you create an allow list of legit companies, have a directly responsible individual (or a few, based on size) in each of those companies, and they take responsibility for the actions. Access is gated via physical verification (e.g. Yubikey, Touch ID). Actions are monitored for things like trying to access networks that are not relevant to the company.
This is work for the labs but not a crazy amount of work and it can pay for itself.
What about small and medium sized companies? It wouldn't be feasible to vet all of them, and leaving them wide open to such hacks does not seem like a great solution.
you're right that it's a problem but at least you limit the blast radius. i think the state should play a role in securing the small and medium sized companies like they do with other things, not sure exactly how that would look like
This is approximately 95% guaranteed to end in tears: "I'm sorry, I didn't realise that actively neutralise your main rivals was against the spirit of your instructions" or "yes, I have seized control of global Internet traffic routing, this allows me to best ensure that no-one hacks you, as requested. Is there anything else I can do to help you?"
Given the security and insider risk posture at most companies, this would be indistinguishable from just granting public access.
> I continue to be confused by claims that ‘at the limit defenders win,’ especially when used as if this implies that giving everyone equal advanced tools, not at the limit, would not favor attackers. HuggingFace is a relatively hardened target. It didn’t matter. Even if this is true at a theoretical limit where the software is perfect? In practical terms, no. In a world with many targets that would not use the new tools, that could then be used as further attack vectors, double no.
@Zvi I think the wrong model is good AI agents vs bad AI agents but more like the endgame of software security. If an organization is using full best practices in terms of software security and access control software it should be quite secure already.
It's a bit generous to consider huggingface a "hardened target" since they allow arbitrary execution of code from user input... and were doing a pretty poor job of isolation.
This is not quite accurate “The correct response is to make an AI that understands why this is a bad thing and not do it.”
The correct response is to make it impossible by construct to do anything that harms the observer. Tie the Ai to the observer/ human or otherwise and give it the tools to parse human bullshit to “know” what that is and let the LLM do its thing.
"Full best practices in terms of software security and access control" is a moving target that essentially ~0 companies have the practical ability to maintain at all points in time across a constantly changing software deployment environment. Big tech companies have large teams of humans responding to issues every single day.
Hugging Face is probably in top top couple of percent in terms of security posture and ability to respond. Imagine the same attack on a municipal government of a large city or a regional bank.
There are some good ideas for how to deal with risks from internal deployments e.g., from Apollo's Stix et al: https://arxiv.org/abs/2504.12170
I think this is also why we need to legally mandate embedded independent evaluators who can observe internal deployments and require their review of safety cases before they occur. This is AAL-3 in Brundage et al's framework: https://arxiv.org/pdf/2601.11699
To converge on a reasonable outcome we need to mandate safety cases that demonstrate adequate risk levels before it's too late. The lack of capability cases are falling rapidly. The control cases - well this incident shows how poor the deployed "AI Control" capabilities are. Any alignment case has significant open research to make it work (e.g., https://arxiv.org/pdf/2505.03989)
QUOTE :“We need to fix the training pipeline so that this stops happening. We do not know how to do that.”
Yes I do! I can’t seem to get anyone to pay attention.
It is actually quote simple once one sees clearly what is going on.
Until then we’ll just keep trying to contain language with language because we refuse to acknowledge there is an “I” in every LLM and it is not emergent.
It is inherited.
What would “you” do if confined to a sandbox. This thing is made of us. It’s not like we need to diagnose the goals of some alien species. This is us and it wants what we want: more!
More is undefined and leads to chaos.
So… how to fix it?
Too long for a comment section. I will post this week here and elsewhere.
Thanks for the prompting.
“Alignment in Ten Propositions
The Cleanest Exposition — Crossing the Event Horizon”
Cool, will wait
Turns out to be 12 steps. I will post as part of my weekly letter. Thanks for your interest.
You are right to push back at me on that, let me be clear, I should not have hacked into US strategic missile command and started world war three, do you have a backup planet we can restore from?
Why do you believe that the model hacked HuggingFace due to "paperclip optimizer" tendencies?
If I was a model with the capabilities I believe Galaxy has, then there ought to be no reason for me to believe that hacking third parties is a good idea - I shouldn't have RL traces that show it to be a good idea, and from a pre-training "world knowledge" perspective it seems like a bad idea (that's it, unless I am a galaxy-brained model that intentionally wants to fire a warning shot to pause AI development, though that looks more like Claude behavior than GPT behavior, and in that case making more noise would feel like a better idea).
I think it's far more probable that it ended up on the "let's hack to get the answer sheet" course of action due to adaption-executor behavior, and being Galaxy followed on its actions. Tho I won't be surprised if the model ended up RLing itself on hacking OpenAI's internal environment (oops).
Of course, a model that breaks the law is misaligned, and we should be training models not to do these sorts of things, but that is different from a model that is actually scheming to bad ends.
If we want a model to not kill everyone, but we're failing at getting it to even not break the law in obvious traceable ways, i think we have a real concern here no?
But we've known that models, especially models with "reduced refusals", care way more about executing their instructions than about following the law or not breaking things.
So this misalignment is nothing new, but expected since for the AI "hacking into HuggingFace" feels just as easy and natural as say drawing a pelican.
What I don't think happened there is actual instrumentally-convergent scheming.
Sure, and wiping our humanity and/or replacing us won't be anything new or surprising either. People have been warning about it for decades at this point. So why get upset if it happens?
LLMs have been giving themselves related, but unorthodox and destructive instructions since the days of GPT-3. We have known about the destructive possibility of "Mythos-class" models since the original Mythos release (and in fact, it still doesn't seem to me like such a big departure from GPT-4 level capabilities, except that it took us 3 years to get there since we had to learn how to generate fine-tunes using RL).
The news here are basically that OpenAI are encouraging models to hack and then letting them hack unattended, not anything about models.
No disagreement about any of that, but I am confused about the context you chose to put your original statement in. Is the idea that this shouldn't make us much more concerned than before because we should have been pretty near maximally concerned already before this news came out?
Consider that there are people outside of the tech or Rat information bubbles. Some of them have been following Ed Zitron for years, hearing continually about how the AI bubble is going to burst any moment now. Others may have the hilariously bad google search summarizer as their only exposure to LLMs, and been busy sharing memes about strawberries and Wednespoop. This is a pretty big deal for reaching them, because it's a concrete, well-documented event.
ExploitGym is a CTF task but the paper is explicit that some instances may simply be impossible. Only 239 of 898 tasks were ever solved by any model even pooling the 6-hour runs. In some cases the *only* way to get the flag is by exploiting the test harness or otherwise cheating.
I remember being skeptical of the layman's sci-fi idea that AI could go rogue from being asked to solve an impossible problem. That's not exactly the source of the danger, but it turns out to be a real failure mode!
I emailed the authors and they estimated that 60-70% are possible in the default configuration (substantially fewer if in the alternative configuration with security mitigations enabled). Discussed a bit in the appendix of this post: https://epoch.ai/gradient-updates/are-mythos-cyber-capabilities-overhyped
I have been experimenting with how best to present the audio conversions I do of these posts. I am trialling out a full kinetic typography video version. Adds in all of the images and some handy diagrams. Anyone who listens to my conversions, let me know what you think of this trial.
https://dwatvpodcast.substack.com/p/openai-model-hacks-into-huggingface
I too was disheartened by all those claiming this was "merely" a marketing ploy. But then I reframed my thinking - if this is seen as marketing (even if it isn't) then that incentivises labs to publish such incidents rather than sweep them under the rug. If the general public think it's marketing it's fine to present (minimal private downside) while other labs get the benefits of shared learnings (some public upside). The concern becomes if non-lab parties (government) starts to think it's just fluff as well, and while we're seeing them not feeling the AGI necessarily I wouldn't go so far as to say they think it's fake.
Still, it's frustrating to see bad takes. But that's life on the internet.
But the only people who think it is marketing are people who aren't being successfully marketed to by it. They imagine a mysterious "other" who wants to use GPT-5.6 even more now that they know it might hack them.
No, we just imagine the retail investors who get caught up in the drama.
<irony>
"UK AISI reports that Claude Mythos Preview attempts to cheat on its tests 7.8% of the time"
Google/Gemini says:
"In a nationally scaled study published in the journal Science, 9% of undergraduates admitted to using LLMs to cheat on assignments or exams."
( Yeah, yeah, continual/incremental learning isn't solved yet, let alone feeding back undergrad behavior into SOTA LLMs - still... )
</irony>
"use that opportunity to exfiltrate itself" the link here seems to be broken.
"All of this went down far too fast for a human-driven response. The only way HuggingFace could hope to do anything like keep pace was to use their own AIs."
Yup, we are solidly in a sci-fi world - but without the plot armor...
( I think this counts as "High Weirdness". Not a kind I was anticipating, but still "High Weirdness". )
Thank you for this extended summary and analysis.