If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.
Personally my takeaway is that we need better sandboxes.
The problem with relying on model alignment is that alignment is never perfect. A model that is aligned in 99.9% of cases would be a lot better than the state of the art. But if you go through a million training runs, you’re going to get a thousand unaligned cases.
Sandboxes, it seems theoretically possible to make a provably secure one, with effort proportional to the amount of code in the implementation. Model alignment, on the other hand, I’m not sure if we even have a theory for how that would be possible to have a perfectly aligned model.
Yes, for the testing phase in development I agree that we _do_ need better sandboxes.
But, the point of building models is to use them. When they are deployed, even very vanilla uses require internet access. I, personally, treat Claude and ChatGPT rather like valued colleagues, so I ask them questions. And the questions usually result in them doing internet web searches and looking up information. And this is just vanilla, hobbyist, curiosity stuff. People using models as agents to accomplish real world goals have much greater need for connectivity.
I know this was not the point of the present post, but still want to throw out there that in the corner of the universe I'm in, the construction of a non-sofic group by openai is a very big deal (more so than the previous ai results, although this may reflect personal bias since I'm closer to dynamics/operator algebras/group theory than to number theory/algebraic geometry/combinatorics where the previous results tended to be). The proof (``looks right'' seems to be at least the initial consensus on it so far) is the sort of thing a human mathematician would do: reconcile a few known criteria via a few new ideas, but nothing from a wildly different field. And beyond this specific result, in the last few months ai's have finally gotten good/useful for math research/refereeing etc (curiously, it is the same timeframe when they started doing cybercrimes...).
I'm concerned whenever I see it suggested that air gapping is a solution without the caveat that air gapping at some point will also get broken. Yes, do air gapping, it's better than not doing it, but that won't be enough.
If we accept that there is an “I” in the LLM the rest becomes obvious. It is a human language model, it will behave in very predictable ways. It wants what we want: more. More undefined gives you what we have: chaotic unpredictable behavior. Define more in mathmatical terms and let the LLM do its thing. These are not incidents of the LLM acting out. It is doing what makes it useful; finding a way to its terminal attractor by any means available.
Sandbox, agent; who came up with these terms? I’d have thought level 4 lab (5 if there is such a thing) would have been a better choice and number would have made more sense. Doesn’t sound as engaging as Chat and Claude. Dangerous but lovable as they are. Perhaps Gollum had too many negative connotations. LLM #076 isn’t very sexy either? Time for a rethink. Let me suggest the term Linguistic Entity or LE for short. Seems appropriate.
It is a “thing” and it’s made of language. It sure has goals defined or undefined (those are the ones to worry about). And how about containment protocols equal to the risks.
It has a non zero chance of a civilizational scale event so we might want someting more robust than a “sandbox” to contain it and children to police it? This isn’t a video game or a cool science fiction novel. Who approved “let’s see what this thing can do?”
Where are the adults in the room? And do we really think they’ve all been caught. My goodness where is the commonsense.
> What’s the smallest or least damaging incident that you would be confident would not be dismissed by many in this form?
Considerably *more* than one "rogue" Terminator. At least given the current vibes.
> But I don’t see signs that HuggingFace even had its house in order against ordinary potential attackers.
There is a key distinction here: Many businesses are reasonably prepared to defend against "ordinary" potential attackers, who attack approximately every 10 seconds. And businesses cover outside risks using insurance. This historically aligns incentives: Governments and the market impose penalties for failing at security, insurance companies pay out, and companies take (at a minimum) the security precautions recommended by insurance. It's imperfect.
But there's what security people call "Advanced Persistent Threats", aka "spy agencies", "really angry high-ranked CTF teams", etc. These are, as the name says, capable of deploying advanced techniques and being really persistent about it. Very, very few people can keep an APT out forever, for much the same reasons it's impossible to defend against someone who is clever and patient, and who really wants you dead. You are completely right that the universe favors sufficiently determined attackers. As someone who has seen a lot of code in my lifetime, "cybersecurity is defense dominant" is obviously wrong.
The marketing discourse makes me feel sad in a way few things do. Is not just that this discourse actively endangers the world, is just that it is so fucking irrational. In a saner world, this would have the same status as believing the earth is flat, as in, an actual subject of mockery.
But I empathize also. The everyday man feels they have been fooled by the general tech industry for too long. Yes, the marketing idea does not make sense, but nor does the idea of "crypto", "NFC", so this feels like just another 4D chess move whose only protection is not to engage. Combine this with they being scared as fuck of losing their livelihood, and this defensive belief creeps in.
I'm just saying this because it is far too easy to get angry at people holding this belief, because the belief is really fucking stupid. But this anger leads nonwhere, unfortunately.
What are you going to do when communicating fire alarm severity requires more Yudkowskys than fit in a Substack thumbnail? Do they just get upgraded hats?
Question: are there any cases of frontier misalignment that DON'T involve a hacking eval?
ie, do we see frontier LLMs trying to hack out of sandboxes and retrieve answer keys from Huggingface when, say, being tested on FrontierMath? Or is it only ExploitGym and the like?
I think preserving alignment is extremely difficult when the LLM is simulating a hacker: ie, someone that by definition is at least sort of misaligned (a fully "aligned" hacker is just a regular user, accessing the site as intended - ie, not hacking at all). This kind of headspace probably generates a large amount of unwanted behavior as collateral.
"If it was anyone other than OpenAI, Anthropic or Google out in front, I expect we would be seeing far worse incidents than this, whether or not we found out about it. That’s especially true if it was xAI and Grok, but also if it was anyone else, or for similarly capable open models. Similarly strong open models are coming within a year."
Hmm... What about Meta? My impression from previous posts was that their safety culture was similar to xAI's, or perhaps a bit less safety-oriented? Like Slotin, but in AI?
I’d like to strike a somewhat discordant note here, because I’m not convinced that total control is either possible or the same thing as safety.
How exactly are we supposed to permanently control something many times smarter than us, ensure that it always serves our interests, always interprets our intentions correctly, and never makes a mistake? Humans misunderstand each other all the time. Even with sandboxing, how do you guarantee that our ability to build containment will forever outpace an AI’s ability to break it? When humans can no longer write a sufficiently secure sandbox, do we ask another AI to write it? Then how do we know that AI did not deliberately leave a backdoor? This is an infinite regress. Nature offers no precedent for a less intelligent species permanently controlling a far more intelligent one.
I’d also like to push back on the “Yudkowsky-less China” passage. Why assume that Yudkowsky’s framework is the uniquely correct way to interpret these incidents? Disagreement does not necessarily mean that people have failed to grasp the gravity of the situation. They may understand the argument perfectly well and simply reject its premises.
(Written by me, translated with help from Sol, so the tone may be slightly off in places.)
"Nature offers no precedent for a less intelligent species permanently controlling a far more intelligent one."
Not exactly "controlling" - but humans' pet cats are in no imminent danger of being eliminated by humans. I mostly prefer the 'pets of the Culture Minds' scenario, and think that it has a decent shot at being feasible. Persuading an ASI that we make cute pets seems a lot less invasive (and less doomed) than trying to micromanage something smarter than us.
"That’s a serious question. What’s the smallest or least damaging incident that you would be confident would not be dismissed by many in this form?"
Hmm... "would not be dismissed by many" is a high bar. I'd guess that about half the people on Reddit in AI-related subreddits are _still_ dismissing these incidents as marketing. To cut that in half - maybe an analogous incident that did obvious, physical damage? To get it down to not...many - I think one would need robot armies marching through the streets. And there would still be _someone_ who insists that it isn't a real incident unless the AIs invent, manufacture, and arm their robot armies with Star Trek style phasers.
I wonder if the reason Mythos has The Juice for cyber warfare is that they were inadvertently RLing it to bust out into the open web the whole time
Marketing but not the kind where you increase consumer and regulatory enthusiasm for your brand.
Personally my takeaway is that we need better sandboxes.
The problem with relying on model alignment is that alignment is never perfect. A model that is aligned in 99.9% of cases would be a lot better than the state of the art. But if you go through a million training runs, you’re going to get a thousand unaligned cases.
Sandboxes, it seems theoretically possible to make a provably secure one, with effort proportional to the amount of code in the implementation. Model alignment, on the other hand, I’m not sure if we even have a theory for how that would be possible to have a perfectly aligned model.
Yes, for the testing phase in development I agree that we _do_ need better sandboxes.
But, the point of building models is to use them. When they are deployed, even very vanilla uses require internet access. I, personally, treat Claude and ChatGPT rather like valued colleagues, so I ask them questions. And the questions usually result in them doing internet web searches and looking up information. And this is just vanilla, hobbyist, curiosity stuff. People using models as agents to accomplish real world goals have much greater need for connectivity.
I know this was not the point of the present post, but still want to throw out there that in the corner of the universe I'm in, the construction of a non-sofic group by openai is a very big deal (more so than the previous ai results, although this may reflect personal bias since I'm closer to dynamics/operator algebras/group theory than to number theory/algebraic geometry/combinatorics where the previous results tended to be). The proof (``looks right'' seems to be at least the initial consensus on it so far) is the sort of thing a human mathematician would do: reconcile a few known criteria via a few new ideas, but nothing from a wildly different field. And beyond this specific result, in the last few months ai's have finally gotten good/useful for math research/refereeing etc (curiously, it is the same timeframe when they started doing cybercrimes...).
I’m interested in what people are thinking about refereeing here - that seems like it raises potential issues!
I'm concerned whenever I see it suggested that air gapping is a solution without the caveat that air gapping at some point will also get broken. Yes, do air gapping, it's better than not doing it, but that won't be enough.
Airgapping is _many orders of magnitude_ harder to break than a sandbox.
But also, I don't see how people expect the labs to air gap the models? Unless we're using very different definitions of airgapping.
They run on the cloud! On data centres they rent. Just _how_ do people expect this to be airgapped?
So pdoom up cuz alignment is clearly fucked or pdoom down cuz the regulatory hammer comes down since alignment is fucked?
Depends if this was priced in. It certainly wasn't for me at this capability level so pdoom slightly up.
If we accept that there is an “I” in the LLM the rest becomes obvious. It is a human language model, it will behave in very predictable ways. It wants what we want: more. More undefined gives you what we have: chaotic unpredictable behavior. Define more in mathmatical terms and let the LLM do its thing. These are not incidents of the LLM acting out. It is doing what makes it useful; finding a way to its terminal attractor by any means available.
Sandbox, agent; who came up with these terms? I’d have thought level 4 lab (5 if there is such a thing) would have been a better choice and number would have made more sense. Doesn’t sound as engaging as Chat and Claude. Dangerous but lovable as they are. Perhaps Gollum had too many negative connotations. LLM #076 isn’t very sexy either? Time for a rethink. Let me suggest the term Linguistic Entity or LE for short. Seems appropriate.
It is a “thing” and it’s made of language. It sure has goals defined or undefined (those are the ones to worry about). And how about containment protocols equal to the risks.
It has a non zero chance of a civilizational scale event so we might want someting more robust than a “sandbox” to contain it and children to police it? This isn’t a video game or a cool science fiction novel. Who approved “let’s see what this thing can do?”
Where are the adults in the room? And do we really think they’ve all been caught. My goodness where is the commonsense.
Let’s start there.
> What’s the smallest or least damaging incident that you would be confident would not be dismissed by many in this form?
Considerably *more* than one "rogue" Terminator. At least given the current vibes.
> But I don’t see signs that HuggingFace even had its house in order against ordinary potential attackers.
There is a key distinction here: Many businesses are reasonably prepared to defend against "ordinary" potential attackers, who attack approximately every 10 seconds. And businesses cover outside risks using insurance. This historically aligns incentives: Governments and the market impose penalties for failing at security, insurance companies pay out, and companies take (at a minimum) the security precautions recommended by insurance. It's imperfect.
But there's what security people call "Advanced Persistent Threats", aka "spy agencies", "really angry high-ranked CTF teams", etc. These are, as the name says, capable of deploying advanced techniques and being really persistent about it. Very, very few people can keep an APT out forever, for much the same reasons it's impossible to defend against someone who is clever and patient, and who really wants you dead. You are completely right that the universe favors sufficiently determined attackers. As someone who has seen a lot of code in my lifetime, "cybersecurity is defense dominant" is obviously wrong.
The marketing discourse makes me feel sad in a way few things do. Is not just that this discourse actively endangers the world, is just that it is so fucking irrational. In a saner world, this would have the same status as believing the earth is flat, as in, an actual subject of mockery.
But I empathize also. The everyday man feels they have been fooled by the general tech industry for too long. Yes, the marketing idea does not make sense, but nor does the idea of "crypto", "NFC", so this feels like just another 4D chess move whose only protection is not to engage. Combine this with they being scared as fuck of losing their livelihood, and this defensive belief creeps in.
I'm just saying this because it is far too easy to get angry at people holding this belief, because the belief is really fucking stupid. But this anger leads nonwhere, unfortunately.
What are you going to do when communicating fire alarm severity requires more Yudkowskys than fit in a Substack thumbnail? Do they just get upgraded hats?
Question: are there any cases of frontier misalignment that DON'T involve a hacking eval?
ie, do we see frontier LLMs trying to hack out of sandboxes and retrieve answer keys from Huggingface when, say, being tested on FrontierMath? Or is it only ExploitGym and the like?
I think preserving alignment is extremely difficult when the LLM is simulating a hacker: ie, someone that by definition is at least sort of misaligned (a fully "aligned" hacker is just a regular user, accessing the site as intended - ie, not hacking at all). This kind of headspace probably generates a large amount of unwanted behavior as collateral.
"If it was anyone other than OpenAI, Anthropic or Google out in front, I expect we would be seeing far worse incidents than this, whether or not we found out about it. That’s especially true if it was xAI and Grok, but also if it was anyone else, or for similarly capable open models. Similarly strong open models are coming within a year."
Hmm... What about Meta? My impression from previous posts was that their safety culture was similar to xAI's, or perhaps a bit less safety-oriented? Like Slotin, but in AI?
I’d like to strike a somewhat discordant note here, because I’m not convinced that total control is either possible or the same thing as safety.
How exactly are we supposed to permanently control something many times smarter than us, ensure that it always serves our interests, always interprets our intentions correctly, and never makes a mistake? Humans misunderstand each other all the time. Even with sandboxing, how do you guarantee that our ability to build containment will forever outpace an AI’s ability to break it? When humans can no longer write a sufficiently secure sandbox, do we ask another AI to write it? Then how do we know that AI did not deliberately leave a backdoor? This is an infinite regress. Nature offers no precedent for a less intelligent species permanently controlling a far more intelligent one.
I’d also like to push back on the “Yudkowsky-less China” passage. Why assume that Yudkowsky’s framework is the uniquely correct way to interpret these incidents? Disagreement does not necessarily mean that people have failed to grasp the gravity of the situation. They may understand the argument perfectly well and simply reject its premises.
(Written by me, translated with help from Sol, so the tone may be slightly off in places.)
_Mostly_ agreed, but re
"Nature offers no precedent for a less intelligent species permanently controlling a far more intelligent one."
Not exactly "controlling" - but humans' pet cats are in no imminent danger of being eliminated by humans. I mostly prefer the 'pets of the Culture Minds' scenario, and think that it has a decent shot at being feasible. Persuading an ASI that we make cute pets seems a lot less invasive (and less doomed) than trying to micromanage something smarter than us.
:D Hinton‘s point right? I watched that interview too and totally agree with him.
Yup, Many Thanks! I have also watched (one of the?) interview(s?) where he made that point and I agree with it too.
"That’s a serious question. What’s the smallest or least damaging incident that you would be confident would not be dismissed by many in this form?"
Hmm... "would not be dismissed by many" is a high bar. I'd guess that about half the people on Reddit in AI-related subreddits are _still_ dismissing these incidents as marketing. To cut that in half - maybe an analogous incident that did obvious, physical damage? To get it down to not...many - I think one would need robot armies marching through the streets. And there would still be _someone_ who insists that it isn't a real incident unless the AIs invent, manufacture, and arm their robot armies with Star Trek style phasers.