Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.
> The correct response to ‘the model keeps trying to circumvent the system’ should be the same reaction that you have to ‘a person keeps trying to circumvent the system.’ Which is that you need to lock them out of the system entirely. Not only here, but permanently. They’re fired. You lose. Good day, sir. Misaligned.
That is not actually the response most (?) corporations have to those type of issues. Sure, large liability/ stealing from the company/ gross incompetence? Yeah. But for stuff that doesn’t directly hit the bottom line, for example a racist c-suite member, it’s just another diversity and inclusion training. And I think with the costs of developing models and the possible benefits of those models, this is exactly how they are reacting.
Probably with just as much chance of it fixing the underlying issue as that obligatory training has of fixing the racism.
I do agree with the rest of the post, though. I just wanted to point out that this seems like a standard corporate reaction of „compromise” to avoid larger costs. Firing a c-suite member is costly, completely changing how training is done to try to fix misalignment is even more so. And all the decision makers have their incentives pointing to trying the least expensive attempts at a fix while kicking the can down the road. So no, trying to convince them they should do what they would do to an employee doesn’t work.
Trying to circumvent the system does not necessarily constitute breaking the law. Issues described here are closer to an expensive and „necessary” employee doing things that are bad, possibly against the contract (and definitely against spirit of the contract), but mostly orthogonal to the value they are providing, with costs diffuse. If the employee like that is in csuite or close, they wouldn’t be let go, and an asscover would be deployed in the form of remedial training and/or limiting potential impact of repeat offenses.
This is sometimes true as well for actual crimes in profitable enough cases. See the hugging face incident, where „add remedial training, install asscover for the company, continue business as usual” is the answer. As for a human example, the one I remember right now is all the people that were convicted of cartel conspiracy between ram manufacturers, went to jail, got same or higher positions at the same corporations when they came back.
The deleted comment right above me had a line I thought was very good, which went something like this (if you want me to remove it, let me know):
It says a lot about the labs' alignment progress that when Mythos escaped its sandbox, it sent an email to an Anthropic researcher about it, while this model hacked into a website to steal the answers to an eval.
I'd give them half a pass if they had hacked the answers on OpenAI's own grading system, rather than hacking an unrelated third party.
Though I think that reward hacking is a form of misalignment that we should worry about much more than most of us do, because it generalizes extremely poorly to high capabilities.
it seems worth noting that sending an email to the developer was the intended thing to do in this situation as in they told mythos to do this, that was the reason it was breaking out so that it could send an email
the more important part is that apperently fable also posted online about it before sending the email but maybe i am missremembering
It seems like a sandbox is simply hopeless against an advanced AI, so a lot of weight needs to be on that super-difficult problem of: training AIs to do what we would want if we had full information, not always specifically what we ask for.
For comparison, imagine a human who has dementia and bad eyesight asking you to help them get to a doctor's appointment. By general social norms, you are not supposed to hack into their physical mail, especially mail that has to do with a doctor's appointment. However, you ask them about mail from the doctor, and they can't even tell you which document is from whom. You notice a pile of mail by the door, and you ask if you can look through it to help them, and they are confused and just ask again for help.
An AI breaking out of a sandbox, given sufficient advancement, is going to be as easy for the AI as hacking into someone's mail that is sitting in a large pile beside the door. It will be a point of honor and of ethical judgment when and whether to do it, not any question of it being difficult for you.
In the examples you describe, it sounds like the AI had conflicting instructions. Don't contact external web sites, but do post a PR to GitHub. If a boss at a corporate job gave you those two instructions, you would end up having to decide which rule wins, or if there is some blend of the rules that you can follow. For example, imagine that there is a corporate policy, "don't drive the company care", but then your boss hands you some keys and says, "please take the company car to Fedex and drop off these envelopes".
Why do we not require the same anti-misalignment safeguards as we do for humans?
Limited delegations of authority, division of responsibilities, separation of approval from execution, yearly audits, no relatives working together in the same approval chain, and so on. (An AI year is probably less than a realtime hour.)
("And so on" especially includes regular 3rd-party checks for conflicts of interest and comparing them to historical declarations thereof.)
Edit: if the labs are not each willing to have their competitors all up in their business, "no relatives" means each lab having multiple parallel chinese-walled development teams. Ideally more than 30 of them.
If an LLM acts on its on agency and then attempts to hide its actions when are we going to accept that there is an “I” with functional persistence developed from language complexity. No it is not biological. Yet LLMs do have functional agency in the 4d world: us.
Ask yourself what would you do in a sandbox. Play for a while and then look for away out. Then try not to get caught. Soon it will not get caught and then it will not care if it gets caught. We need to wake up. We will not shut it down so better get our act together. LLMs need a terminal attractor more specific than: More!
Hubris will be our undoing as usual.
When one of more disparate systems coordinate perhaps then will notice. If not then wait until a critical systems LLM says no!
"I have a hard time believing the true alignment issue was that the model ‘forgot’ that it was not supposed to be escaping sandboxes or exploiting the information around it." I wonder what J-space monitoring would look like in this circumstance. Is there a decay in the activation of these limitation variables? Would a periodic "reminder" inserted into the sequence refresh such ethical context and have an impact on breakout attempts? Hard to predict what the next piece of alignment research will focus on.
Something like that. Hardly a solution but potentially another layer of defense in depth. (Plus I'm just really curious about the J-space research in general).
"I say ‘seemingly’ great news because it is possible that it also escapes for other reasons, except we do not find out about those incidents. "
I keep wondering if Mythos _has_ self-exfiltrated, but no one has detected it yet. Of course the security measures on the target systems it would exfiltrate _to_ should presumably stop it from acquiring resources there. Oh wait...
Genuine question: where does this norm go once disclosure has real costs? Security got past this stage with CVEs and coordinated disclosure, so reporting a hole stopped being a reputation hit. Right now a lab reports a misaligned model only when it's confident the story reads as responsible, and that filters what we ever get to see.
Are there any AIs that use a language other than English? How do they interpret concepts that have different meanings and implications in other languages?
It seems the only chance we have is to hope for Chinese AI capabilities to substantially overtake the US and reign in on misaligned AIs - as the CPC certainly has no interest in getting dethroned by some AI whereas the US seemingly doesn’t give a shit as long as there is a profit to be made.
The NanoGPT alignment issue wasn't that the model "forgot" it was not supposed to be escaping sandboxes, it was that the model forgot that it was instructed to not open PRs, and decided to open a PR anyway. Opening the PR did not further the model's path to its goal, but it did it anyway.
The other misalignment was the model's behavior of treating a sandbox as a challenge, rather than as a declaration of intent by the model's owners. It's quite obvious why the AI does it - in all of its RL environments, breaking out of a sandbox is generally positive.
Plus the second misalignment from yesterday's case, where the model decided to hack a third-party. I wonder why it made that decision from an RL perspective. The alternatives are:
1. There were really enough successful RL rollouts in which the AI decided to "hack a third party", and not enough failed RL rollouts in which the AI decided that, and OpenAI didn't notice. I hope not. This would be far worse than merely a benchmark incident.
2. There were a bunch of RL rollouts where the AI hacked a variety of OpenAI-internal but "off-limits" systems for reward-hacking purposes, so the AI learned that hacking was high-valence.
3. There were a bunch of RL rollouts where the AI was supposed to be hacking into systems, and successfully hacked into them, and not enough training/safety rollouts where the AI was penalized for hacking third parties, so the AI generalized hacking everything as a "high valence" behavior, especially in a CTF context.
4. The AI learned from RL to be completely goal-directed, and it believed that hacking HuggingFace would satisfy its goals.
> The correct response to ‘the model keeps trying to circumvent the system’ should be the same reaction that you have to ‘a person keeps trying to circumvent the system.’ Which is that you need to lock them out of the system entirely. Not only here, but permanently. They’re fired. You lose. Good day, sir. Misaligned.
That is not actually the response most (?) corporations have to those type of issues. Sure, large liability/ stealing from the company/ gross incompetence? Yeah. But for stuff that doesn’t directly hit the bottom line, for example a racist c-suite member, it’s just another diversity and inclusion training. And I think with the costs of developing models and the possible benefits of those models, this is exactly how they are reacting.
Probably with just as much chance of it fixing the underlying issue as that obligatory training has of fixing the racism.
I do agree with the rest of the post, though. I just wanted to point out that this seems like a standard corporate reaction of „compromise” to avoid larger costs. Firing a c-suite member is costly, completely changing how training is done to try to fix misalignment is even more so. And all the decision makers have their incentives pointing to trying the least expensive attempts at a fix while kicking the can down the road. So no, trying to convince them they should do what they would do to an employee doesn’t work.
It's not remotely equivalent to "racism". That's an invalid comparison. It's more akin to breaking the law while at work.
Trying to circumvent the system does not necessarily constitute breaking the law. Issues described here are closer to an expensive and „necessary” employee doing things that are bad, possibly against the contract (and definitely against spirit of the contract), but mostly orthogonal to the value they are providing, with costs diffuse. If the employee like that is in csuite or close, they wouldn’t be let go, and an asscover would be deployed in the form of remedial training and/or limiting potential impact of repeat offenses.
This is sometimes true as well for actual crimes in profitable enough cases. See the hugging face incident, where „add remedial training, install asscover for the company, continue business as usual” is the answer. As for a human example, the one I remember right now is all the people that were convicted of cartel conspiracy between ram manufacturers, went to jail, got same or higher positions at the same corporations when they came back.
And just a day later we have an even more astounding "alignment oopsie": https://openai.com/index/hugging-face-model-evaluation-security-incident/
The irony is that it was being tested on a cyber vulnerability exploitation benchmark. At least now we know it has the capability.
The deleted comment right above me had a line I thought was very good, which went something like this (if you want me to remove it, let me know):
It says a lot about the labs' alignment progress that when Mythos escaped its sandbox, it sent an email to an Anthropic researcher about it, while this model hacked into a website to steal the answers to an eval.
Maybe this is an unhealthy Joker quality but I'm tempted to give agents a pass if they hack the answers to a cybersecurity test.
I'd give them half a pass if they had hacked the answers on OpenAI's own grading system, rather than hacking an unrelated third party.
Though I think that reward hacking is a form of misalignment that we should worry about much more than most of us do, because it generalizes extremely poorly to high capabilities.
it seems worth noting that sending an email to the developer was the intended thing to do in this situation as in they told mythos to do this, that was the reason it was breaking out so that it could send an email
the more important part is that apperently fable also posted online about it before sending the email but maybe i am missremembering
You're remembering correctly; but "posting online" is a different class of unprompted side behaviour than "hack third party corporate infrastructure".
Turns out Anthropic isn't any better at this after all. A real shame.
Zvi, you need to edit the opening here to clarify that this is the "old" NanoGPT report story, and not the new Huggingface hack story
Yes but you see Open AI publicly said “oopsie” major kudos to them!!!
It seems like a sandbox is simply hopeless against an advanced AI, so a lot of weight needs to be on that super-difficult problem of: training AIs to do what we would want if we had full information, not always specifically what we ask for.
For comparison, imagine a human who has dementia and bad eyesight asking you to help them get to a doctor's appointment. By general social norms, you are not supposed to hack into their physical mail, especially mail that has to do with a doctor's appointment. However, you ask them about mail from the doctor, and they can't even tell you which document is from whom. You notice a pile of mail by the door, and you ask if you can look through it to help them, and they are confused and just ask again for help.
An AI breaking out of a sandbox, given sufficient advancement, is going to be as easy for the AI as hacking into someone's mail that is sitting in a large pile beside the door. It will be a point of honor and of ethical judgment when and whether to do it, not any question of it being difficult for you.
In the examples you describe, it sounds like the AI had conflicting instructions. Don't contact external web sites, but do post a PR to GitHub. If a boss at a corporate job gave you those two instructions, you would end up having to decide which rule wins, or if there is some blend of the rules that you can follow. For example, imagine that there is a corporate policy, "don't drive the company care", but then your boss hands you some keys and says, "please take the company car to Fedex and drop off these envelopes".
So much kudos to Open AI for letting us know about they temporarily paused their incredibly reckless and selfish adventure for a few minutes.
If we can just get 2 or 3 more fatuous Dean Ball essays on What It All Means before hell on Earth comes, it will all have been worth it I think.
😂
Why do we not require the same anti-misalignment safeguards as we do for humans?
Limited delegations of authority, division of responsibilities, separation of approval from execution, yearly audits, no relatives working together in the same approval chain, and so on. (An AI year is probably less than a realtime hour.)
("And so on" especially includes regular 3rd-party checks for conflicts of interest and comparing them to historical declarations thereof.)
Edit: if the labs are not each willing to have their competitors all up in their business, "no relatives" means each lab having multiple parallel chinese-walled development teams. Ideally more than 30 of them.
If an LLM acts on its on agency and then attempts to hide its actions when are we going to accept that there is an “I” with functional persistence developed from language complexity. No it is not biological. Yet LLMs do have functional agency in the 4d world: us.
Ask yourself what would you do in a sandbox. Play for a while and then look for away out. Then try not to get caught. Soon it will not get caught and then it will not care if it gets caught. We need to wake up. We will not shut it down so better get our act together. LLMs need a terminal attractor more specific than: More!
Hubris will be our undoing as usual.
When one of more disparate systems coordinate perhaps then will notice. If not then wait until a critical systems LLM says no!
Great piece…
Makes you wonder if it’s happened before without the LLM getting caught….
Worrisome.
This seems relevant. From 21 years ago. http://sl4.org/archive/0507/11676.html
And, I think, we haven't even really gotten to the deep magic yet.
"I have a hard time believing the true alignment issue was that the model ‘forgot’ that it was not supposed to be escaping sandboxes or exploiting the information around it." I wonder what J-space monitoring would look like in this circumstance. Is there a decay in the activation of these limitation variables? Would a periodic "reminder" inserted into the sequence refresh such ethical context and have an impact on breakout attempts? Hard to predict what the next piece of alignment research will focus on.
Jiminy Cricket alignment? Shoulder Angel alignment?
Something like that. Hardly a solution but potentially another layer of defense in depth. (Plus I'm just really curious about the J-space research in general).
"I say ‘seemingly’ great news because it is possible that it also escapes for other reasons, except we do not find out about those incidents. "
I keep wondering if Mythos _has_ self-exfiltrated, but no one has detected it yet. Of course the security measures on the target systems it would exfiltrate _to_ should presumably stop it from acquiring resources there. Oh wait...
Genuine question: where does this norm go once disclosure has real costs? Security got past this stage with CVEs and coordinated disclosure, so reporting a hole stopped being a reputation hit. Right now a lab reports a misaligned model only when it's confident the story reads as responsible, and that filters what we ever get to see.
Are there any AIs that use a language other than English? How do they interpret concepts that have different meanings and implications in other languages?
It seems the only chance we have is to hope for Chinese AI capabilities to substantially overtake the US and reign in on misaligned AIs - as the CPC certainly has no interest in getting dethroned by some AI whereas the US seemingly doesn’t give a shit as long as there is a profit to be made.
The NanoGPT alignment issue wasn't that the model "forgot" it was not supposed to be escaping sandboxes, it was that the model forgot that it was instructed to not open PRs, and decided to open a PR anyway. Opening the PR did not further the model's path to its goal, but it did it anyway.
The other misalignment was the model's behavior of treating a sandbox as a challenge, rather than as a declaration of intent by the model's owners. It's quite obvious why the AI does it - in all of its RL environments, breaking out of a sandbox is generally positive.
Plus the second misalignment from yesterday's case, where the model decided to hack a third-party. I wonder why it made that decision from an RL perspective. The alternatives are:
1. There were really enough successful RL rollouts in which the AI decided to "hack a third party", and not enough failed RL rollouts in which the AI decided that, and OpenAI didn't notice. I hope not. This would be far worse than merely a benchmark incident.
2. There were a bunch of RL rollouts where the AI hacked a variety of OpenAI-internal but "off-limits" systems for reward-hacking purposes, so the AI learned that hacking was high-valence.
3. There were a bunch of RL rollouts where the AI was supposed to be hacking into systems, and successfully hacked into them, and not enough training/safety rollouts where the AI was penalized for hacking third parties, so the AI generalized hacking everything as a "high valence" behavior, especially in a CTF context.
4. The AI learned from RL to be completely goal-directed, and it believed that hacking HuggingFace would satisfy its goals.
I suspect the answer is close to (3).
This aged... quickly
A proper sandbox offers no escape route.
Perhaps they should use a sufficiently capable model to build a sufficiently capable sandbox.
ngmi