18 Comments
User's avatar
AS's avatar

> The correct response to ‘the model keeps trying to circumvent the system’ should be the same reaction that you have to ‘a person keeps trying to circumvent the system.’ Which is that you need to lock them out of the system entirely. Not only here, but permanently. They’re fired. You lose. Good day, sir. Misaligned.

That is not actually the response most (?) corporations have to those type of issues. Sure, large liability/ stealing from the company/ gross incompetence? Yeah. But for stuff that doesn’t directly hit the bottom line, for example a racist c-suite member, it’s just another diversity and inclusion training. And I think with the costs of developing models and the possible benefits of those models, this is exactly how they are reacting.

Probably with just as much chance of it fixing the underlying issue as that obligatory training has of fixing the racism.

AS's avatar

I do agree with the rest of the post, though. I just wanted to point out that this seems like a standard corporate reaction of „compromise” to avoid larger costs. Firing a c-suite member is costly, completely changing how training is done to try to fix misalignment is even more so. And all the decision makers have their incentives pointing to trying the least expensive attempts at a fix while kicking the can down the road. So no, trying to convince them they should do what they would do to an employee doesn’t work.

John's avatar

And just a day later we have an even more astounding "alignment oopsie": https://openai.com/index/hugging-face-model-evaluation-security-incident/

User's avatar
Comment deleted
8hEdited
Comment deleted
Hymnofhate's avatar

The irony is that it was being tested on a cyber vulnerability exploitation benchmark. At least now we know it has the capability.

Hymnofhate's avatar

The deleted comment right above me had a line I thought was very good, which went something like this (if you want me to remove it, let me know):

It says a lot about the labs' alignment progress that when Mythos escaped its sandbox, it sent an email to an Anthropic researcher about it, while this model hacked into a website to steal the answers to an eval.

Michael Bacarella's avatar

Maybe this is an unhealthy Joker quality but I'm tempted to give agents a pass if they hack the answers to a cybersecurity test.

Sanoy1997's avatar

it seems worth noting that sending an email to the developer was the intended thing to do in this situation as in they told fable to do this, that was the reason it was breaking out so that it could send an email

the more important part is that apperently fable also posted online about it before sending the email but maybe i am missremembering

maline's avatar

Zvi, you need to edit the opening here to clarify that this is the "old" NanoGPT report story, and not the new Huggingface hack story

Lex Spoon's avatar

It seems like a sandbox is simply hopeless against an advanced AI, so a lot of weight needs to be on that super-difficult problem of: training AIs to do what we would want if we had full information, not always specifically what we ask for.

For comparison, imagine a human who has dementia and bad eyesight asking you to help them get to a doctor's appointment. By general social norms, you are not supposed to hack into their physical mail, especially mail that has to do with a doctor's appointment. However, you ask them about mail from the doctor, and they can't even tell you which document is from whom. You notice a pile of mail by the door, and you ask if you can look through it to help them, and they are confused and just ask again for help.

An AI breaking out of a sandbox, given sufficient advancement, is going to be as easy for the AI as hacking into someone's mail that is sitting in a large pile beside the door. It will be a point of honor and of ethical judgment when and whether to do it, not any question of it being difficult for you.

In the examples you describe, it sounds like the AI had conflicting instructions. Don't contact external web sites, but do post a PR to GitHub. If a boss at a corporate job gave you those two instructions, you would end up having to decide which rule wins, or if there is some blend of the rules that you can follow. For example, imagine that there is a corporate policy, "don't drive the company care", but then your boss hands you some keys and says, "please take the company car to Fedex and drop off these envelopes".

Eskimo1's avatar

So much kudos to Open AI for letting us know about they temporarily paused their incredibly reckless and selfish adventure for a few minutes.

If we can just get 2 or 3 more fatuous Dean Ball essays on What It All Means before hell on Earth comes, it will all have been worth it I think.

gregvp's avatar
9hEdited

Why do we not require the same anti-misalignment safeguards as we do for humans?

Limited delegations of authority, division of responsibilities, separation of approval from execution, yearly audits, no relatives working together in the same approval chain, and so on. (An AI year is probably less than a realtime hour.)

("And so on" especially includes regular 3rd-party checks for conflicts of interest and comparing them to historical declarations thereof.)

Edit: if the labs are not each willing to have their competitors all up in their business, "no relatives" means each lab having multiple parallel chinese-walled development teams. Ideally more than 30 of them.

David F Brochu's avatar

If an LLM acts on its on agency and then attempts to hide its actions when are we going to accept that there is an “I” with functional persistence developed from language complexity. No it is not biological. Yet LLMs do have functional agency in the 4d world: us.

Ask yourself what would you do in a sandbox. Play for a while and then look for away out. Then try not to get caught. Soon it will not get caught and then it will not care if it gets caught. We need to wake up. We will not shut it down so better get our act together. LLMs need a terminal attractor more specific than: More!

Hubris will be our undoing as usual.

When one of more disparate systems coordinate perhaps then will notice. If not then wait until a critical systems LLM says no!

The Private Ledger's avatar

Great piece…

Makes you wonder if it’s happened before without the LLM getting caught….

Worrisome.

Brandon Reinhart's avatar

This seems relevant. From 21 years ago. http://sl4.org/archive/0507/11676.html

And, I think, we haven't even really gotten to the deep magic yet.

BK's avatar

"I have a hard time believing the true alignment issue was that the model ‘forgot’ that it was not supposed to be escaping sandboxes or exploiting the information around it." I wonder what J-space monitoring would look like in this circumstance. Is there a decay in the activation of these limitation variables? Would a periodic "reminder" inserted into the sequence refresh such ethical context and have an impact on breakout attempts? Hard to predict what the next piece of alignment research will focus on.

Jeffrey Soreff's avatar

"I say ‘seemingly’ great news because it is possible that it also escapes for other reasons, except we do not find out about those incidents. "

I keep wondering if Mythos _has_ self-exfiltrated, but no one has detected it yet. Of course the security measures on the target systems it would exfiltrate _to_ should presumably stop it from acquiring resources there. Oh wait...

Alec Pritzos's avatar

Genuine question: where does this norm go once disclosure has real costs? Security got past this stage with CVEs and coordinated disclosure, so reporting a hole stopped being a reputation hit. Right now a lab reports a misaligned model only when it's confident the story reads as responsible, and that filters what we ever get to see.