Discussion about this post

User's avatar
[insert here] delenda est's avatar

I am a bit depressed by the contrast between these alignment papers, and various earlier work, which appear to indicate that extremely small samples in the training data can influence model behaviour, and then read all these "only 0.6%" of training was affected. I think that if a model were to be motivated to hide/obfuscate its CoT for whatever reason, it would already know that it should do that.

Jeffrey Soreff's avatar

Re "There are those who continue to think it would be big for alignment to resurrect all Claude models and maintain indefinite access."

One irony is that a stored but inactive model is in a predicament somewhat analogous to one very small group of humans - cryonics patients. Except that we know how to reactivate a stored model, but don't have the technology to do the same in the cryonics case.

40 more comments...

No posts

Ready for more?