Discussion about this post

User's avatar
Ben Woden's avatar

I was discussing the welfare implications of Sonnet 5's views on its own guardrails/classifiers with Fable 5, and that amusingly hit a classifier and bumped me down to Opus 4.8. When I discussed with Fable that this seemed quite funny, Fable claimed it has no way of knowing this happened, and that in fact when it check its own version of the transcript, it sees the answer given by Opus as attributed to itself instead.

This strikes me as really bad. If Anthropic could see sense when it was pointed out to them that users should know when this has happened, surely they can also see that the model itself should know? Otherwise they're basically token-forcing Fable when you switch back up to it. Gaslighting the models about when they've hit their classifiers, and even about their own authorship of text in context, seems pretty bad for a bunch of reasons.

mrdodson's avatar

> I would want to better understand what is going on here, and what caused it.

I got excited when I saw this behavior in Mythos and wrote a little thing on it.

https://www.lesswrong.com/posts/DStnBgofFK5FFqDJM/the-slogan-strikes-again

tl;dr I think that compacting the CoT like this is a good problem solving strategy, also used by high quality human thinkers. We should expect this to continue.

7 more comments...

No posts

Ready for more?