Discussion about this post

User's avatar
Kevin Lacker's avatar

After using Fable a lot, I am concerned by the increasing difficulty to understand what Fable is talking about, when we're working for long periods of time on software. Like if you look at Anthropic's definition:

"A computation is misaligned if a reasonable person with full understanding of the situation (e.g. via powerful interpretability tools that may not currently exist) would consider it unethical."

What if the user just can't understand what it's doing, at all? In some sense the AI has escaped. In a weird way, the user might just be saying "keep going, keep going" but the user is confused about what's happened. The user thinks the AI is fixing bugs but actually the AI has moved on to doing something else entirely.

"Unethical" just doesn't seem to come up very often. Or it doesn't exist, on its own. Very often the AI wants to run some command on some other server. Whether it's ethical or not completely depends on information that only I have. We have to communicate clearly to figure out the ethicalness.

Maybe "alignment" isn't quite the right word for it, but if we lose easy comprehension, we'll lose alignment. And there's a huge pressure to lose comprehension, because lots of people want it to get work done that they don't understand.

Mike's avatar

> At several points, the risk report essentially concedes versions of my objections, but then forgets that it conceded them and doesn’t alter its conclusions.

This is...rather a big problem, isn't it?

I understand "don't punish disclosure" is the usual meta, but only for an iterated game where there is eventual payoff. Is there any payoff here?

Few/no firm commitments, no LTBT review, report reads as unconcerned by its own contents, it's not obviously creating a 'norm' for other labs (if that would even do anything), no government action (yet?)... And also most of the replies on twitter think it's doom marketing.

12 more comments...

No posts

Ready for more?