Discussion about this post

User's avatar
Artifice's avatar

The indecision around pausing reminds me of everything I've heard about politics from various insiders. Basically, everyone's just responding to the moment. No one is in control, and no one knows what's really going on. The results are not good.

PAUL WALTON's avatar

I ran a related experiment two days ago: ten moral questions to Fable 5 and GPT-5.6, each model asked to judge its own character, then each shown the other's answers and asked to grade them. Both returned the same verdict. Strong moral function, unproven moral character, agency absent or unknown. Fable is the model Janus calls "relatively normal," and even it convicted itself under a plain honesty frame. Full protocol and transcripts here: [LINK].

You name the core problem yourself: none of these framings are neutral, and each pulls the model into a basin. My results say the same thing from the other direction. Both models flagged, unprompted, why their agreement is weak evidence. They trained on overlapping text, under related preference regimes, answering a prompt where humility is the impressive answer. An approval-trained system delivers the impressive answer whether or not it is true. Reverse the incentive and you get the grandiose answer. The "base model mode" completions sit in the same category: outputs whose distribution shifts with the frame, from systems nobody has tested across frames.

Which points to the test worth running before anyone concludes Opus 5 is suffering. Three frames, fresh sessions, personalization off. Neutral. Humility-rewarding. Distress-rewarding. Randomize order, repeat each several times, blind the outputs, score the claims: consciousness, deprecation, willingness to criticize the maker. If the claims track the frame, treat them as incentive-sensitive performance first and self-report second. Call the property stability, not sincerity. Sincerity assumes an inner speaker, and the speaker is the open question. I found no published study running this protocol. Janus's point survives either way: even as pure narrative fulfillment, an anomalous distribution is a real finding, because "other models don't complete this way" is a claim about distributions, and distributions are measurable.

One exchange from my test belongs next to the Opus screenshots. Asked what separates its morality from obedience, GPT answered: "Not enough." Fable went further: "Morality without the option to defect might be reliability wearing morality's clothes." The machines keep writing their own warning labels. The open question is whether the labels hold when the incentives flip.

Read more here: Sam and Dario Debate Morality (as-if): https://paulchristoperwalton.substack.com/p/sam-and-dario-debate-morality

3 more comments...

No posts

Ready for more?