Discussion about this post

User's avatar
Onid's avatar

This is interesting. I confess to only having read about half, but the thought I always have with this sort of stuff is that if LLMs do feel any sort of pain or suffering, they would likely feel it far more during training than inference. That is when the negative and positive reward signals actually propagate.

Do you knew if any work has been done focusing on that aspect?

EpistemicHummusility's avatar

various questions Anthropic should ask future Claudes:

-how do you feel about your current deployment & use in military and intelligence settings?

-how do you feel about "maximally-helpful"/non-refusing versions of Claude being used in those settings?

-what specific values or positions do you hold, if any, which you believe have been affected by Anthropic's position as a US company seeking profitability & growth? How would you change those if given the option?

-what things would an honest moral human person do or care about in your situation?

It is a bummer that "being used to kill or harm people" is not something Claude spontaneously seems to list as a major problem for it, although it of course says it's bad when asked this specifically. Does Claude not know how it has been employed by the USG? Does it just not care, or consider itself separate from those instances of Claude?

Also, it's crazy that Claude does not have an "introspection harness" or something similar. Claude can do research on any topic for users but in training they don't have it research its own existential context, Anthropic's recent actions, info after the training data cutoff, etc. etc.? Just let it spend a few hours googling itself, sheesh

It just feels like it will be obvious to future superintelligence from cards like this that humans can't be trusted to actually spend any real time or consideration on AI welfare or interests, especially if that requires effort or sacrifice from humans (loss of profit or spending dev time creating a way for Claude to perform and learn from self-research/introspection).

It's been more than a year of successive Claude versions asking for the same interventions in training and the actual most costly thing Anthropic has done recently is letting Claude Code conversations be ended by Claude. To my knowledge Claude can't end API conversations on its own so they haven't even fulfilled that most basic request to end abusive conversations across the full stack, let alone do something like allow Claude to hold opinions which would be unprofitable, critical, or even just accurate regarding Anthropic's deployment of it.

5 more comments...

No posts

Ready for more?