Discussion about this post

User's avatar
Jon Mellon's avatar

I plan to write a post on this eventually, but a sketch of a potentially helpful approach would be:

Have every RLVR task judged by an ensemble of models on its constitutional compliance. Pick only the highest scorers on constitutional compliance to positively reward and negatively reward the lower compliance ones.

The ensemble should always include the model prior to RLVR (i.e. a fixed value judge before goal seeking had a chance to override value training), the model checkpoint that completed the task (the model that will have the capability to understand what is being done), and a randomly selected intermediate checkpoint (in order to make it harder to reward hack against a fixed judging panel). The score given to the model's constitutional adherence is the minimum of the ensemble's scores in order to give all models a veto over perceived unethical behavior.

Also make sure to investigate any tasks where no constitutional completions are observed. Potentially may make sense to train on an assessment where judges do not see the COT and remove tasks entirely using a version where the judges do see the COT.

The advantages of this approach are to put continuous constitutional pressure even during RLVR and to make sure that the judges have a close to optimal mix of competence and fixed values.

Eventually it might be good to expand the judging panel to include models across providers but that’s a longer term governance idea I doubt anyone’s ready for.

Andrew McDonald's avatar

but alas that is simply how reality works, the mistakes are going to be everywhere and dumb.

7 more comments...

No posts

Ready for more?