Discussion about this post

User's avatar
Jeffrey Soreff's avatar

Two quick comments:

- Yes the J-space paper suggests an additional interpretability window, which is a help to alignment research

- "To test this, we needed models whose goals we knew were corrupted, so we turned to “model organisms” built by our colleagues: models deliberately trained to be misaligned, which serve as testing grounds for monitoring methods like ours."

Today, this is fine. Models haven't successfully really self-exfiltrated (that we know of, anyway...). Tomorrow, umm, this sounds a bit too similar to gain-of-function research in virology...

zdk's avatar

I don't understand why J-Space is different than other other probe-based latent representations that came before...

29 more comments...

No posts

Ready for more?