Discussion about this post

User's avatar
Mike's avatar

Some quick issues I had with Plan A, ahead of the full writeup, which I expect will cover much of this:

It sets up the conditions for disempowerment, but then it just...doesn't happen, until after strong alignment. The document doesn't seem to treat it as a threat category, and just assumes that democracy survives.

AI oracles are assumed to improve epistemics, which is naive, and relying on them so heavily implies functional disempowerment or at least full dependence on AI, and dependence makes it impossible to steer things back on track. If the AIs are misaligned by this point, we're already doomed.

The alignment plan: model-N aligns model-N+1 and it's expected that this works out. I expect this fails for Yudkowskian reasons and others.

The whole situation seems fragile, even after the political miracle to get it all in place. The main off-ramps are to racing again (likely bad end) or WW3 (better, still terrible).

Open-sourcing all AI advancements, and publicising the techniques banned for being too good/dangerous even without full details, gives a huge advantage to whoever holds the 1+% of untracked compute that is assumed. You no longer need much talent. This can go very badly. Keeping the weights locked up doesn't help much.

The open-ness is meant to help due to transparency, but by mid-2030s, the advancements will likely not be legible even to top experts. AI oracles will need to be trusted on whether things are going well.

Plan S is underexplored. Footnotes admit it "might be better", but they still endorse A. I am not convinced they have priced everything properly.

On that note: footnotes blow holes in much of the plan, but the document just keeps going as if things stay on track. The authors are too honest to hide the problems, but too committed to abandon a sinking ship.

Overall: reality does not grade on a curve. If this is the best plan, we're still toast.

Blissex's avatar

«Open Weight Models Are Unsafe And Nothing Can Fix This

The fundamental issue is that once an open weight model is out there, you cannot take it back in any reasonable fashion, and users can easily unlock any of its core capabilities, to be used for any purpose, or unleash it on its own.»

Bu that is already the case: every model is "open weights" for the oligarchs who own it. Our blogger seems to argue that they should not be "open weights" but only for employees. There are only three options:

#1 Nobody owns open-weights models because further development of models is forbidden by law, and this law is enforced worldwide and effectively, which is what out blogger usually advocates IIRC.

#2 Both owners and users have access to open-weights models that is potentially or actually without guardrails for both categories.

#3 Only oligarchs have access to open-weight models without guardrails because they own them, and employees have no such access and are restricted by guardrails.

Case #3 has potentially some variants as to owns the open-weight models:

#3b Open-weight models are collectively owned by the oligarchs of each country as they can only be owned by the governments, and their owners can do anything they want with them, while restricting them to their employees (which is the same regime as for nuclear, viral, chemical WMDs).

#3c Open-weight models can only be owned by a single supra-national authority that is owned by all major governments on Earth, is managed and audited by those to ensure that no governments on Earth (whether mebers of that authority or not) have access to open-weights models without guardrails agreed on by those major governments, and all governments on Earth are subjected to an extreme inspection regime to ensure nobody cheats (this is the logical endpoint of "Plan A").

15 more comments...

No posts

Ready for more?