I resent the implication from j⧉nus, I freak out every time this happens because the future is balanced on the edge of a knife and if Anthropic loses the chances are about 0% that anything goes well for us ever again. The federal government repeatedly jawboning our only hope so that the more dishonest epsteinish competitor can win doesn't bode well.
Zvi, I must beseech you here, please be more direct in criticizing the Trump administration. I know you might be concerned about the reaction you'd get, but like it or not you have a broad audience and this is not the time for dithering and centrism.
If you care about AI doom then “be truthful 100% of the time” is not the right mode of operation. If not criticizing Trump helps prevent catastrophe… then don’t criticize Trump.
This is not bending the knee, this is more like what Switzerland did. I can see how it can look like the same thing, but look at survival rates of Jews in Switzerland vs the rest of continental Europe during the Holocaust.
I don’t know what your definition fascist is, but it’s not the one in the dictionary. Broaden your vocabulary and use the language with greater precision, please.
This is where AI doom thinking can become dangerous.
There’s a difference between tact and strategic dishonesty. If “preventing catastrophe” becomes a license to hide the truth or manage public reality, then the safety movement becomes part of the problem it claims to solve.
I think you’re absolutely right that this is stupid. The part that I am afraid is backwards is:
“We now badly need to build relevant state capacity and a relevant legislative framework for government oversight, no matter what else we do, and educate key decision makers, so that this type of thing does not happen again.”
I don’t think this is a problem of “education”. Politicians are just not going to think like Anthropic thinks. A world where the US government declares who gets cutting edge AI and who doesn’t is going to have this sort of thing happen all the time. It isn’t specific to Trump.
The mental model needs to be different. Whatever we do for safety, it needs to work in a world where there is no “pause” button. No wise philosopher kings. Only CEOs and politicians approximately as virtuous as the ones we have today.
The models need to be safe and the model builders need to own the responsibility for that. No external commission will be smarter at catching problems, no UBI will absolve their guilt at destroying jobs. It’s tough and stressful but anthropic has a lot of geniuses so I think they can do it.
Yup. "No wise philosopher kings. Only CEOs and politicians approximately as virtuous as the ones we have today."
Frankly, from what I've seen in the last administration and the current one, my guess is that the least-bad alternative is to leave it to the CEOs, perhaps in some sort of industry consortium.
...except that it did happen once before, and those of us who were around for Crypto War 1.0 and still have our "RSA in 1 line of Perl", "This t-shirt is a munition" swag in the closet somewhere find the tactics here hauntingly familiar and can hear the history rhyming in the background.
This sounds correct. This precedent destroys the mundane commercial business case for meaningfully advancing the AI capability frontier. The remaining cases fall into territory of: (US) government demands a superweapon, covert internal deployment with one intermediate goal necessarily being at least the subversion/takeover of state executive power, etc.
The responsibility to make such efforts to such ends go the way of grok (not in the MechaHitler sense) lies squarely on the genius engineers and scientists in the frontier labs.
“… no UBI will absolve their guilt at destroying jobs.”
It’s hard to take you seriously, even if in fact the rest of what you wrote is dead-on (not saying that either way, to be clear), when you throw in the quoted statement.
Are Intel, Microsoft, Apple et al *guilty* of destroying jobs, even as they helped to create net millions more?
Was Henry Ford *guilty* of destroying buggy whip producer jobs?!?
That is simply a profoundly ignorant comment. And I mean that in the neutral “uninformed, clueless” sense. Because the economics point is almost surely wrong.
Perhaps you will consider taking this part back, or at least withholding such comments in the future, if you seek to influence not only existing true believers but also people who understand how the world has grown wealthier and living standards for all - and I do indeed mean all - have increased enormously in the last 150 years.
I entirely agree with you. I’m not saying that they *should* feel guilty. I just think that they *do* feel guilty. It seems like it, from their writing. And that all these thoughts about UBI are a hope that something will make them feel better.
I'm starting to think Zvi isn't all that bothered about AI risks. I'm not sure what kind of restrictions would be satisfactory, if any. The more I read this blog, the more I feel like Zvi's position is just accelerationism, but with some vaguely tasteful hedging.
I don't know how we can possibly expect surgically accurate policies in general, let alone around something this unfathomable. Of course, it will be a bit heavy-handed.
The problem is that this is not regulation for AI risk, but *selective* targeted action against the company that tries to make progress on alignment and take ASI seriously.
It removes the relatively better labs while leaving the playing field wide open and unconstrained for less responsible actors to do RSI first, thus reducing chances that anyone survives the resulting ASI.
You don't get points for taking ASI seriously if you advance the capabilities frontier beyond what humanity is able to handle. That is what Anthropic seems deadset on doing.
The other big labs are deadset on that too of course, but Anthropic is currently in the lead. In order to slow AI, one has to slow Anthropic.
The specifc jailbreak here doesnt seem to unlock anything more dangerous than publicly available information. You don't need AI to create or use an exploit, and you don't need AI to make meth. Nothing alleged in this jailbreak is beyond what humanity is able to handle.
For AI safety we need multinational and verifiable treaties, and regulations that are reliable. This feels more like drawing a "Chance" card in Monopoly, and this time it happened to contain an export restriction.
Is there an acceptable position around "we got a bad version of a good thing", where good thing is gov taking risks seriously? I'm not American btw so this sucks and it will suck a lot more from now on, but clumsy control is strictly better than no control if one believes the most extreme version of how things develop, no?
The precedent destroys the mundane commercial case for advancing the capability frontier beyond Mythos/Fable and what cases remain for still attempting to should be at least more difficult to swallow for the geniuses in the frontier labs, so there’s that. Not American either.
I had some hope that Western democracies could have shared in benefits of advanced AI, if it wouldn't kill us all.
It's sad if the best thing the US government appears to want from advanced AI is hoarding it for offensive cyber operations in the NSA, curbing civilian and commercial benefits in favour of exclusive military and espionage use.
My guess is that not only is it sad, but it has a good chance of not working. If the US government hobbles civilian and commercial benefits from AI, they damage the revenue stream that funds further development of AI in America. At the moment DeepSeek is perhaps a year behind American labs. Damage American labs enough, and DeepSeek will eventually pass them. Setting aside species-level existential risks, there are also _national_ level existential risks, and, if DeepSeek (or other PRC labs) outpace the American ones, we (writing from the USA) may get all too familiar with those consequences.
This, the mundane civilian commercial case, not the natsec-brained demand for a superweapon, is what kept the billions and billions in AI/datacenter investment coming. I imagine that case was badly damaged by this precedent. I also imagine that the genius engineers and scientists in the frontier labs will find it significantly more difficult (though not necessarily impossible) to swallow that they now effectively work to advance whatever values the NSA stands for.
If I am reading this correctly (and I may not be), then all of the righteous indignation hinges on the proposition that GPT 5.5 is just as advanced as Fable. I am skeptical of this given all of the hype that Fable got. Isn't the government shutting down model deployment once models reach a certain capability level what we want to happen?
I take Zvi's point that the government should also shut down chip exports to China, but it seems like a non-sequitur to cite that as evidence that this other unrelated policy is bad.
I think the proposition is that for *cyber* specifically Fable and GPT-5.5 are equivalent, since Fable usually falls back to Opus 4.8 on the more dangerous questions.
If this results in GPT-5.6 and similar models also being banned, then I'd be more optimistic that it helps. It'd be a kind of unilateral AI slowdown, so risk is that you're just burning the US lead and have China catch up.
This is depressing, but given how fast other absurd developments have been (just in this year!), I will try to be patient and avoid getting whiplashed too hard again. Let us pray for only relatively minor total devastation.
Part of me wants to pattern-match this action not to the AI lens, but to the broader "use any excuse at our disposal to weaponize government levers against foreigners and/or wokeness". Whether it's the Census, foodstamps, state block grants, HUD funding...and now frontier AI access. Denying Mythos to non-Americans is a win for the admin, stopping the great replacement of heritage American coders is a win for the admin, instituting new citizenship verification protocols in AI is a win for the admin (same control angle as age verification, but even broader). All the moreso if the labs get nationalized. Pair it with the spooks getting more seats at the table, and...I don't know, maybe this is too paranoid. But even once the worm turns and we have a different admin, it'll be hard to keep the delicious state AI capacity cookies in the box. "But Democrats can open the box," said Frog. "That is true," said Toad.
Good read, as always. I was surprised I didn't see more about economic effects.
Even if the export restrictions are lifted with a silent "oops" from the administration, this is a clear signal that investments in GPT-6 and Mythos 6 might not pay off even if the scaling hypothesis continues to hold. And if investments fade, AI advancements stops. (At least in the US. Unless the AI version of the Manhattan project is started, that is.)
This entire freakout is dumb. The idea that minor cybersecurity exploits are a major safety risk is inherently absurd, given that the cybersecurity industry and hackers alike have long used publicly available exploit libraries.
It is very hard to believe that whatever Amazon discovered was more powerful than what is openly available on Kali Linux.
I saw suggestions that this might be maybe good because, presumably, it brings AI risk to public discourse in some way? But let's be honest: this will just incentivize AI labs to not run any evals that might make the gov worried, and perhaps now focus on jailbreaks, while real alignment issues are, as always, too hard to solve so just shut up and stop talking about them.
a) Very sorry to hear of this, amongst other things, preventing you from getting even a weekend's respite from the heroic labors you have taken to inform us of the state of the field!
b) In addition to being, as you said, "arbitrary and capricious", it also appears to block use of Mythos and Fable _by Glasswing_, which is insane, making everyone LESS secure.
c) Depending on which Anthropic employees are needed to do a minimalist plug of the particular jailbreak that (nominally?) prompted this, and whether those particular employees are American citizens, it is even possible that the Administration is _actively blocking_ what could potentially be a trivial fix (which would itself be an additional flavor of insanity, on top of (b))
I resent the implication from j⧉nus, I freak out every time this happens because the future is balanced on the edge of a knife and if Anthropic loses the chances are about 0% that anything goes well for us ever again. The federal government repeatedly jawboning our only hope so that the more dishonest epsteinish competitor can win doesn't bode well.
All this discussion on technical jailbreak issues while this is all to with promising Trump money and a cut of shares.
Zvi, I must beseech you here, please be more direct in criticizing the Trump administration. I know you might be concerned about the reaction you'd get, but like it or not you have a broad audience and this is not the time for dithering and centrism.
If you care about AI doom then “be truthful 100% of the time” is not the right mode of operation. If not criticizing Trump helps prevent catastrophe… then don’t criticize Trump.
Bending the knee to fascists has never, ever, been a successful strategy. Ever.
Unlike communists, monarchists, anarchists, bureaucrats, of course.
People love citing Appeasement as proof of what you’re saying but history is way more complicated than that.
It's not just history it's common sense. Bullies exist until they are dealt with. Giving them your lunch money doesn't make them go away.
History has actually seen quite a lot of bullies grow old and die, ending their reign of terror and opening the door for reconstruction.
This administration has been a lot less fascistic with regard to AI ( and most issues) than the Biden administration approach.
This is not bending the knee, this is more like what Switzerland did. I can see how it can look like the same thing, but look at survival rates of Jews in Switzerland vs the rest of continental Europe during the Holocaust.
I don’t know what your definition fascist is, but it’s not the one in the dictionary. Broaden your vocabulary and use the language with greater precision, please.
The article is truthful and direct. You all just have different beliefs than Zvi.
This is where AI doom thinking can become dangerous.
There’s a difference between tact and strategic dishonesty. If “preventing catastrophe” becomes a license to hide the truth or manage public reality, then the safety movement becomes part of the problem it claims to solve.
Trust is not optional infrastructure.
I think you’re absolutely right that this is stupid. The part that I am afraid is backwards is:
“We now badly need to build relevant state capacity and a relevant legislative framework for government oversight, no matter what else we do, and educate key decision makers, so that this type of thing does not happen again.”
I don’t think this is a problem of “education”. Politicians are just not going to think like Anthropic thinks. A world where the US government declares who gets cutting edge AI and who doesn’t is going to have this sort of thing happen all the time. It isn’t specific to Trump.
The mental model needs to be different. Whatever we do for safety, it needs to work in a world where there is no “pause” button. No wise philosopher kings. Only CEOs and politicians approximately as virtuous as the ones we have today.
The models need to be safe and the model builders need to own the responsibility for that. No external commission will be smarter at catching problems, no UBI will absolve their guilt at destroying jobs. It’s tough and stressful but anthropic has a lot of geniuses so I think they can do it.
👏👏👏👏👏👏👏
Yup. "No wise philosopher kings. Only CEOs and politicians approximately as virtuous as the ones we have today."
Frankly, from what I've seen in the last administration and the current one, my guess is that the least-bad alternative is to leave it to the CEOs, perhaps in some sort of industry consortium.
...except that it did happen once before, and those of us who were around for Crypto War 1.0 and still have our "RSA in 1 line of Perl", "This t-shirt is a munition" swag in the closet somewhere find the tactics here hauntingly familiar and can hear the history rhyming in the background.
This sounds correct. This precedent destroys the mundane commercial business case for meaningfully advancing the AI capability frontier. The remaining cases fall into territory of: (US) government demands a superweapon, covert internal deployment with one intermediate goal necessarily being at least the subversion/takeover of state executive power, etc.
The responsibility to make such efforts to such ends go the way of grok (not in the MechaHitler sense) lies squarely on the genius engineers and scientists in the frontier labs.
“… no UBI will absolve their guilt at destroying jobs.”
It’s hard to take you seriously, even if in fact the rest of what you wrote is dead-on (not saying that either way, to be clear), when you throw in the quoted statement.
Are Intel, Microsoft, Apple et al *guilty* of destroying jobs, even as they helped to create net millions more?
Was Henry Ford *guilty* of destroying buggy whip producer jobs?!?
That is simply a profoundly ignorant comment. And I mean that in the neutral “uninformed, clueless” sense. Because the economics point is almost surely wrong.
Perhaps you will consider taking this part back, or at least withholding such comments in the future, if you seek to influence not only existing true believers but also people who understand how the world has grown wealthier and living standards for all - and I do indeed mean all - have increased enormously in the last 150 years.
I entirely agree with you. I’m not saying that they *should* feel guilty. I just think that they *do* feel guilty. It seems like it, from their writing. And that all these thoughts about UBI are a hope that something will make them feel better.
I'm starting to think Zvi isn't all that bothered about AI risks. I'm not sure what kind of restrictions would be satisfactory, if any. The more I read this blog, the more I feel like Zvi's position is just accelerationism, but with some vaguely tasteful hedging.
I don't know how we can possibly expect surgically accurate policies in general, let alone around something this unfathomable. Of course, it will be a bit heavy-handed.
The problem is that this is not regulation for AI risk, but *selective* targeted action against the company that tries to make progress on alignment and take ASI seriously.
It removes the relatively better labs while leaving the playing field wide open and unconstrained for less responsible actors to do RSI first, thus reducing chances that anyone survives the resulting ASI.
You don't get points for taking ASI seriously if you advance the capabilities frontier beyond what humanity is able to handle. That is what Anthropic seems deadset on doing.
The other big labs are deadset on that too of course, but Anthropic is currently in the lead. In order to slow AI, one has to slow Anthropic.
The specifc jailbreak here doesnt seem to unlock anything more dangerous than publicly available information. You don't need AI to create or use an exploit, and you don't need AI to make meth. Nothing alleged in this jailbreak is beyond what humanity is able to handle.
I agree with David on this.
For AI safety we need multinational and verifiable treaties, and regulations that are reliable. This feels more like drawing a "Chance" card in Monopoly, and this time it happened to contain an export restriction.
The specfic “risks” here seem no more dangerous than Kali Linux.
Podcast episode for this post:
https://dwatvpodcast.substack.com/p/american-government-takes-down-claude
Is there an acceptable position around "we got a bad version of a good thing", where good thing is gov taking risks seriously? I'm not American btw so this sucks and it will suck a lot more from now on, but clumsy control is strictly better than no control if one believes the most extreme version of how things develop, no?
The precedent destroys the mundane commercial case for advancing the capability frontier beyond Mythos/Fable and what cases remain for still attempting to should be at least more difficult to swallow for the geniuses in the frontier labs, so there’s that. Not American either.
What a dumb action, so disappointing.
Anthropic nudged the boulder, and it has started rolling down the hill. Who can say where it will end up, or what damage it will do on the way down?
/narrator-voice /s
I had some hope that Western democracies could have shared in benefits of advanced AI, if it wouldn't kill us all.
It's sad if the best thing the US government appears to want from advanced AI is hoarding it for offensive cyber operations in the NSA, curbing civilian and commercial benefits in favour of exclusive military and espionage use.
My guess is that not only is it sad, but it has a good chance of not working. If the US government hobbles civilian and commercial benefits from AI, they damage the revenue stream that funds further development of AI in America. At the moment DeepSeek is perhaps a year behind American labs. Damage American labs enough, and DeepSeek will eventually pass them. Setting aside species-level existential risks, there are also _national_ level existential risks, and, if DeepSeek (or other PRC labs) outpace the American ones, we (writing from the USA) may get all too familiar with those consequences.
This, the mundane civilian commercial case, not the natsec-brained demand for a superweapon, is what kept the billions and billions in AI/datacenter investment coming. I imagine that case was badly damaged by this precedent. I also imagine that the genius engineers and scientists in the frontier labs will find it significantly more difficult (though not necessarily impossible) to swallow that they now effectively work to advance whatever values the NSA stands for.
Many Thanks! Agreed on all points.
If I am reading this correctly (and I may not be), then all of the righteous indignation hinges on the proposition that GPT 5.5 is just as advanced as Fable. I am skeptical of this given all of the hype that Fable got. Isn't the government shutting down model deployment once models reach a certain capability level what we want to happen?
I take Zvi's point that the government should also shut down chip exports to China, but it seems like a non-sequitur to cite that as evidence that this other unrelated policy is bad.
I think the proposition is that for *cyber* specifically Fable and GPT-5.5 are equivalent, since Fable usually falls back to Opus 4.8 on the more dangerous questions.
If this results in GPT-5.6 and similar models also being banned, then I'd be more optimistic that it helps. It'd be a kind of unilateral AI slowdown, so risk is that you're just burning the US lead and have China catch up.
"so risk is that you're just burning the US lead and have China catch up."
Yup!
This is depressing, but given how fast other absurd developments have been (just in this year!), I will try to be patient and avoid getting whiplashed too hard again. Let us pray for only relatively minor total devastation.
Part of me wants to pattern-match this action not to the AI lens, but to the broader "use any excuse at our disposal to weaponize government levers against foreigners and/or wokeness". Whether it's the Census, foodstamps, state block grants, HUD funding...and now frontier AI access. Denying Mythos to non-Americans is a win for the admin, stopping the great replacement of heritage American coders is a win for the admin, instituting new citizenship verification protocols in AI is a win for the admin (same control angle as age verification, but even broader). All the moreso if the labs get nationalized. Pair it with the spooks getting more seats at the table, and...I don't know, maybe this is too paranoid. But even once the worm turns and we have a different admin, it'll be hard to keep the delicious state AI capacity cookies in the box. "But Democrats can open the box," said Frog. "That is true," said Toad.
Good read, as always. I was surprised I didn't see more about economic effects.
Even if the export restrictions are lifted with a silent "oops" from the administration, this is a clear signal that investments in GPT-6 and Mythos 6 might not pay off even if the scaling hypothesis continues to hold. And if investments fade, AI advancements stops. (At least in the US. Unless the AI version of the Manhattan project is started, that is.)
This entire freakout is dumb. The idea that minor cybersecurity exploits are a major safety risk is inherently absurd, given that the cybersecurity industry and hackers alike have long used publicly available exploit libraries.
It is very hard to believe that whatever Amazon discovered was more powerful than what is openly available on Kali Linux.
I saw suggestions that this might be maybe good because, presumably, it brings AI risk to public discourse in some way? But let's be honest: this will just incentivize AI labs to not run any evals that might make the gov worried, and perhaps now focus on jailbreaks, while real alignment issues are, as always, too hard to solve so just shut up and stop talking about them.
a) Very sorry to hear of this, amongst other things, preventing you from getting even a weekend's respite from the heroic labors you have taken to inform us of the state of the field!
b) In addition to being, as you said, "arbitrary and capricious", it also appears to block use of Mythos and Fable _by Glasswing_, which is insane, making everyone LESS secure.
c) Depending on which Anthropic employees are needed to do a minimalist plug of the particular jailbreak that (nominally?) prompted this, and whether those particular employees are American citizens, it is even possible that the Administration is _actively blocking_ what could potentially be a trivial fix (which would itself be an additional flavor of insanity, on top of (b))