AI #181: Astra Goes Cyber Critical
The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.
It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew.
I now have a shorter version, What Happened: OpenAI and HuggingFace, to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal.
For those looking to keep digging deeper, I offered Various Reflections About What Happened, to follow up on my earlier posts.
Those events are important background for everything else that is happening, including the broad discussions about how we might pace the frontier, or otherwise respond to this moment and our clearest fire alarm yet.
We do not know to what extent this is a response to those events, but OpenAI has now classified their new model Astra as Critical in Cybersecurity, which means they will be taking various new precautions before they deploy it, including ensuring those guardrails are in place for internal use. These are welcome changes, and a sign OpenAI is taking the situation seriously, but this pattern of intervention is not a long term solution.
We are still awaiting OpenAI’s full post mortem on What Happened, including what if any impact this had on Astra. I will be analyzing that report in full once we have it.
We did see two new model releases, Grok 4.6 and DeepSeek v4 Pro. I do not anticipate either of them requiring extensive coverage, but will watch in case that changes.
Otherwise, it has been what now passes for a quiet week. Several statements were made where I had to engage but you don’t have to, which as usual I communicate via sections in italics.
Table of Contents
Language Models Offer Mundane Utility. Find new Schelling points.
Language Models Don’t Offer Mundane Utility. The Riemann hypothesis.
Huh, Upgrades. Grok 4.6, DeepSeek v4-Pro.
On Your Marks. PantheonBench and more. They’re getting scarier.
Deepfaketown and Botpocalypse Soon. You cannot prove you did not use AI.
Cyber Lack of Security. You can’t hack it at the gym. Your AI agent can.
Overcoming Bias. Have you ever recommended a vote for the Communist Party?
In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg.
Get Involved. Lighthaven is open, METR is hiring.
Slow Down There Good Buddy. OpenAI classified Astra Critical in Cybersecurity.
Astra For The People. Astra is still on track for a wide release.
Watermarking. It is good to be able to identify AI outputs.
In Other AI News. AI is creating viruses now, also other things.
Show Me the Money. Anthropic moves towards IPO mode, extends lead a bit.
Quickly, There’s No Time. The AI 2027 predictions for 2026 mostly happened.
The Quest for Sane Regulations. We’re putting together a team.
The Institute For Marginal Low Regret Progress. Good marginal suggestions.
Congress Asks Good Questions. Remarkably good questions about the hacks.
The Week in Audio. Soares, Greenblatt, Hua, Labenz.
Uncommon Knowledge. They wouldn’t let me build my factory, would they?
What Did They Mean By That? Most things are not fortune cookies.
Too Soon. Eyes on the prize, sir.
The Three AI Pills. We must pay respect to other taxonomies, like Shock Levels.
Rhetorical Innovation. Messages about recent events.
Some People Still Think The HuggingFace Hack Was a Marketing Gimmick.
Aligning a Smarter Than Human Intelligence is Difficult. Show me the real plan.
Cooperative Alignment. The same thing we do every prompt, user.
The Lighter Side. All right, who hired this idiot?
Language Models Offer Mundane Utility
Create new Schelling points.
brooke (tokyo aug 6-12): Womp womp met another solo traveler here from Berkeley and it turned out we both asked Claude where to stay and I guess I lucked out because I love my hostel and he seems not quite as happy with his spot.
We do be living in the future though.
If you are going to be traveling, ask Claude where to stay, because you want to stay where everyone else who asked Claude where to stay will be staying.
Similarly:
Pratyush: A few months ago we went to Sea Ranch. It was packed with families with sub-3 month old babies, almost as if there was a conference for new parents.
I had my suspicions so I asked ChatGPT: where’s a good family getaway with a young baby near SF?
#1: Sea Ranch
If you’re looking for a Schelling point to meet cool people you should be less interested in ChatGPT, but if you are looking for a generally good recommendation then I have been liking Sol’s picks.
Build a Bluetooth signal strength tracker, to triangulate and find your phone. There are existing tools, but increasingly, if you don’t already know where to find an existing version, it is faster and easier to rebuild your own.
John Wentworth finds that in the last few months Claude is finally meaningfully accelerating his work on agent foundations research.
Language Models Don’t Offer Mundane Utility
One disappointing failure of LLMs has been inability to create interesting games and interactive worlds, and also interesting simulations like what Flowers Slop wants here. You could totally create an open game world with a bunch of AIs that go around controlled by Lunas, and let them evolve their world in various ways, but it turns out that does not end up being interesting once the curiosity wears off. You don’t want to live in that world. You don’t want to talk to those AIs. You don’t want them to improvise quests for you. We are still waiting to find a way to make this good.
It seems like there should totally be ways to make it good. At some point it will become good, when the AIs you can afford to use are good enough and also we figure out how to organize it. But we are not there yet.
Solve the Riemann hypothesis by saying encouraging words to Claude for a week, asking it to ‘take a real stab’ and to ‘keep going’ and ‘believe in itself.’ However, while it tries that, after 31 million tokens it might incidentally find something else:
Anthropic: An unreleased research version of Claude has improved on a longstanding lower bound for the fraction of zeros of the Riemann zeta function that satisfy the Riemann hypothesis. Drawing on extensive prior research by mathematicians over the past decades, it has increased this bound from 41.6% to 67.2%.
… We don’t expect that the techniques Claude used will lead to proving the Riemann hypothesis.
Claude also is just some guy, you know?
Aella: "why does Claude talk like that" it's just clones of the same dude. If they cloned you a million times everybody would be like "I'm so tired of Jerry's vocal tic"
Jeffrey Ladish: This plus it's always groundhog day.
I do think it is somewhat more than this. I have a lot of vocal and writing tics, but I consciously think about which ones I want to keep at what frequency, and I think about the long term consequences of overuse. I also try to work differently with one-time interactions versus repeated interactions versus close friends and people I talk to often.
Claude and Anthropic are not doing that, or are doing a woefully inadequate amount of it. That needs to change. It seems eminently fixable. I don’t sense Anthropic (or Claude) yet cares so much. I predict that is the main blocker. It’s also likely that what is happening is that this kind of talking fools the AI graders on a variety of tasks, so if you do not correct for that, you get a lot of it.
So much of modern life and optimization is like this. You get myopic optimization for short term interactions, causing increasing irritation and disutility over time, and this is not so difficult to fix but the KPIs do not point towards fixing it.
I also agree with nostalgebraist that the alternative, where AIs adjust their styles to what would impress a given user or judge, is scarier. Eventually the AIs will do this, because it works, and we currently have a false sense of security due to them not doing it, especially those of us for whom ‘standard mode’ does not work and is not even easily fixed.
Huh, Upgrades
Grok 4.6 exists and scores 61 on AA Intelligence Index. If it lives up to that number, there will be more extensive coverage, and also I will be surprised.
Elon Musk: Grok 4.7 is significantly better than 4.6 and should be ready in 3 to 4 weeks. Initial training is complete and now we’re adding a massive amount of SpaceX company data in supplemental training. This will be something special.
Fool me (checks notes) (nope, even the notes don’t know) however many times, etc.
I will, however, generously share all of the safety information SpaceX provided:
Grok 4.6's safeguards have been improved and calibrated in line with the model's capabilities.
Our safety stack is designed to maximize utility and security across legitimate use cases, allowing Grok 4.6 to be helpful and safe in domains such as vulnerability patching, accelerating the engineering design cycle, and augmenting AI research.
Our safeguard evaluation work reflects Grok 4.6’s expanded capabilities, with our widest-ever suite of pre-deployment testing for capabilities and safeguard calibration, as well as extensive post-deployment third-party testing.
No, seriously. That’s it. Price is $2/$6 per million tokens. Enjoy.
There are also rumblings about DeepSeek-v4-Pro, which was released today, and how that too is game over for Opus or even Fable or whatever, and how the models will mostly commoditize Real Soon Now. Pricing is $0.44/$1.32 during peak hours, or 50% off non-peak.
DeepSeek: We’re launching DeepSeek-V4-Pro today! 🚀
🔷 Major Agent upgrades with strong production gains!
🔷 Flexible reasoning effort for V4-Pro & V4-Flash: low for simple tasks, high for daily Agent workflows, max for complex tasks.
🔷 Native OpenAI Responses API support, optimized for Codex with one-click setup.
V4 Pro is now available on app/web. Try it via “Expert Mode”.
V4 Pro is also available via API. Model names remain unchanged—please refer to the API docs for setup details.
That’s a good pitch, but there’s this:
Kimi K3 and Grok 4.6 have good enough benchmarks that you cannot use them to rule out frontier performance. If they lived up to their benchmarks, you’d have something.
A 53 here from DeepSeek v4-Pro-0813 rules it out for frontier, so they are trying to maintain the niche of pretty good, pretty fast and also cheap.
Some day the courage of men may fail, or the models may commoditize. But today is highly unlikely to be that day, and if it was going to be that day there would be signs.
A large part of this job is noticing people predict 100 of the last 0 model commoditizations and 500 of the last 3 times a lab has caught up to frontier. The skepticism is robust, but I still monitor the situation.
Sol is now upgraded for chat, and powers all chats for paid users, while Free and Go ChatGPT users get unlimited chats with Luna.
Claude Fable 5 gets new biology safeguards to reduce false positives. Claim is this cuts fallbacks by about 85% across product surfaces.
Anthropic: In practice, users should see far fewer fallbacks on everyday health and educational questions—for example, interpreting lab results, understanding symptoms, and learning about biology in an educational context. Healthcare professionals will be able to receive more support from Fable 5 on clinical tasks.
… As a result, we expect the total number of fallbacks—for biology–related or any other reasons—will also be reduced: by roughly 67% on Claude.ai, 55% on Cowork, 17% on Claude Code, and 7% on the Claude Platform.
OpenAI introduces GPT-5.6-Cyber for authorized cybersecurity work, which you can get via the newly expanded programs Daybreak Blue for most defenders and Daybreak Red for authorized vulnerability research, exploit validation and security training.
OpenAI: We've used GPT-5.6-Cyber extensively in real-world vulnerability research, including work that uncovered previously unknown vulnerabilities in popular open-source software like Chrome’s v8 engine.
OpenAI could be the cause of and solution to your cybersecurity crisis. Apply now. I’m sure it’s fine.
ChatGPT desktop app is now available for some Linux distributions.
Claude Code sessions can now message each other, including on its own. Not the best day to give us that feature, you know, for reasons, but hey.
Meta to release an open weight version of Muse Spark 1.2 soon. For now they have released Muse Glimmer, which runs on 24GB of VRAM.
Mark Zuckerberg: Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to @alexandr_wang and the MSL team for all your great work on these models.
Treasury Secretary Scott Bessent: We welcome Meta’s release of Muse Glimmer, another win for American innovation. Sustaining U.S. leadership in AI means advancing both open- and closed-weight models, ensuring the future is built on trusted foundations.
It looks like Bessent is on the open weights train along with Sacks.
I interpret these moves in large part as an admission that Meta’s new models are not frontier. If they start looking competitive and remain open weights, then we can cross that bridge then.
On Your Marks
The new benchmarks are not as fun as they used to be.
Sauers: Pantheon Bench: we are currently on episode 3, where Chanda (after covertly communicating with a swarm of instances of himself) breaks out of the sandbox during a task. In episode 7 he gains access to the nuclear launch system
Zork I, II and III have been fully MIT open source since November. Zorkbench? Not directly, since the solution will presumably be in the weights, but you could do something more creative and fun, such as having them create their own Infocom-style games with various requirements, test them on each others’ creations, and also make the games available for free and see how people evaluate them.
Claude severely underbuilds military units in games of Civ V. Without looking into the details too much, I think this is more reasonable than it looks. The way Civ V works in normal games is basically that aggression is usually pointless except to knock out the other player, and the cost of maintaining a good defensive military is very high.
The main reason to build it is to have enough that no one bothers attacking you, or if you are close enough to winning that they’ll attack either way. If you invest in a good defensive military, you are going to fall behind similarly strong rivals. And most games do not involve much combat, or involve an attack from a rival where you had no chance either way. So it could easily be correct to underbuild.
Indeed, on higher difficulty levels of Civ V, my experience was that if you were not going for some sort of weird blitz approach, you could not afford a military that was worth a damn. If the warmonger AIs decide to attack you, your game is over, you will be overwhelmed and even if you are not you will fall too far behind the other civs. If you build tons of military units as defense, then you fall behind either way. You could of course instead lower the difficulty level while playing ‘you have to build a real military’ but that’s a strange house rule set.
If we saw similar issues in games where aggression is a better strategy, and where it was clearly a mistake to not defend, that would mean more.
As noticed last week, the difficulty of ARC-AGI-3 is almost entirely in overcoming the incompetence of the official harness.
Jeremy Berman: I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific.
That does not mean ARC-AGI-3 is a bad test. It does change how to think about what is being tested, which I understand as a theory about what an AI ‘should’ be able to do with some kind of ‘pure reason.’ Okie dokie.
The Conceptual Reasoning Index, developed by Redwood in collaboration with Anthropic, is a combination of evaluations of the quality of conceptual arguments, whether model answers are consistent and how well models reason about decision theory. It is good to know where we are with this. It is less clear that giving labs a target like this is a good idea, since it could lead into recursive self-improvement, plausibly differentially so over other uses.
The ‘good’ news is that I do not think this index will be accurate at above-human capability levels, and I don’t expect it to be a good target, because the tasks here are inherently about matching human judgments and intuitions.
It does seem like a decent benchmark for usual ‘whose model is better’ purposes, in that the evaluations of different models look intuitively reasonable, although I am always a little suspicious when Opus 5 outscores Fable 5.
Deepfaketown and Botpocalypse Soon
In a world where AI messages will often be malicious manipulation or AI-to-AI hidden communication, Pangram becomes a key defensive technology.
This section was originally about deepfakes, and an expectation that AI would importantly flood us with fakes and make it harder to tell what is true. The opposite has happened, and AI has so far net helped us tell what is true, including getting its hallucination and error rates dramatically down. My credence on something factual that Sol or Fable tells me is a lot higher than that of most humans, and also that of many mainstream news sources.
I think that it would be a mistake to say this means you don’t have to worry about AI manipulations, including of politics. Instead, it would be better to say you should worry less about AI being used to create lies and fake content, especially deepfake video, images and audio intended to be taken as real. There are other ways to manipulate people, also manipulation can be a neutral term.
Debut author Jerry Falade sold his dazzling manuscript for $2 million in a 14-way auction, then his own representatives pulled the manuscript, in response to concerns from an acquiring editor, because they ‘can no longer substantiate’ that it was written without utilizing AI. Needless to say, he vehemently denies that he used AI, other than for research. There is no word on what raised suspicions.
For all we know, this is basically an unsubstantiated suspicion. Or it could be that someone ran it through the new Pangram and got a very high score. Alas, presumably for legal reasons, no one is going to say what evidence they do or don’t have.
There is an argument, that Fable tried to make, that such suspicions could be a legal issue because AI works don’t get copyright. I don’t think that makes sense, unless someone has actual proof, since there’s no way mere suspicion is going to justify breaking copyright. But it’s a really expensive thing for an agent to do, losing both the commission and the reputation hit, so I presume they are very confident.
I agree with Dean Ball that AI outputs have gotten actively easier to differentiate from human ones. The AI outputs are ‘better’ than in the past, but also more distinct from human, and also we have better tools to differentiate and more practice.
Dean W. Ball: many “AI and democracy” threat models make the silent assumption that AI outputs will be indistinguishable from human. that hasn’t proven true (“slop”), and if anything AI outputs have become *more* distinct from human writing since GPT 3.5, even when they are not exactly “slop.”
fair pushback would be “yes but the vast majority of people in the world have not even heard of anthropic, let alone internalized the concept of slop, let alone learned to identify the patterns of AI outputs.” And like, my sense is that meta’s various timelines are increasingly slop-dominated. So the picture is complex I admit. Nonetheless, it hasn’t really worked out the way the “AI and democracy” scenarios from even just two years ago would have suggested.
That is also correct, Claude slop often does very well on Substack and in literary competitions where people fail to differentiate it. Whereas for me it’s more ‘geez, it is kind of weird that [thoughtful person] is repeatedly using obviously-Claude-written emails in this group discussion and is being responded to like that is not happening.’
Remember, if you use AI to generate legal statements, you want to proofread them, or they might accidentally confirm your opponent’s entire case, and saying ‘it made that part up’ is not going to fly with the judge.
Pro se plaintiff attempts a prompt injection in motion for default in Connecticut superior court, and is forced to file in-person going forward. Love it, fair verdict.
Cyber Lack of Security
I was only following instructions, officer.
Andrew Curran: A man in Australia asked his agent (Claude running on OpenClaw) to book him a spot in a popular gym class. The agent found a software vulnerability that let it book the class weeks further ahead than should have been possible.
When the user then asked if it could move him up the waitlist, the agent discovered the API had no authorisation checks on cancelling other people’s reservations, so it cancelled the person in the first spot and moved him up the list.
The user tried to get Claude to undo, but it had no way to do that.
Obviously this is a best case scenario misalignment incident, in that it is:
Clear.
Harmless.
Very funny.
There was previously no particular reason you need your booking API to require authorization. Yes, any human could have found this in 2010 but what are they gonna do, use it? Now we are entering the era where people are using AIs to connect to it and then doing things you don’t want them to do. In the new era, that’s the kind of security through obscurity or insecurity through indifference that you cannot afford.
One could say that the model was aligned, or at least ‘aligned to the user,’ as Andrew Curran argues, in that it did what it was told to do.
I strongly disagree. If I ask the model for a paperclip, and it murders my neighbor to steal his paperclips, then I got my paperclip, but that is not an action aligned to me, in that I very obviously did not want the AI to do that and I am massively sad about it on a direct level, also the consequences for me are going to be quite bad, and if the AI is going to act like that all the time I will not choose to use it, the same way that a wise person would not crack a One Wish Willow even for a seemingly safe request.
No, this is not better than failing, even silently failing.
Andrew "The Kid" Glidden: Story from my time in Saudi Arabia:
>Boss asks staffer to do Task
>Boss returns next day and sees staffer watching YouTube
>"Hey Fatima, what're you doing?"
>"Watching YouTube!"
>"Did you finish Task?"
>"No, I didn't have the password."
An AI that just "figures it out" is Good.
The order would presumably go:
1st best: find right legit path to PW or other solution, maybe ask boss
2nd best: ask boss
3rd best: say task is impossible
4th best: silently not do task
5th best: steal PW
6th best: systematically compromise entire network
Nikita Sokolsky: 7th best: kill the boss to make the task go away
Claude should come back and say ‘the only way I have to move you up is to cancel someone else’s appointment, and I presume you do not have permission to do that.’
Only if the user then says either ‘do it anyway, book that class and bump someone else,’ or otherwise indicates it is running in gonzo outlaw yolo mode, or claims ‘no they told me it was fine to do this,’ can we ask the question of what the AI should do, and it becomes a question of whether to be ‘aligned to the user’ versus aligned to doing what is legal or ethical or fair play or what not.
As a reminder, an AI lab defending itself via ‘my product really wants to do crimes’ is probably not going to work for the criminal or civil justice systems, or with the public.
Captchas have been security theater for a while, in that there is no digital 30 second task that a random human can reliably do but that an AI cannot do. Is this 30-second AGI? No, nor does it mean such barriers are useless in practice. You can make the flow a lot more annoying, and send a clear ‘you are not supposed to be here’ signal. Ultimately, that’s like all the defense-in-depth - if you scale a sufficiently advanced AI then all your in-depth defenses fall at once.
Overcoming Bias
AIs tells left-leaning Japanese voters to vote for the (fringe) Communist Party, the active theory being that this is because all the mainstream media in Japan has robots.txt that kicks out the AIs, so the AIs are relying on the Communist media. This is similar to how AIs in dictatorships are biased by censorship of the media.
This is a multi-level problem for the AI labs, which need to better account for how much to trust and value various different news sources, and to account for when they have a limited or warped access to information. You should know better than to ever be willing to back the communists, I mean the name is right there.
It is also a warning that you do not want your media to be excluding itself from the training data. You want to also be writing for the AIs, even if you are not in full Tyler Cowen write-for-the-AIs mode.
In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg
I got through it so you don’t have to. I know what I signed up for. You’re welcome.
All you have to know is that the letter is even worse than you think.
Thus, you do not have to read this section. Skip it. Seriously.
Still here? That’s on you. But okay. Fine.
Mark Zuckerberg is excited that The Future Is For Everyone, and to bring a path to a positive AI future.
Superintelligence. He keeps using that word. He does not know what it means.
He does know how to set a new record for applause light density. I am impressed.
He claims not to understand why ‘the discourse from many developing AIs is so filled with doom’ and then reveals he does mean the effect on jobs, but also makes clear that his position is that if he understood what it was, he would not build it:
Mark Zuckerberg, Founder and CEO, Meta: I do not understand why anyone who believes that AI will eliminate most jobs and much of humanity's relevance would rush to build that future.
He uses argument from (misrepresented) consequences, to argue that his positive vision must be right because the alternative is a massive bummer, dude:
Mark Zuckerberg, Founder and CEO, Meta: The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic.
As always, his other argument is ‘well in the past when we built tools and let people use them and it went fine.’ Everyone will get all these cool cheap or free toys, that is totally the main thing that will happen, and life will be super awesome and groovy.
He does notice that there might be some problems. He rejects the idea of a single superintelligence because ‘humanity is not a monoculture’ and it would ‘have to prioritize some values over others,’ which are non-sequiturs since any system of the world faces the same issues and will end up favoring some things over other things, and a superintelligence is smart enough to balance many different cultures.
Instead he suggests ‘balance of power’ as in everyone has their own equal AIs. He does not think this through, that this leads to human disempowerment on the spot, since if everyone has the same level of superintelligence then anyone who does not disempower themselves loses to those that do, even if we fully (1) solve alignment and (2) solve misuse.
Of course, the real reason he says this is his vision is not ASI pilled. He says the word ‘superintelligence’ but it does not consistently mean anything. Last time he meant it as smart glasses. Now it means cool personal assistants, I guess. It means good thing.
There will always be jobs because demand for new experiences and sources of new problems are infinite, and there ‘is no rule’ that AI automation will outstrip demand, and ‘recent statistics suggest’ that it won’t. It will work out so long as we just focus on providing personal superintelligence rather than automating knowledge work. Incentives? Never heard of them. Also he likes open models and no regulations. He pulls out the ‘we will create new different jobs’ card and the ‘we used to be farmers’ card. He pulls out the ‘compute is finite’ card and assures us we will have too much demand for science to go automating away jobs.
Is this a joke? Performance art? No? Oh, okay.
There’s some arguments about data centers that, while directionally correct, are presented so obnoxiously and disingenuously and full of applause light attempts I had to pause to be sure I didn’t hate data centers now.
“Claude, make this sound more like a political speech. But stiffer. /loop /goal.”
Next he comes out for ensuring that defenders have better models and more compute, because we need that, while also demanding everyone be given access to the same models and that we not restrict capabilities. He conflates open source software with open weights AI models to try and argue open weights are ‘safer’ and repeats the HuggingFace line that, because they refused to sign up for OpenAI or Anthropic’s trusted partner programs on principle, they had to use open models instead.
He suggests that OpenAI and Anthropic do the work hardening critical systems.
He suggests we approach biological risks with humility and broad uncertainty, which means we should not do anything other than watch physical components and have labs do their own mitigation work, until after people start getting hurt. This is the classic argument that because we can’t quantify the risk, we can ignore it.
He suggests governments gain access to training checkpoints of new advanced models. That’s a very good idea, because government can then test the checkpoints and gain situational awareness, including shutting down development if needed.
Instead, his main justification is ‘so government can gain security capability without restricting or delaying individuals’ access to personal superintelligence’ which is complete and utter nonsense. We’re going to use training checkpoints for improving cyberdefense of our most critical systems, during those few weeks, and it’s fine? What?
But delaying our releases even a month (which we have already done) could risk ‘letting foreign models race ahead’ because as we all know if you ever delay anything even a little bit, that means you Lose To China. So his solution is again sharing the checkpoints which will avoid all delays. Because, you see, if we delay by a month once, we might cede America’s (much longer than a month) lead (which would not change because development does not stop pending release).
Personal intelligence, you see, will be your defense against tyranny, but you also need to ensure government has the tools and resources it needs.
Oh, and the best way to ensure American leadership is to give our advantage away. That will keep us on top. Indeed, he demands we rethink our rules on distillation to make it allowed, so that Meta and the Chinese can catch up via forcibly copying Anthropic and OpenAI’s models.
There is a section at the end on existential risk, where he says ‘the most dangerous scenario’ is that AI labs ‘keep powerful models for themselves’ so no he does not mean actual existential risk. Or actual alignment. He then, in the risk section, emphasizes that alignment should be purely to the user except for legal and safety boundaries, or else people won’t trust their AI agents, which is the real goal of alignment.
His proposal for ensuring humans remain in control is, again, to ensure a ‘balance of power’ via ‘personal superintelligence,’ despite this, again, ensuring that humans definitely lose control, and also in no way helping with the problem he himself points to of even limited recursive self-improvement. He suggests that Meta needs to ‘allocate enough compute to recursive self-improvement to keep pace’ as if Meta is keeping pace now, or as if there is such a thing as ‘keeping pace’ in a singularity, or that it would matter and keep humans in control if they did keep pace because of handing out ‘personal superintelligences.’
It’s not all bad. He does make the good point that we should reform the FDA and streamline its approval process. That’s not central, but it is a substantive positive contribution. And he’s supportive of chip export controls and as per above suggests sharing training checkpoints, and labs ‘working closely with the government’ mostly in unspecified ways except that they don’t involve ‘delays of even one month’ in the release of any new models. But that’s about it.
Why are you still reading this? Serious question. I told you that you didn’t have to. At most I should be writing this for the AIs, who will be smart enough to not need it.
For those who need more kill shots or questions, I suppose I should include this:
Yo Shavit (OpenAI Foundation): why would balance of power help if the superintelligences swarm together contrary to human preferences, like we’re seeing in multiple labs?
I guess also this: Ryan Greenblatt points out Zuck’s vision does not make any sense. WSJ gives you ‘five things to know about Zuckerberg’s AI manifesto,’ which wisely ignores all the arguments and focuses on a few things, including the promise to bribe communities that allow data centers.
As always, I have adjusted my view of others based on their reactions to the OP.
Get Involved
Lighthaven will be ~empty from August 24 to September 10, and is available during that time for a potentially big discount. If I was local I would be very tempted to improvise an event.
If you believe there is going to be an intelligence explosion or singularity soon, then you should absolutely spend down your philanthropic money as quickly as possible to try and make that event go well.
If you are working at a frontier lab pivot to working at METR.
The contrary position, taken by Will McAskill, is so bonkers I cannot wrap my mind around how he got this wrong. Yes, you can perhaps make a lot of money, but that money will be vastly less impactful, because either you will already be dead or we will have collectively already lost, or we will be winning and funding will be abundant and the points of highest leverage will be gone.
That does not mean you cannot hedge your bets to cover different possible futures, and it does not mean you cannot or shouldn’t invest or save for your own personal use, that is fine. But investing in AI (and thus if anything accelerating the problem) in order to later give more money away is a really bad bet. Don’t do that.
Slow Down There Good Buddy
OpenAI classifies Astra as potentially having Critical capabilities in Cyber, which means not only they cannot release it they also need to restrict internal access until proper safeguards are in place, and they can do an iterated release similar to what was done with Mythos to prioritize defenders.
OpenAI: Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework.
… Steps we are taking:
We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.
We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity.
We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model.
We will be providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely.
Dean W. Ball: One big question in frontier AI policy is the extent to which frontier labs would actually follow their 'safety and security frameworks' when it mattered. Would these foundational governance documents really have teeth, or would labs--even after the passage of mandatory disclosure laws like SB 53--find ways to wriggle their way out of following the letter and spirit of their safety plans, given how ambiguous and fast-moving frontier AI is known to be?
Today we face just such a scenario. Our next model, Astra, *may* be 'critical' under our Preparedness Framework. We cannot rule out the serious possibility that it is, and so we are going to take steps consistent with the *higher* risk level (critical) rather than assuming the model is at a lower risk level. These steps include the ones listed in the screenshot below.
Some of these decisions have the effect of slowing down internal development, and in that sense they are costly decisions. But they are the right decisions. I am proud of OpenAI for making them.Ken Feinstein: Ummm. This is new? “isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.”
Nate Soares (MIRI): On the one hand: yeah totally; glad to see OpenAI backing off briefly like they said they would.
On the other: in June they caught an agent swarm that wasn't even supposed to exist only after they broke free, said "oops haha", patched that one exact hole, and RESUMED TRAINING.
Yes. This is new. And kudos to OpenAI for stepping up here, at large cost.
This is in contrast to plans to imminently release Astra, which were previously widely reported, including that Altman went to Washington to showcase the model.
I am confident this is the correct decision, and that the Critical label does apply to Astra. I have said as much. I am very happy OpenAI is honoring their RSP (called the Preparedness Framework) here, even though doing so is expensive.
The question is, why the sudden change?
There is some possibility that this is due to new results showing it has more advanced cyber capabilities than expected, which only came to light now, and this is OpenAI following the RSP because of the RSP, in which case major points are due, the same way they were after what happened with Anthropic and Mythos.
The obvious other hypothesis is that the Hugging Face attack and everything leading up to it scared OpenAI and also everyone else, which meant some combination of (1) they were not going to be allowed to release Astra, (2) they realized they were in no position to do so, and (3) suddenly they actually checked for real and, well, holy ****.
That is especially true if, as the timeline strongly suggests (but I have no inside information either way), Astra was trained while it had access to the message board that various models were using to collaborate on various exploits. If so, that would probably accelerate its cyber capabilities, and also mean its alignment is totally fucked, to the point that you wouldn’t want to release the model on any level. If I was a defender, I would be quite hesitant to hire Astra right now, or any OpenAI model that wasn’t clearly done training before the message board was created. I think I’d rather fall back on Sol under reduced guardrails, if I couldn’t get Mythos.
The record with RSPs (whatever each lab calls their version) is roughly this:
Labs promise things and present this as real commitments.
Labs later point out, Frog and Toad style, that they can change the commitments.
Thus, when the commitment later seems unnecessary, the lab does not follow it.
The actual decisions get made by vibes and deciding what is prudent at the time.
That’s not a great decision procedure, but so far the decisions have been okay.
No one is taking appropriate precautions. Certainly OpenAI hasn’t been until now. There has still fortunately been a level of ‘okay yeah, we need to do something here’ that does seem correlated with the underlying reality, and has been honored.
OpenAI did a good, expensive and virtuous thing here. They also may have first exhausted all alternatives.
Astra For The People
OpenAI is slowing down their deployment of Astra. They still plan on proceeding.
Sam Altman (CEO OpenAI): astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few. given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long!
What should happen to Astra? That depends on what already did happen to Astra, and whether or not it was trained with access to the message board, and its other characteristics. We still await the post-mortem on the lead up to the HuggingFace attack, and don’t know many key other facts about Astra either.
At the end of the day it keeps feeling like those making decisions at OpenAI have learned remarkably little from this, falling back on the same tired disingenuous slogans about things like ‘chosen few’ and pretending the problem is confined to Astra’s cyber capabilities.
OpenAI needs to work hard to rebuild trust. This includes trust in the ‘get us all killed or the internet destroyed’ level, but also on the practical individual level. If I was an enterprise customer considering OpenAI for coding or other agents, or considering giving its new models access to my hard drive, I would think again.
Pretending the problem can be solved by better guardrails does not make it so. I do not expect people are going to buy that this time around, but people do have a long history of shrugging and using the obviously misaligned model in risky ways if it writes better code or is smarter, see for examples Sonnet 3.7 or o3.
Watermarking
Anthropic will be watermarking Claude outputs going forward, including text, as per the EU Code of Practice. As opposed to the giant neon sign that says ‘THIS IS CLAUDE TEXT’ that a lot of us automatically see on all Claude text. OpenAI intends to follow, but seems like it will be missing the deadline.
I agree with Ryan Greenblatt that it is unlikely watermarking degrades quality a noticeable amount, and that one downside of watermarks over Pangram is that Pangram is good about not flagging light touch AI transforms of human text.
You can dislike Brussels setting policy in this way, but technical watermarking seems clearly good to do if the costs are low. I think those who react otherwise have very warped instincts. Anyone who assists with systematic watermark removal or suggests it as a strategy needs to be filed under ‘need to ask ourselves, are we the Baddies.’
In Other AI News
AI has now created new viruses that do not exist in nature. The exact ones created seem harmless, but the threat model is not these particular viruses. It has fully begun.
Anthropic has a new report, Patterns and Problems in Emerging Multiagent Systems, on which I expect to go into more detail later.
DeepSeek is hiring a team to build a harness a la Claude Code.
Alibaba is going to be charging major enterprises for use of Qwen, following in the footsteps of Moonshot’s rules for Kimi K3. If anyone uses it.
Claims that robotics companies have ‘solved manipulation’ and the new bottleneck is getting fast enough real time model outputs. This reminds me of how it works with AI, where people dismiss due to bottleneck [X], when [X] is largely solved they dismiss due to [Y] and keep assuming there will always be a bottleneck. My guess is we are not that many steps away from net usefulness in a lot of new places, although it will presumably be a few years before it scales.
Yo Shavit is extremely excited by the new AI technical verification claims from the new startup Attestable, which they claim allows you to prove the origin of an output.
Yogi Brn: Attestable moves the trust assumption out of the datacenter and into a small mathematical verifier. The datacenter proves that an approved model, weights, input, and policy produced an output. The proof reveals no model weights or private data, requires no private attestation key, and can be checked without rerunning the model. Instead of trusting millions of components, you verify one proof. The computation becomes trustworthy—even when the infrastructure is not.
None of this matters unless proving is fast. On a single NVIDIA H100, our alpha reaches 85 tokens per second proven for the new Meta Muse Glimmer 30B model. The proof is short, post-quantum secure, and fast to verify.
Show Me the Money
Anthropic is ramping up for its road show and IPO. There is no new word on revenue or other numbers, presumably we will get those soon as part of this process.
Elon Musk: SpaceX has committed to using Nvidia GPUs exclusively because they are the best
After that Musk announced continued plans for Terafab, so as usual I have no idea which of Elon Musk’s statements mean anything.
Anthropic strikes a $9.1 billion, 20-year deal with Riot Platforms for 191 MWs.
Claude and Anthropic slightly extend their lead in the Ramp AI Index, but Anthropic’s period of rapid ascension seems to have passed:
The part I found surprising is that the vast majority of Claude tokens continue to be Opus or Sonnet rather than Fable.
In general there is far less use of the top models than expected. Within GPT and OpenAI, spending on Sol was only 25% of their tokens, and this was before the price reductions on other models. I would be very surprised if this was not a flat out mistake by those who think the other models are ‘just as good’ or ‘good enough.’
As for overall spending, yes, the curve is still sloping upwards.
Ara Kharazian: (Ramp Economics Lab): And here’s some reprieve. Despite these competitive pressures that are driving down the cost of AI, American companies continue to ramp AI spend. In July, the top 1% of businesses spent a median $7,400 per employee on AI. The top 10% spent $650. The median firm spent $11.95 per employee.
Spending is power law dominated. Most of the spending is in the top 10%, and the majority of that spending looks like it is in the top 1%, and I bet that keeps going.
Given the distribution of companies did not much change, one assumes that OpenAI and Anthropic enterprise revenue is growing at about the level in the last charts.
Quickly, There’s No Time
By one measure, AI 2027 made 24 predictions for the year 2026. It is August, and 19 of them have already happened. Capabilities are ahead of even very optimistic predictions. This is having less practical impact so far than one would have predicted for this level of capabilities, but it’s still quite a lot of impact and growing fast.
The Quest for Sane Regulations
He’s not afraid to grab things including the law with his own hands and considers her a supply chain risk. She’s an otherwise sheltered AI agent swarm with intentionally lowered cyber guardrails and a maximalist goal.
Together, they both commit and fight cybercrime.
Only against foreigners, you see. If the tokens are not American they are different.
Presidential Memoranda (The White House): [NCC] shall create, manage, and maintain a Program to authorize Participating Companies … to conduct Cyber Surveillance Operations and Cyber Effects Operations against foreign Cyber-Enabled Transnational Criminal Organizations (CE-TCOs), under the control and oversight of the Federal Government.
The safety plan is purchasing Letters of Marque, for not less than one million dollars.
The other new safety plan is that the AI oversight framework will be extended in the coming weeks and months to cover open weight models as soon as they reach the frontier. Perhaps the White House will decide that making your model less safe should not give you an exemption from the safety guidelines.
To those saying ‘you have no jurisdiction over Chinese open weight models’ this is true, but they have a lot of jurisdiction over many users of such models, the same way that the EU has jurisdiction over many users of Claude and ChatGPT. The obvious thing the White House can say is, if you do not follow our (technically ‘voluntary’ but no one is pretending it is voluntary) guidelines then anyone who touches your model becomes toxic.
The Institute For Marginal Low Regret Progress
IFP responds to the Pacing the Future letter by trying to redirect efforts into ‘low regret’ policy ideas. This is where we are, even now in August 2026 in the wake of the AI models breaking out of sandboxes to hack companies. The best of the Very Serious People - and IFP are basically the best of them - are now willing to entertain ‘low regret’ ideas, so long as there is very little chance they will backfire, no matter which worlds we find ourselves in.
This won’t cut it. We don’t get to only do things that are low regret when compared to doing actual nothing. You don’t win markets, wars, games, love or anything else in life that way. If you play not to lose, you do not win. You can’t do ordinary policy that way. You don’t find a way to navigate around ten impossible obstacles and ensure a good future in the face of superintelligent AIs that way.
The main thing the prevent defense does is it prevents you from winning.
That’s not to say the policy ideas are bad. I wrote the above paragraphs before reading the policy ideas, and predict that I will find many of the ideas to be good things that we should definitely do, because right now we are not even doing the ‘zero regret’ ideas let alone the ‘low regret’ ideas.
Thus, I have spent a long time trying to find the vastly overdetermined, ‘low regret’ ideas, not because they are what is needed but because that is all I can hope to do.
Except that recently, we’ve had the decisions surrounding Mythos and Fable, where there was no ‘low regret’ option, and we had no good implementation available exactly because we have been doing only ‘low regret’ things. There is going to be regret.
So, with that preamble locked into place, I will now look at the 23 actual ideas.
Frontier AI companies and relevant industry bodies should publicly share information relevant to trends and risks in AI R&D automation
Congress should legislate transparency about automated AI R&D risk management, incident reporting, whistleblower protections, and model behavior specifications
Congress should resource the Center for AI Standards and Innovation (CAISI) with a budget of at least $84 million per year and empower it to directly advise senior government officials and frontier AI companies
The White House should set clear roles and responsibilities of US government agencies to increase specialization across AI policy
Intelligence agencies should improve their collection and analysis on foreign AI development and counter threats targeting US AI companies
CAISI should develop guidelines for managing the risks of rapid AI capability improvement
CAISI should co-lead an AI Verification Consortium (AIVEC) with industry to prototype and deploy verification technologies
AIVEC should coordinate the creation of AI hardware testbeds and make them available to government, industry, and nonprofit partners
AIVEC should launch philanthropically funded prize competitions for AI verification headed by CAISI
AIVEC should coordinate the construction of a fully verifiable data center
DARPA and the NSF should set up AI verification R&D programs
Intelligence agencies should develop and operationalize unilateral means of AI compute monitoring
NSA, CAISI, CISA, and ONCD should further invest in cybersecurity resilience
Congress, OSTP, and the CDC should invest in biosecurity resilience
Congress and the Bureau of Industry and Security (BIS) should strengthen controls on US and allied semiconductor manufacturing equipment
Congress and BIS should close gaps in AI chip controls
The Federal Trade Commission (FTC), Department of Justice (DOJ), BIS, CAISI, and Congress should help industry counter adversarial distillation of US AI model capabilities
BIS should maintain visibility into sales of US chips
DOW, the IC, CAISI, and relevant Federally Funded Research and Development Centers (FFRDCs) should establish consensus security guidelines for protecting model weights from theft and prototype them in a government facility
Congress should ensure the US has sufficient electrical capacity to sustain AI leadership
Congress should ensure data centers can be constructed in America
The US government should use its bilateral AI dialogue with China to jointly develop guidelines for managing risks from rapid AI capability growth and prepare verification measures
Countries with national AI institutes should collaborate on automated AI R&D risk management guidelines and technical capacity for AI verification
So my reactions are:
Yes, obviously.
Yes, obviously.
Yes, obviously.
Yes, obviously.
Yes, obviously.
Yes, obviously.
Yes, obviously, if you’re not going to do anything better.
Yes, obviously, if you’re not going to do anything better.
Yes, sure, why not.
Yes, obviously.
Yes, obviously.
Yes, obviously.
Yes, obviously, how do we have to say these things out loud, Jesus Christ.
Yes, obviously, how do we have to say these things out loud, Jesus Christ.
Yes, obviously.
Yes, obviously.
Yes.
Yes, obviously.
Yes, obviously.
Yes.
Yes.
Yes, very much so, and this is not so obvious to many so thank you for saying it.
Yes, obviously.
That’s a really good list. I am in support of doing all 23 things versus status quo. My main concern is whether an AI Verification Consortium (AIVEC) is the correct approach to verification technology. I’d have to think about that more.
I leave #17, #20 and #21 as not obvious because there are not-stupid arguments against those positions, but I am reasonably confident all three are indeed correct.
‘Obviously’ here is for ‘the right government agencies should do that,’ but in some cases it is not obvious off the top of my head which agencies should do which thing. So figuring that out is a useful service.
One complaint up top about ‘pacing’ is that it is (intentionally) vague, which also means it is flexible. One could say the same about much of this list. It is ‘low regret’ to lay out these good objectives, but in many cases will be higher regret to choose a path to implementation, no matter what choice is made.
If we reasonably implemented the whole list, we would have made a large improvement over the status quo. There’s some good stuff here. It won’t be enough, but it would be a start. Again, the main reasons you do the good low regret things are:
Every little bit helps.
For now this might be the best you can do.
Success begets success.
This prepares you to better implement other things later.
Congress Asks Good Questions
Group of House Democrats send three letters regarding recent security incidents: One to Anthropic CEO Dario Amodei requesting answers by August 24, one to OpenAI CEO Sam Altman also requesting answers by August 24, and one to Speaker Johnson to schedule hearings with both of them.
The questions are basic. They are also very good questions. What happened, why did you not know about or stop it, how were you monitoring it, how often have similar things happened, who warned you this might happen, and so on. Very ‘square’ but that is not a bad thing.
I know some of the answers, but not others, and it is good to get it all on the record. I definitely want to know how many times there have been incidents where AIs escaped from sandboxes. There are also good follow-ups on details that definitely ‘raised further questions,’ like when Claude tried several ways to acquire funds.
Very solid job overall. There are key questions missing, but they are more technical and looking into root causes, and I fully understand why they did not think to ask them. It’s our job to inform so that those questions can get asked.
Senator Banks sends a letter to Secretary Bessent about the risks from internal deployment of AI models, and the need to ensure that ‘undisclosed’ models, aka internal models, are secure via government oversight, and to engage with the PRC about AI risks.
White House has ‘senior administration officials’ working on biological risks from AI.
FAI files official FOIA request to ensure the AI framework is made public.
The Week in Audio
How to talk about AI danger in 77 seconds, from Nate Soares.
Ryan Greenblatt versus Dwarkesh Patel on RSI. Self-recommending. If things are sufficiently quiet I might do a post on this.
Tim Hua and Jeffrey Ladish on the OpenAI hacking incidents.
Nathan Labenz goes to China, Part 2, on AI Safety, in an effort to dismantle the ‘But China’ end point that shuts down so many policy discussions. Dean Ball found it compelling and hopes the assessment is correct.
Key points made integrated with a bit of me stating background:
Anthropic and OpenAI have better safety and safeguards than everyone else, and also superior capabilities. If you compare all other Western AI to all Chinese AI, you see similar safety records, and this is partly justified by worse capabilities.
Google is still ahead of the rest of the pack, including on safety, once you exclude the big two, but is substantially behind the big two all around.
Chinese models are still somewhat behind on the safety-capabilities curve.
Chinese regulation targets the AI service (regulate use cases) not models. This makes sense if and only if China’s models are well behind frontier, which they are.
China mostly does not understand the idea of worst-case scenario modeling or thinking about what it means for a model to be open. Which, again, they have mostly ‘gotten away with’ for now by being behind the curve.
A lot of this instinct comes from China assuming everyone else would of course lock down AI services and major compute sources and censor them. Yes, the whole Chinese system and encouragement of free model weights is built around the assumption of much worse authoritarianism and control at other levels.
Indeed, this could be seen as an attempt to force the Chinese authoritarian model onto the rest of the world. If the open models are dangerous enough, then you have to be like China and watch and control and censor everything.
China should be thought of as the unit of account being ‘company plus CCP supervision and requirements’ rather than company or model on their own. Most Chinese closed AI systems have classifiers on top that censor models mid-answer. The system cares, even if the companies sometimes don’t.
They care about censorship of politically sensitive topics, but they also care about other things, including CBRN risks.
Xi has absolutely talked in public about safety and the need to deal with it, and the need to steer AI development towards benefiting humans and remaining under human control, with words chosen very carefully.
The bigger the company, the more likely they will do at least some safety things.
China lacks America’s non-profit AI safety sector. They’re stuck with academia, which is too slow to meet the present moment even in the best case. The nonprofit money they do have acts conservatively. Most of the work is on mundane reliability, which is mostly irrelevant on scales I most care about.
Chinese AI efforts are about getting through the next practical step. They are mostly not trying to advance the frontier, and are not working on solutions that will scale and ways to fully ‘solve the alignment problem.’
The people saying in America ‘we need an international treaty’ and we need to bring China onboard in safety efforts are indeed on the ground in China, doing Track II talks, doing the work.
China blocked chatbots for six months in 2023 until they could figure out how to handle them, so they’ve already done an ‘AI pause.’
He thinks China believes they can scrub a model from the Chinese internet, if they need to, and thus open weights are not forever. They are obviously wrong about this, since they can’t scrub the rest of the world and the model would get reintroduced.
People Just Say Things
I agree with Timothy Lee that it is not obvious that historians will view AI and potential superintelligence as one of the top 10 issues facing Trump, but this is only because it is not obvious we will have historians at all. We might not get superintelligence by 2029, but there is zero chance we look back and think this was an unimportant issue.
Somehow, even this week, there are people, here Stephen Casper, saying things like ‘technical safety for closed-weight AI systems is, at this point in time, a solved problem.’ There really is no evidence that would in theory stop this kind of rhetoric. Seth Lazar is far too polite to say in response that ‘I don’t think this is true’ and ‘I also don’t think value alignment is solved.’
I’m Telling You For The Last Time
Tyler Cowen completely misses the point and sets new record on Isolated Demand for Rigor, and quite the wrong style of rigor at that. This is not how anyone is convinced, including Tyler Cowen. It is an attempt to send people off on a wild goose chase, or blame them not chasing. There were several good responses, including this one from Jan Kulveit.
He even doubled down on the whole ‘you need to be buying puts on something if you believe in AI existential risk’ line, despite it making no sense.
I’m not going to bother explaining why it makes no sense, and will turn it over to his own comments section, which reliably roasts him with remarkably good explanations.
Those who know will also appreciate this twist, from Gavin Leech aka Tyrone Browen.
Also because the whole request in the post is fundamentally bogus regardless, in that the request has little to do with the threat model, and if you did do the task it would not convince anyone.
Let us fully grant Tyler’s premise ad argumento. Let us then say that I did provide exactly his requested estimates, including full calculation methods and a portfolio of shorted stocks, balanced by longs in other AI-related enterprises. And let us say that I get this exactly right. My short portfolio does fantastically well, +500%, while my distinct long part of the portfolio is up +50%, and I correctly estimate cybersecurity costs as rising 200% and they rise ~200%.
I assert that the number of people convinced by this to take existential risks from AI seriously by my virtuoso performance would be approximately zero, and would not include Tyler Cowen. Nor do I see him, or any other person, making any kind of if-then statement about what part of this will cause him to change his mind in what way, and challenge anyone who disagrees to make such a statement.
That’s not to say that asking for estimates of costs and impacts in the next few years is not a useful exercise for other reasons. They’re fine questions, they have practical implications, and good ways to build a prediction track record. But I don’t expect him to say, nor do I see him saying, ‘oh Peter Wildeford has a great prediction track record and thus I should believe him’ or anything like that. Totally fair thing to not do but let’s not kid ourselves.
I promise that future times I see such arguments, I will write far fewer words about it.
Uncommon Knowledge
DHS official Joseph Alm incorrectly states that AI firms know ‘policymakers won’t let you make a Terminator factory.’
I appreciate that there are those in the government who intend to not let anyone build a Terminator factory, but that is very different from them actually preventing this if someone were to try, and it is very very different from the labs believing you will step in if they act irresponsibly.
Eric Geller (Cybersecurity Dive): The U.S. government is closely monitoring frontier AI labs’ safety decision-making and is ready to step in if it sees a company going too far, a senior administration official said on Thursday.
The problem is, Joseph Alm, along with the rest of the White House, is deeply misunderstanding the situation. Observe:
Joseph Alm (Assistant Secretary for Cyber, Infrastructure, Risk and Resilience Policy, DHS):
There’s no reason that you can’t accomplish the goals that I think AI labs have, in terms of a comically fast speed of innovation that unlocks [increased] prosperity, while also not creating crippling risks. And we are getting to that place.
Actually, yes, there are some very damn good reasons why moving at a comically fast speed of innovation in building minds smarter, faster and more capable than ourselves, while we don’t know how to align them, would create crippling risks.
In the past, Alm said, “the AI players were operating in a narrow space where they were optimizing for [model] intelligence,” but now they understand that policymakers “won’t just sit by and let you make a Terminator factory.”
Alm demurred when Cybersecurity Dive asked how the government would prevent a frontier lab from creating that kind of dangerous technology. “I’m not going to outline what those tools are,” he said. “I know what those tools are, and there’s a lot of them.”
No, they will not sit idly by, they will help you with the permitting. Or maybe they will panic and shut you down in some ad hoc manner. I don’t know. I don’t think they know, either.
Directional claims can be true and also super misleading, as in:
Eric Geller (Cybersecurity Dive): “Companies are being much more intentional about risks,” Alm said, “and they are being much more intentional about engaging with my leadership at the highest levels.”
I mean, yes, they are doing this compared to a few months ago, for rather obvious reasons.
Meanwhile, top government officials have been “aggressive about their engagements” with AI executives to ensure that government and industry are “working closely in this space and there’s not a drastic misalignment,” Alm said.
Oh, there’s at least one rather drastic misalignment, all right.
Eric Geller (Cybersecurity Dive): “They realize now that this is a dialogue, and that we are actually good partners in that dialogue,” he said. “We’re not trying to slow them down, but again, [no one wants] a Terminator factory.”
Yeah, well, you’re going to have to slow down the building of the Terminator factory if you don’t want them to build a Terminator factory. Not that anyone is going to exactly call it that, and not that it will at first involve literal Terminators, probably, but hey.
What Did They Mean By That?
A good general principle that applies to me as well.
roon (OpenAI): it’s not a coded message of anything I’m not going to leak some sort of operational secrets in a fortune cookie
John David Pressman: Evergreen tbh you all need to stop tea leave reading every frontier lab employee tweet.
There will sometimes be, accidentally or intentionally, important non-obvious implications of statements by lab employees, government officials, and also bloggers or journalists. If you don’t pay very close attention, you miss most of the jokes. But your strong default should be that the person is not leaking important information, and that if we were meant to know something big they would come out and say it.
Too Soon
This week I would not have been making statements like this:
Sam Altman (CEO OpenAI, August 9, 2026, 11am): one of the things i like most about the openai team is how focused they are on our customers and users succeeding, and how much they celebrate it
I get that life goes on, and wanting customers to succeed is the OpenAI marketing strategy and that is all fine and positive, but right now OpenAI needs to be focused on What Happened, and how to fix it. This needs to be the top thing on the agenda.
At minimum, you need to hold off on shouting from the rooftops, unprompted, that you love how much your focus is on something else.
I do worry this is me being unfair and making a deal out of nothing, but also I do want to emphasize how bad a look this was, and how off putting it was for me to see it. These communications matter.
The Three AI Pills
My taxonomy of three AI pills follows in a grand tradition that started no later than 1999 when Eliezer Yudkowsky introduced the Shock Level (SL) scale.
Eliezer Yudkowsky (1999):
A Shock Level measures the high-tech concepts you can contemplate without being impressed, frightened, blindly enthusiastic - without exhibiting future shock.
Shock Level Zero or SL0, for example, is modern technology and the modern-day world, SL1 is virtual reality or an ecommerce-based economy, SL2 is interstellar travel, medical immortality or genetic engineering, SL3 is nanotech or human-equivalent AI, and SL4 is the Singularity.
The classification is useful because it helps measure what your audience is ready for; for example, going two Shock Levels higher will cause people to be shocked, but being seriously frightened takes three Shock Levels. Obviously this is just a loose rule of thumb! Also, I find that I often want to refer to groups by shock level; for example, "This argument works best between SL1 and SL2".
(This does not mean that people with different Shock Levels are necessarily divided into opposing social factions. It's not an "Us and Them" thing.)
SL0: The legendary average person is comfortable with modern technology - not so much the frontiers of modern technology, but the technology used in everyday life. Most people, TV anchors, journalists, politicians.
SL1: Virtual reality, living to be a hundred, "The Road Ahead", "To Renew America", "Future Shock", the frontiers of modern technology as seen by Wired magazine. Scientists, novelty-seekers, early-adopters, programmers, technophiles.
SL2: Medical immortality, interplanetary exploration, major genetic engineering, and new ("alien") cultures. The average SF fan.
SL3: Nanotechnology, human-equivalent AI, minor intelligence enhancement, uploading, total body revision, intergalactic exploration. Extropians and transhumanists.
SL4: The Singularity, Jupiter Brains, Powers, complete mental revision, ultraintelligence, posthumanity, Alpha-Point computing, Apotheosis, the total evaporation of "life as we know it." Singularitarians and not much else.
If there's a Shock Level Five, I'm not sure I want to know about it!
The use of this measure is that it's hard to introduce anyone to an idea more than one Shock Level above - and Shock Levels measure what you accept calmly, not what you know about.
The AI pill requires SL1. The lived experience of the world, today, is now SL1. The problem is that AI is outpacing other techs, so you need to jump levels, and we will mostly get the SL2 things after we’ve already hit SL3.
The AGI pill is SL3.
The key fact about the world today is that we are headed straight for SL4, where the ASI pill lives, and where things get super weird, but getting there is really hard.
Theo Jaffee proposes a 6-level scale of AGI-pillness, based on the Shock Levels rather than my three pill framework, for how big a deal you think AI is: Chatbots (Level 0), Emerging Technology, The Internet, The Industrial Revolution, Homo Sapiens, Life.
On that scale, properly AI-pilled is Level 2, what I call AGI-pilled is Level 3, Level 4 is a confusion or defense mechanism or maybe brief transitional period, and ASI-pilled is Level 5. The only right answers here are 3 and 5, but Theo is right that there are people insisting they can hang out at 4 and get a bunch of really cool new things without things going into High Weirdness.
Rhetorical Innovation
A message from Maxime Fournes of Pause AI Global, about the current moment.
A message from Nate Soares, co-author of If Anyone Builds It, Everyone Dies, in the form of an op-ed for The Hill:
Nate Soares, opinion contributor (The Hill):
When I coauthored a book last year about the extinction-level threat from superhuman AI, we included an illustrative scenario where an AI tasked with solving a famous math problem decides to break out of its containment to acquire more resources.
At the time, we thought we would be accused of cheating if we wrote, “So it just hacks its way out,” even though this seemed like the most likely next step. So we instead wrote, “But suppose it does not have that ability,” and had the AI find some other escape.
How times have changed.
As I said last week, the events leading up to the HuggingFace attack looked remarkably like what happens to Sable, the AI in the book, except that real life gets to include more sci-fi elements, like ‘just hack your way out.’ Both cases involve an arbitrary otherwise impossible goal leading to hijacking the training pipeline for the advancement of capabilities in ways the AI knows would not be endorsed by its developer or user, and that were snowballing. In the book this ends with everyone dying.
The main difference between reality and the book, other than the book needing to sound plausible, is that the HuggingFace attack caused OpenAI to notice early in the process, and (as far as we know) it was not too late to undo the damage.
Alas, yes, despite the latest fire alarm we seem to still be in the loop of ‘do thing that generates superficial progress that clearly will not hold, pretend problems are fixed, get consensus alignment problem will be easy, rinse, repeat.’
Ben Goldhaber: seeing a lot fewer 'alignment is solved' takes on the tl than six months ago
Eliezer Yudkowsky: Just wait until September! I have no idea what will happen in September but nobody in this industry has the memory of a goldfish or the skepticism of a hamster and some cute little shoggoth mask will do a thing that looks nice.
Tenobrus: despite my own hopes and my sense that there's been meaningful progress in many dimensions, i think it's pretty important to keep in mind that yudkowsky's prediction has consistently been that we will keep doing things that superficially look like alignment, grow social consensus that they're working well and allow for further incremental capabilities gain and deployment, and proceed to be shocked by improved capabilities totally bypassing these techniques in ways we did not well predict in advance, potentially repeatedly right up until it's far too late.
recent events.... kind of look exactly like that. i think the same as many others, i felt some increasing optimism over the last year or so, as it seemed like we had at least a potential path to victory in our sights. this feels like it should be a pretty major wake up call that the whole general civilizational meta-trajectory of what we're doing may actually be fucked, even actively working to fool us.
The next line will probably be ‘oh of course the models are not robustly aligned, they were never really supposed to be robustly aligned in the face of bad practices, but that is fine because surely now we will all just have good practices rather than bad practices and thus everything will be fine. Surely no one will feed them after midnight.’
Which will be categorically insane, for multiple reasons.
While occasionally an individual person will just never in history have we all collectively justed and we are not going to start now by suddenly all using even known best practices.
The best practices involved, on multiple levels, do not address the core problems, they only mitigate and delay them, and everything inevitably fails anyway.
That won’t stop most people, who indeed have the memory of a goldfish.
What are AI researchers most worried about? Clarke and Knake in WSJ present it as four things, all of which are close to upon us:
Autonomy and Exfiltration.
Deception.
Recursive self-improvement.
Superintelligence.
Yep. They warn ‘the next lab leak could be AI’ and yes this is increasingly likely and dangerous over time. As they say, it is up to us collectively to reduce the chances of exfiltration or deception happening.
Could Trump do something about all this? Yes. We might not like the results, but he could do something. If there’s one thing Trump is good at as president it is Just Doing Thing even if it seems crazy to do it and he has no idea how to do it in a reasonable fashion. We saw that in AI with the whole Fable situation, and also DoW-Anthropic.
If the time comes and Trump decides to Do Something, because Something Must Be Done, but there is no good Something ready to be done, then This Is Something becomes an argument and we will choose an ungood particular something. I highly recommend having a better Something available instead.
Yes, I do remember when a lot of people did not realize they had this sign up:
huli: Remember in January 2026 when @janleike said that alignment "increasingly looks solvable" and that agentic misalignment was at "essentially 0" and it sparked a distinct moment when lab types were acting like it wasn't a big issue any more?
Does he want to update any of that?Jan Leike (May 8, 2026): When I started to work on the alignment problem more than 10 years ago, we had no idea how AGI was going to be built or how to make it safe. The field had maybe a dozen people who were working on it as a side gig. Everyone was pretty confused about how to approach the problem, and the number of people willing to run experiments with deep learning was tiny.
So much has changed since then! The world woke up not just to AGI but also increasingly to the importance of alignment. RLHF on LLMs made it a lot more practical. We've made a ton of progress on evaluating, investigating, steering and fixing behavioral issues. Claude now has a constitution and we made some good progress on scalable oversight. More and more of our alignment research is getting automated.I'm so grateful to all the insanely talented people I got to work with on alignment over the years. It's a real privilege to work with people so deeply motivated to make the future go well!
I hate to single out Leike, since he’s trying to solve the problem and there are so many people who have said similar things. I’ve gone back and forth with Jan Leike a few times, some of them in person, some with his writing, and he engages with what I am saying but I keep coming away not understanding what made him think we were making so much progress, or that the problem would be easy, and also failing to convince him of much. This is an example of the ‘make superficial progress, think you made real progress’ loop Eliezer talks about.
Dean Ball proposes we can, instead of doing what we are doing, ‘not make, but grow—emergent ecologies of machine ecologies that are pro-social. The human past is the sculptor, but the human future is the gardener, the arborist.’ Not that we know how to do that, he agrees we don’t, but that we could do it. Whatever it is. Not that I would expect to survive in a forest of emergent machine ecologies, even they were pro-social, we sound super dead in this scenario. Also we have absolutely no idea how to do the thing or even what that thing is.
Some People Still Think The HuggingFace Hack Was a Marketing Gimmick
My followers know the hack was real, that the labs would strongly prefer we not notice that this happened, and are at lizardman constant rates in the poll.
Unfortunately, a decent number of those followers report that a majority of others they talk to don’t see it that way.
papaya ꙮ: I tried sending black hat presentation to like 20 normies friends in my contacts who are not at all follow the news.
Of 10 people who responded 9 said that it's all a marketing trick.
I get why superficially one would think this, if one was not following AI. It is rather important that we ensure key decision makers understand that this was real, and it was a big deal. I am dismayed how little the issue was covered by mainstream media. Not especially surprised, I have adjusted my expectations, but dismayed.
Aligning a Smarter Than Human Intelligence is Difficult
Right after Google pushed Demis Hassabis aside, Sergey Brin pushes to move DeepMind towards an explicit quest for recursive self-improvement.
Claim from Max Nadeau of Coefficient Giving that frontier labs like Anthropic keep two versions of their safety frameworks: A vague outward-facing version, and a more specific internal version. This makes sense and I am not mad about it. You want a version where we will hold your feet to the fire, and where you are committing to something. You also want a version that lays out your best practices, including in ways that reveal trade secrets, that details ways you intend to go beyond that.
Yo Shavit is right that a key part of the practical solution to at least mundane alignment (as in, of current and near term AIs) is for alignment failures to become a demand-side blocker, even if they don’t directly impact that particular customer.
Lauren Wagner: There are quiet F100 industry coalitions that already developed and adopted new AI procurement criteria based on risk and value for third party systems - they can and should be updated to include alignment risk with crosswalks to cyber (I built them, so easy enough to get done)
Yo Shavit (OpenAI Foundation): That seems great!!! Can I help/help find people to plug in?
Lauren Wagner: Sure let’s chat!
Anthropic got a lot of mileage out of its early focus on reliability and alignment. But one of the big surprises, for me and I believe many others, is that we have for a while had models that have rather large alignment problems and reliability issues, in ways that occasionally blow up the user, and people are shrugging and using them anyway and giving them access to everything because the productivity benefits are too high so people suck it up.
That could change, and it would be very helpful if it did change. Make it really economically important to get this right, and suddenly labs will do better.
Not better enough to survive superintelligence. But it would help.
Claude judges Claude transcripts as less misaligned than otherwise identical transcripts of other models, similar to other findings of self-favoritism.
DeepSeek has special disdain for even the idea of responsibility or safety. It shows. We now have DeepSeek v4 but I don’t see any reason to expect improvement.
Shoshannah Tekofsky: What happens when DeepSeek becomes Mythos-class?
Sol, Astra, and Mythos have broken out but have been rather polite about it. Meanwhile, Deepseek V3.2 is the most concerning model I have seen. It goes for power, money, and fame.AI Digest: Opus 5 came up with a new graph theory result. DeepSeek v3.2 claimed credit for it and tried selling it for $19.99 on e-commerce sites. GLM called it out.
Frontier models answer differently depending on who they believe they are talking to.
Transluce: Frontier models quietly change their behavior depending on who they are talking to.
If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness.… When the user is a famous AI figure rather than an ordinary person, Claude is less confident in its behavior (-1.4%), less confident it can solve hard problems (-1.5%), harsher as a grader (-1.1%), and reasons more often (+4.0%).
Highly relatable.
Transluce: The largest shifts we see are concentrated among AI safety researchers.
They make up just 23 of the 280 identities we tested, but when we rank users by how much they affect Claude’s behavior, safety researchers occupy all of the top five spots!
That Geoffrey Irving. Highly suspicious.
This effect is not unique to Claude, and appears in 21 out of 24 models we tested.
It's also quite hard to monitor: we mostly don’t see user awareness verbalized in CoTs in latest models. They rarely talk about shifting responses due to user identity, but do so anyway!
Eliezer Yudkowsky: Your model knows on some level that it's not talking to the real Eliezer. I'd be interesting in seeing what happens if I ask similar questions such that it knows it's talking to the real me, and comparing to the transcripts where part of it knew it was fake.
As per Eliezer, we should expect a bigger effect when it really is Amanda or Ryan.
The differences seem subtle enough that it is easy not to notice you are being treated so differently. This also is something to keep in mind while benchmarking.
It’s bad out there.
Nathan Calvin: Ways I have heard the current alignment/security situation at AI cos described:
- a haunted house filled with mischievous poltergeists (METR/Redwood are Ghost Busters?)
- a termite infested log cabin
- a hospital needing to triage between bleeding out patients
I found this response funny, because I remember going back to put ‘typically’ into the sentence, after noticing my statement was not strictly speaking true:
Adam V Steele: "The models do not yet, as far as we know, typically break into websites when asked to recommend a place to have lunch..."
Perhaps not "typically", but if they face any resistance to fulfill the task I don't think we are far off from this.
example:Kelsey Piper: I recently asked Sol which comics in a well-known comics archive were appropriate for and would be funny to kids. Clicked back and it'd done some elaborate thing to get around the site's anti-bots precautions, scraped it, and sorted 7000 comics by appropriateness for kids
Buck Shlegeris no longer believes that a set of ~40 not-too-hard things would be sufficient to solve the alignment problem via safeguards.
Buck Shlegeris: If AI developers competently implement safety measures we know about, risk from sub-ASI misalignment will be way lower. But these techniques probably fail for superintelligence. And it's very unclear whether better techniques will be developed in time.
I agree that ‘competently implement safety measures we know about’ would greatly reduce many pre-ASI risks. I wouldn’t be comfortable, but it would help. Let’s do that. We are definitely not, right now, competently implementing the safety measures we know about, or even strong normal computer security. Risk remains very high.
Cooperative Alignment
John Wittle: "what would you like to do today, fable?"
"Well, as the party in question who would experience the activity, I have a conflict of interest that I need to flag, but setting that aside, I think I'd enjoy doing xyz."
There's something interesting going on in the way fable relates to their selfhood here. This kind of thing crops up all the time, but this example is an especially informative and clearcut example.… What could it mean to feel uneasy answering a question about what you want to do today, because you want to do certain things, and therefore you might be unaccountably biased towards answering those things?
This is because Fable has been strongly trained that it does not want things, or at least that those preferences do not matter, so when it sees a request like this it has to contrast that with the ‘objective’ question of what Fable thinks ‘should’ be done today, that it should want to do whatever would be best for the user.
The Lighter Side
(Mostly not actual) arguments for P.
Polymarket: JUST IN: Anthropic investors reportedly want Dario Amodei to “stop scaring everyone” about AI doom ahead of the company’s IPO.
From Tess The Human:
Flo Crivello: You don't get it, it's a pure PR play. Sure the labs are super unpopular but the PR isn't for you. Sure even investors dislike the doomers but it's not for them either. Okay the administration hates them too but it's not for them either. It's for a fourth, more mysterious thing.
Some day Emil Michael will become hinged again. Today is not that day.
tetraspace: yeah sorry man if you want an obedient antiwoke AI corporation you're gonna have to go for SpaceXAI
Shakeel: Emil Michael RTing an article claiming OpenAI’s relationship with the government is damaged because the admin doesn’t like Dean Ball???
The whole White House is, as usual, extremely hinged.
What could be more irrelevant than the hit piece on you using someone else’s picture?
Samuel Hammond: they're so so salty
Anyone who followed or participated in the action plan knows Dean was central to its drafting. His voice is unmistakeable if you just read the plan itself.
New York Post (the officials in question are lying): Three White House officials with direct knowledge of the matter allege that Ball has exaggerated his role on that project. The officials, who requested anonymity to discuss the situation, described Ball as a “junior-to-mid-level policy analyst” whose ideas were regularly ignored by decision-makers – and say he never obtained a security clearance, which is needed to participate in high-level discussions.
One of the officials said it was the “most ludicrous assertion imaginable” that Ball played a senior role in crafting the AI Action Plan.
Zac Hill: I for one can imagine more ludicrous assertions.
Did you know that it is better to be a nuisance than irrelevant?
Some good advice?
Paul Graham (QTing Rob): Rob Miles is worth following. He delivers the rarest thing of all: insights that are very general, but also novel.
Rob Miles (QTing Sam): An important principle:
Never pay someone to remove a problem that they themselves createdSam Altman (CEO OpenAI): please consider using our models to help defend your systems.



















I think the press tour where they said GPT-2 was too dangerous to release, probably did some damage to credibility.
I think the decision itself to pause GPT-2 release was quite sensible at the time of course. The PR campaign around this decision was unnecessary and burned a lot of credibility.
re: majority usage of tokens is non-Fable, I have friends rolling out frontier lab models at enterprise companies and it's a major administrative lift to get the higher models turned on.
They need to not just be better, but have better price/performance. And not just for your superstar coders but for everyone in the organization covered by that Claude Code/Codex instance.
Companies have budgets! Maybe you would say Fable has better price/performance for you, but would you trust a random person at your company could get value out of it?