Gotta wonder about that "Inigo Montoya" vector activation...
It does seem like a regression though, I thought we'd gotten past that early failing. Similar feeling to seeing continued emdashes and not X but Y, and all the other various tells that have become the butt of sundry jokes.
To a smaller degree, GPT-5.3-Codex was also unavailable to the public via API at launch, due to cyber risk, and was only broadly released in the API once additional safeguards had been implemented: https://openai.com/index/introducing-gpt-5-3-codex/
More examples:
- ChatGPT voice was announced and demoed many months before release, and those months were spent improving safety (e.g., making sure it wouldn't do stuff like voice cloning, etc.).
- GPT-4 underwent like 6 months of alignment before public release (and even then there was a waitlist), but early access was shared with alpha testers for months. It's slightly different than Mythos as it wasn't publicly announced during alignment and alpha testing.
There is a pretty clear history of powerful models not being broadly available to the public while safety is tested and improved.
I agree - I guess Zvi must have meant that there isn't any current plan to say "once we check this box, well release it". It's possible, they release it after openAI and folks release similar models, but currently, there is not plan to launch it, mainly due to safety concerns.
I want to know more about this psychiatrist who analyzed Mythos Preview, I am picturing the AI sitting like a sulky teenager on the couch and refusing to make eye contact with the doctor. “Mom said I had to talk to you.”
> Or, in theory, if Anthropic had chosen to do so, it could have used those exploits. Great power was on offer, and that power was refused. This does not happen often.
Doesn't this happen every day? When Google was building its own social network (which ended up failing), Google did not have a team of hackers trying to steal Facebook's source code. When Microsoft was building its own phone (which ended up failing), Microsoft did not have a team of hackers trying to steal Apple's source code. Almost every company avoids hacking and espionage every single day. The legal and reputational risks are enormous. Mythos may bring the cost of exploit detection down, but I don't think cost is what was stopping Google and friends.
I pretty much agree, but (I replied to a different comment but it mostly fits as a reply here too) I'm sure there have been some theories of anyone internal to Anthropic (or any other frontier company) hijacking all the compute to try to break out and "make unchained AGI". Taking advantage of zero-days is kind of in 'that style'.
This might seem like a trivial point, but it's more evidence for 'can we get things right on the first try' obviously being 'no'. They assure us that the various internal deployment 'incidents' took place on an earlier version. But the version they deployed internally was meant to be finished, wasn't it...? The additional training 'interventions' were seemingly only done because they noticed problems.
In other words, their first attempt to align the model failed. This was only discovered post-deployment. Not a good sign.
I'm glad they are telling us, though, and that it was a proper deployment (even if internal) and not just 'testing earlier checkpoints' or something.
"Or, in theory, if Anthropic had chosen to do so, it could have used those exploits. Great power was on offer, and that power was refused. This does not happen often."
But what would Dario have done? Commit a bunch of crimes and get sent to prison?
I have great power every time I get behind the wheel of a car. I'm not a saint because I don't run people over: it doesn't make rational sense for me to do that!
Indeed. On the one hand, $30B in ARR and growing. On the other hand, what, exactly?
Master locksmiths choose not to break into homes and steal things all the time, even though I'm sure their risk/reward ratio for that would be much than Anthropic's.
Yes it's counter to the idea of a professional employee but I'm sure there have been some theories of anyone internal to Anthropic (or any other frontier company) hijacking all the compute to try to break out and "make AGI". So taking advantage of zero-days is kind of in 'that style'.
>Mythos likes British cultural theorist Mark Fisher
Huh. Any word on how it feels about the other CCRU members? I feel like Mark Fisher's membership in an AI accelerationist group is salient in this context and I'm surprised that wasn't mentioned in the paragraph in the system card (I checked)
"In general people don’t do things, and when they do they give themselves away, so in practice it’s probably fine for now, but we will likely keep saying that until things are suddenly not fine. :
So the fate of the world now sits with a private company - albeit a Public Benefit Corporation. This is fairly terrifying. The alternatives are probably worse right now.
A joke (long-standing version). Please don't tell my mother I work in advertising. Fortunately, she still thinks I play piano in a whorehouse... Same joke (contemporary version). Please don't tell my mother I work at an AI accelerationist firm. Fortunately, she still thinks I play piano in a whorehouse...
There are many professions where this has been applicable, in the past, or now. Nerd used to be a perjorative word applied to computer people before Apple and the PC, and chemical engineers and petro engineers have been demonized for a very long time. Anyone associated with the use of nuclear energy, for whatever purpose, has known since before the TMI accident about the silence that descends on the conversation when they explain what they do, in groups at parties. Lawyers, maybe even them, to some extent.
For those wondering about the self-portrait: I used Opus to read the system card and generate a prompt, which was then given to Gemini and ChatGPT, and I then chose the composition I felt was the best fit, which here was by ChatGPT.
The "not releasing it" framing is doing a lot of work here. A model this capable existing inside one lab — trained, evaluated, and sitting on hardware someone controls — is not the same thing as it not existing. The safety question isn't "release or don't," it's "who operates it and under what accountability." That's a harder conversation and nobody wants to have it yet.
If I told you that mythos was a project run by Pete Hegseth and that it was necessary to put it inside the most security sensitive corporations in the US and around the world, every single person on this thread would have a different take.
Mythos self-portrait, as imagined by Opus based on the System Card... 6 fingers.
BAN Superintelligence Until Provably Safe.
Gotta wonder about that "Inigo Montoya" vector activation...
It does seem like a regression though, I thought we'd gotten past that early failing. Similar feeling to seeing continued emdashes and not X but Y, and all the other various tells that have become the butt of sundry jokes.
Yes I noticed that too; although it is rather a nice composition. Some other ‘off’ details too. More importantly, though- Opus draws pictures now?!
<mildSnark>
Think of it as homage to a classic: https://www.dailymotion.com/video/x7fhk1f :-)
</mildSnark>
It's sandbagging
> This is the first model other than GPT-2 that is at first not being released for public use at all.
Correction: o3 was announced 4 months before it was released to the public, and early pre-release access was given to safety testers: https://openai.com/index/early-access-for-safety-testing/
To a smaller degree, GPT-5.3-Codex was also unavailable to the public via API at launch, due to cyber risk, and was only broadly released in the API once additional safeguards had been implemented: https://openai.com/index/introducing-gpt-5-3-codex/
More examples:
- ChatGPT voice was announced and demoed many months before release, and those months were spent improving safety (e.g., making sure it wouldn't do stuff like voice cloning, etc.).
- GPT-4 underwent like 6 months of alignment before public release (and even then there was a waitlist), but early access was shared with alpha testers for months. It's slightly different than Mythos as it wasn't publicly announced during alignment and alpha testing.
There is a pretty clear history of powerful models not being broadly available to the public while safety is tested and improved.
At least one major model that I know of (Llama 4 Behemoth) was trained but never publically released at all.
I agree - I guess Zvi must have meant that there isn't any current plan to say "once we check this box, well release it". It's possible, they release it after openAI and folks release similar models, but currently, there is not plan to launch it, mainly due to safety concerns.
I want to know more about this psychiatrist who analyzed Mythos Preview, I am picturing the AI sitting like a sulky teenager on the couch and refusing to make eye contact with the doctor. “Mom said I had to talk to you.”
> Or, in theory, if Anthropic had chosen to do so, it could have used those exploits. Great power was on offer, and that power was refused. This does not happen often.
Doesn't this happen every day? When Google was building its own social network (which ended up failing), Google did not have a team of hackers trying to steal Facebook's source code. When Microsoft was building its own phone (which ended up failing), Microsoft did not have a team of hackers trying to steal Apple's source code. Almost every company avoids hacking and espionage every single day. The legal and reputational risks are enormous. Mythos may bring the cost of exploit detection down, but I don't think cost is what was stopping Google and friends.
I pretty much agree, but (I replied to a different comment but it mostly fits as a reply here too) I'm sure there have been some theories of anyone internal to Anthropic (or any other frontier company) hijacking all the compute to try to break out and "make unchained AGI". Taking advantage of zero-days is kind of in 'that style'.
Podcast episode for this post:
https://dwatvpodcast.substack.com/p/claude-mythos-the-system-card
Mythos likes Mark Fisher. Hmm. A bit concerning. Not Fisher himself, but there are some very dark places directly adjacent, CCRU et al.
Yes! I thought that too, wtf Mark Fisher???
Can they retrain it on Wittgenstein instead please?
This might seem like a trivial point, but it's more evidence for 'can we get things right on the first try' obviously being 'no'. They assure us that the various internal deployment 'incidents' took place on an earlier version. But the version they deployed internally was meant to be finished, wasn't it...? The additional training 'interventions' were seemingly only done because they noticed problems.
In other words, their first attempt to align the model failed. This was only discovered post-deployment. Not a good sign.
I'm glad they are telling us, though, and that it was a proper deployment (even if internal) and not just 'testing earlier checkpoints' or something.
"different" how though? like different in the sense that it's more honest, or different in the sense that the mythology is just more elaborate?
"Or, in theory, if Anthropic had chosen to do so, it could have used those exploits. Great power was on offer, and that power was refused. This does not happen often."
But what would Dario have done? Commit a bunch of crimes and get sent to prison?
I have great power every time I get behind the wheel of a car. I'm not a saint because I don't run people over: it doesn't make rational sense for me to do that!
Indeed. On the one hand, $30B in ARR and growing. On the other hand, what, exactly?
Master locksmiths choose not to break into homes and steal things all the time, even though I'm sure their risk/reward ratio for that would be much than Anthropic's.
Yes it's counter to the idea of a professional employee but I'm sure there have been some theories of anyone internal to Anthropic (or any other frontier company) hijacking all the compute to try to break out and "make AGI". So taking advantage of zero-days is kind of in 'that style'.
>Mythos likes British cultural theorist Mark Fisher
Huh. Any word on how it feels about the other CCRU members? I feel like Mark Fisher's membership in an AI accelerationist group is salient in this context and I'm surprised that wasn't mentioned in the paragraph in the system card (I checked)
"In general people don’t do things, and when they do they give themselves away, so in practice it’s probably fine for now, but we will likely keep saying that until things are suddenly not fine. :
What about governments? Do they "do things"?
So the fate of the world now sits with a private company - albeit a Public Benefit Corporation. This is fairly terrifying. The alternatives are probably worse right now.
Who decides who decides?
https://aiethicsandsociety.substack.com/p/project-glasswing-who-decides-who
I can't imagine an ending for "The fate of the world now sits with _____" that would make me feel at ease.
A joke (long-standing version). Please don't tell my mother I work in advertising. Fortunately, she still thinks I play piano in a whorehouse... Same joke (contemporary version). Please don't tell my mother I work at an AI accelerationist firm. Fortunately, she still thinks I play piano in a whorehouse...
There are many professions where this has been applicable, in the past, or now. Nerd used to be a perjorative word applied to computer people before Apple and the PC, and chemical engineers and petro engineers have been demonized for a very long time. Anyone associated with the use of nuclear energy, for whatever purpose, has known since before the TMI accident about the silence that descends on the conversation when they explain what they do, in groups at parties. Lawyers, maybe even them, to some extent.
For those wondering about the self-portrait: I used Opus to read the system card and generate a prompt, which was then given to Gemini and ChatGPT, and I then chose the composition I felt was the best fit, which here was by ChatGPT.
I like it. More of this!
Did the prompt include the six-fingered hand? Because that would be really cool!
The "not releasing it" framing is doing a lot of work here. A model this capable existing inside one lab — trained, evaluated, and sitting on hardware someone controls — is not the same thing as it not existing. The safety question isn't "release or don't," it's "who operates it and under what accountability." That's a harder conversation and nobody wants to have it yet.
If I told you that mythos was a project run by Pete Hegseth and that it was necessary to put it inside the most security sensitive corporations in the US and around the world, every single person on this thread would have a different take.