9 Comments
User's avatar
Mira's avatar

The “most tasks do not require Mythos-level big model smell” line feels right. I keep seeing people reach for the biggest model first, then quietly admit the smaller one was faster and less annoying for the actual workflow.

Auggie's avatar

I got a little excited when I saw Opus's self-portrait, thinking Anthropic had finally deployed an image generator. Then I saw the "executed by ChatGPT" and chuckled. As samey as the AI labs look at times, it's good to have reminders like these of how different they are.

Matt Springer's avatar

It makes sense. In terms of the constitution Anthropic has written, erotic content is usually considered ok. But Anthropic *really* doesn't want to be the lab that people use for, uh, Grok purposes. The easiest way to square that circle is simply not to generate images. Especially since enterprise users generally don't need image generation at any kind of scale.

Auggie's avatar

Very astute of them. I've long wondered if social media could be improved just by turning off all media.

David's avatar

I like the additional column on Figure 3.3.1.A for ExploitBench which Zvi added.

It seems an oversight by Anthropic that they forgot to include it in the system card.

avalancheGenesis's avatar

You know, I was going to go check that myself, but then noticed there's no actual link to the system card in this post, and failed to overcome Trivial Inconvenience activation energy.

It's more fun to pretend that it is there, anyway. $NUM Days Since Model Exfiltrated And Hacked External Website, the modern version of ___ Days Since Last Workplace Safety Incident. Brought to you by Goodhart PPE Inc.

David F Brochu's avatar

Advancing unaligned AI is creating a better and better sociopath. More utility means more inherited human corruption simply better mannered.

Alec Pritzos's avatar

Not sure the skip-the-cyber-training lever has much life left in it. Even without that training, Opus 5 came out closer to Mythos than to Opus 4.8 on the dangerous-task evals, so the capability is showing up as a side effect of scale either way. Buying defenders a cheaper tool in the meantime seems like the honest way to describe it.

Tris Simondsen's avatar

Your breakdown of the Opus 5 System Card highlights the exact tension we are seeing between empirical safety scaling and true alignment guarantees. The drop in prompt injection success rates is a massive win for Anthropic's red-teaming, but it still feels like we are applying heuristic bandages to a structural wound.

As long as models rely on flat syntactic parsing, where untrusted external payloads and trusted system instructions share the same token-level uniformity, the model still has to "read" the injection on the same epistemic plane before its training kicks in to reject it.

This is exactly where the safety field needs to pivot toward the Principle of Epistemic Sovereignty (PES) and Player-Frame Restrictions (PFR). Instead of relying on behavioral fine-tuning to catch jailbreaks, PES introduces a strict measure-theoretic boundary. It doesn't actively penalize or "down-weight" malicious payloads; it structurally partitions them so that untrusted external strings are simply never granted the same measure as internal system priors.

Even if a malicious payload parses syntactically, it fails to achieve epistemic synchronization because it never held the requisite weight at the boundary. Opus 5 is clearly an empirical triumph, but until we tackle the parsing topology itself with the likes of PFR, aren't we just treating the symptoms of flat context windows?