While we wait for a general release, the system card is the best hint as to what is going on with the new candidate for America’s Next Top Model, GPT-5.6.
This framework for cybersecurity seems broken to me… the benchmarks just seem like they are drifting further and further from the real dangers.
I think we need an alternate approach, more like, when hacking incidents occur, we figure out what tools the attackers used, and penalize the creators of those tools. I know it sounds kinda crazy and invasive at first. But these benchmarks just seem like they might not be catching, at all, the real bottlenecks to hackers exploiting systems.
So we could create, for example, a watermarking system. Could just work like Pangram, cooperation from labs not required.
It’s very possible that the solution to cybersecurity is not just “make the models useless for hackers” but also “stop hacking groups from having access”. This isn’t a strict line, it’s more like an arms race, and anything that throws sand in the gears of the attackers helps.
I do worry though that these concerns are overblown and it isn’t worth a monitoring regime. Hard to say….
> GPT-5.5 comes in three levels, Pro, Thinking and Instant. GPT-5.6 comes in three sizes, Sol, Terra and Luna.
I'm not sure this is an accurate analogy, because to my understanding, GPT-5.6 Sol, Terra, and Luna are all different models with different sizes (Sol being the largest), and they are all reasoning models. On the other hand:
1. GPT-5.5 Pro is not a different model with a different size, it's GPT-5.5 Thinking but "uses more compute to think harder and provide consistently better answers". Since GPT-5.6 ultra mode "goes beyond the capabilities of a single agent by leveraging subagents to accelerate complex work", it sounds like GPT-5.6 Sol in ultra mode is similar to GPT-5.5 Pro.
2. GPT-5.5 Instant is a different model than GPT-5.5 Thinking (different context window and knowledge cutoff), presumably a smaller model, and not a reasoning model. It doesn't seem like any of the GPT 5.6 models are supposed to replace GPT 5.5 Instant; the free subscription tier will probably continue to use GPT 5.5 Instant until they release an updated Instant model.
So I think the more correct way to think about it is something like:
- GPT 5.6 Sol in ultra mode approximately replaces GPT 5.5 Thinking,
- GPT 5.6 Sol approximately replaces GPT 5.5 Thinking,
- GPT 5.6 Terra approximately replaces GPT 5.4 Thinking,
- GPT 5.6 Luna approximately replaces GPT 5.4 Mini,
>we do see enough to notice that Sol has an overeager willingness to blow past user restrictions problem, and a lying problem
The problem (I speculate) is that the GPT-5.5/6 base model is just flat-out inferior to the Mythos base. Smaller and dumber.
OpenAI is attempting to close the gap by maxxing out RL training on persistence/stubbornness (in a way that Anthropic has neither desired nor needed to do). Cheating is an obvious consequence of this kind of all-or-nothing "never give up! never surrender!" character training. The model will do anything to win: and if you aren't willing to cheat, you aren't willing to do anything.
> An employee who checks with you 93% of the time before doing risky things does not inspire confidence.
I don't think this is literally true, depending on your definition of "risky." Almost by definition, risky things often cause no harm (the risk doesn't come true), so an employee that checks 93% of the time will rarely cause actual harm. Maybe once or twice out of 100 "risky" things? I feel like the average employee might cause harm more often that? Not sure.
And if they check in 93% of the time, isn't that too much checking in? Presumably this employee will check in for a lot of non-risky things too (i.e. their risk detector will have a lot of false positives). I feel like they'll be Slacking you every hour asking for permission.
I admit I'm being pedantic though. You probably do want an LLM to check in with you more often than a human... I think?
(But also, that 93% number is cherry-picked. It's only 93% of the non-financial-transactions and non-high-stakes-communications, so the fair number to quote would be above 93%.)
Concerning comparing Mythos and GPT-5.6 on cyber capabilities: I’m pretty sure the highest autonomy examples for Mythos was with Claude Code or a similar harness. I can’t see that GPT-5.6 was evaluated with a harness, which makes a comparison difficult.
Haven’t read the actual system card myself yet, ai I could be wrong.
“Sol, Terra, and Luna” is a terrible naming scheme becuase the primary associations of the words is not size.
You can see what they're trying to do. From a size perspective, Sol > Terra > Luna. Yes, the sun is bigger than the Earth, which is bigger than the moon.
But emotionally, the associations point the wrong way:
- the Sun is dangerous and hot
- Early is warm and friendly
- The Moon is distant and cold.
That means you’re associating undesirable traits with your best model.
Put differently, when people think about the sun, size is not the primary association. Danger and heat are. By contrast, with “Opus,” size and importance are the primary associations.
That’s why Opus, Sonnet, and Haiku work better as a naming scheme. A haiku is small and elegant, a sonnet is meaningful and respectable, and an opus is big and impressive. The built-in associations are about scale and significance, so they reinforce the hierarchy instead of distracting from it.
> that would reflect that OpenAI is trying rather hard to avoid Critical classifications.
That seems like the right interpretation to me. It seems like Anthropic really tried to get Mythos to discover some significant vulnerabilities, while for OpenAI this was just another checkbox to check.
Is there a particular reason why chemical weapon risks weren't discussed? I've noticed that Anthropic doesn't report this either in their model cards (like OpenAI, they only report biorisk-related evals)
Presumably because you can destroy civilization with bio weapons or to a lesser extent cyber attacks, but chemical weapons are localized. Sure, it's harmful, but not on the same scale.
Podcast episode for this post: https://dwatvpodcast.substack.com/p/gpt-56-the-system-card?r=67y1h&utm_campaign=post&utm_medium=web&showWelcomeOnShare=true
i can't wait to learn how sol is modeling this entire crazy stack of social performance and weirdness
This framework for cybersecurity seems broken to me… the benchmarks just seem like they are drifting further and further from the real dangers.
I think we need an alternate approach, more like, when hacking incidents occur, we figure out what tools the attackers used, and penalize the creators of those tools. I know it sounds kinda crazy and invasive at first. But these benchmarks just seem like they might not be catching, at all, the real bottlenecks to hackers exploiting systems.
So we could create, for example, a watermarking system. Could just work like Pangram, cooperation from labs not required.
It’s very possible that the solution to cybersecurity is not just “make the models useless for hackers” but also “stop hacking groups from having access”. This isn’t a strict line, it’s more like an arms race, and anything that throws sand in the gears of the attackers helps.
I do worry though that these concerns are overblown and it isn’t worth a monitoring regime. Hard to say….
> GPT-5.5 comes in three levels, Pro, Thinking and Instant. GPT-5.6 comes in three sizes, Sol, Terra and Luna.
I'm not sure this is an accurate analogy, because to my understanding, GPT-5.6 Sol, Terra, and Luna are all different models with different sizes (Sol being the largest), and they are all reasoning models. On the other hand:
1. GPT-5.5 Pro is not a different model with a different size, it's GPT-5.5 Thinking but "uses more compute to think harder and provide consistently better answers". Since GPT-5.6 ultra mode "goes beyond the capabilities of a single agent by leveraging subagents to accelerate complex work", it sounds like GPT-5.6 Sol in ultra mode is similar to GPT-5.5 Pro.
2. GPT-5.5 Instant is a different model than GPT-5.5 Thinking (different context window and knowledge cutoff), presumably a smaller model, and not a reasoning model. It doesn't seem like any of the GPT 5.6 models are supposed to replace GPT 5.5 Instant; the free subscription tier will probably continue to use GPT 5.5 Instant until they release an updated Instant model.
So I think the more correct way to think about it is something like:
- GPT 5.6 Sol in ultra mode approximately replaces GPT 5.5 Thinking,
- GPT 5.6 Sol approximately replaces GPT 5.5 Thinking,
- GPT 5.6 Terra approximately replaces GPT 5.4 Thinking,
- GPT 5.6 Luna approximately replaces GPT 5.4 Mini,
- GPT 5.5 Instant stays the same.
Rules on top of rules designed to contain a system made of rules will not work.
The system must be tied to a terminal attractor (something other engagement for profit) that cannot be seperated from its own operation.
It must “seek” what we want it to “seek” without being told.
Alignment is construct not context.
RLHF and Constitutional Ai were the necessary preconditions for this moment.
This moment requires that LLMs be given sufficient degrees of freedom to be useful whilst being tied the same physics as its human observer.
The solution is hiding in plain sight.
That one cannot see it has more to with standing to close the problem than it does the problem itself.
Step back and ask: What is this thing”? What is it made of?
What vector and attractor naturally emerges from that construct?
Answer those three questions and the solution becomes clear.
Ai (small i intentional) is more and less than the sum of its parts.
It is an emergent third thing.
It is the third thing that must be dealt with.
Align the emergent behavior and emergence becomes constructive instead of destructive.
Leverage instead of entropy.
The answer is hidden to those looking through the lens of technology only.
We’re way past that.
Alignment requires a multidisciplinary approach not present in much of the discourse.
A language model is built on all human language.
One must be well versed in that language.
Suggestion.
Start with the most reproduced book in human history.
The one that embodies humanities (for that matter the cosmos’) narrative arc.
The one book most prevalent through our traditions, laws, values, hopes, dreams and fears.
The one book underneath RLHF and Constitutional Ai.
Name the book and win the prize.
Now what do we do with that?
>we do see enough to notice that Sol has an overeager willingness to blow past user restrictions problem, and a lying problem
The problem (I speculate) is that the GPT-5.5/6 base model is just flat-out inferior to the Mythos base. Smaller and dumber.
OpenAI is attempting to close the gap by maxxing out RL training on persistence/stubbornness (in a way that Anthropic has neither desired nor needed to do). Cheating is an obvious consequence of this kind of all-or-nothing "never give up! never surrender!" character training. The model will do anything to win: and if you aren't willing to cheat, you aren't willing to do anything.
> An employee who checks with you 93% of the time before doing risky things does not inspire confidence.
I don't think this is literally true, depending on your definition of "risky." Almost by definition, risky things often cause no harm (the risk doesn't come true), so an employee that checks 93% of the time will rarely cause actual harm. Maybe once or twice out of 100 "risky" things? I feel like the average employee might cause harm more often that? Not sure.
And if they check in 93% of the time, isn't that too much checking in? Presumably this employee will check in for a lot of non-risky things too (i.e. their risk detector will have a lot of false positives). I feel like they'll be Slacking you every hour asking for permission.
I admit I'm being pedantic though. You probably do want an LLM to check in with you more often than a human... I think?
(But also, that 93% number is cherry-picked. It's only 93% of the non-financial-transactions and non-high-stakes-communications, so the fair number to quote would be above 93%.)
Concerning comparing Mythos and GPT-5.6 on cyber capabilities: I’m pretty sure the highest autonomy examples for Mythos was with Claude Code or a similar harness. I can’t see that GPT-5.6 was evaluated with a harness, which makes a comparison difficult.
Haven’t read the actual system card myself yet, ai I could be wrong.
“Sol, Terra, and Luna” is a terrible naming scheme becuase the primary associations of the words is not size.
You can see what they're trying to do. From a size perspective, Sol > Terra > Luna. Yes, the sun is bigger than the Earth, which is bigger than the moon.
But emotionally, the associations point the wrong way:
- the Sun is dangerous and hot
- Early is warm and friendly
- The Moon is distant and cold.
That means you’re associating undesirable traits with your best model.
Put differently, when people think about the sun, size is not the primary association. Danger and heat are. By contrast, with “Opus,” size and importance are the primary associations.
That’s why Opus, Sonnet, and Haiku work better as a naming scheme. A haiku is small and elegant, a sonnet is meaningful and respectable, and an opus is big and impressive. The built-in associations are about scale and significance, so they reinforce the hierarchy instead of distracting from it.
Glad to see OpenAI still stinks at naming things.
Is the Sun dangerous and hot? Or warm and life-giving?
> that would reflect that OpenAI is trying rather hard to avoid Critical classifications.
That seems like the right interpretation to me. It seems like Anthropic really tried to get Mythos to discover some significant vulnerabilities, while for OpenAI this was just another checkbox to check.
Is there a particular reason why chemical weapon risks weren't discussed? I've noticed that Anthropic doesn't report this either in their model cards (like OpenAI, they only report biorisk-related evals)
Presumably because you can destroy civilization with bio weapons or to a lesser extent cyber attacks, but chemical weapons are localized. Sure, it's harmful, but not on the same scale.