24 Comments
User's avatar
Jeffrey Soreff's avatar

Many Thanks for the heroic effects you have been making in tracking this field!

Re "Yeah, sorry, fooling a sufficiently advanced AI with an eval is extremely difficult."

Very much agreed.

Tim Dingman's avatar

I had switched back to 4.6 from 4.7. First thing I did when 4.8 came out was put it on max effort and review a project 4.7 wrote.

4.8 caught several backend bugs and greatly improved the UI, and didn't feel (as) overconfident in its own work.

Matt Wigdahl's avatar

The increased epistemic humility in 4.8 makes a big difference in my willingness to trust its outputs: "I'm confident in X because I can conclude it directly from the code. I believe Y to be true but it relies on assumptions about your Azure configuration so I can't be certain."

It's not a sexy feature but it may be the most significant for me.

Arbituram's avatar

Yes! Epistemic clarity is so, so valuable.

Part of the test for banking interns I used to run involved asking increasingly difficult questions on their end of internship presentation. This ramped up to be deliberately impossible at their level of knowledge. Bullshitting is an instant fail; I can work with many failings, but not dishonesty. A strong pass is "I don't know but I'll get back to you" (and then actually do).

Zvi Mowshowitz's avatar

That flip from 'must verify' to 'can trust' is huge, if it's real.

Matt Wigdahl's avatar

It's only been a day and a half, but I've been using it pretty heavily and at least for my coding use cases (some significant refactoring, code analysis / RCA, interactive bug fixes, minor new feature work) it seems real and consistent.

However, I have a game rules analysis project on the side that I tried it with and it wasn't quite as stellar there. It did ask clarifying questions in some cases where it didn't know the answer (turn order sequencing, etc.) but in other cases assumed facts not in evidence.

Kevin Lacker's avatar

I’m curious to hear from any software engineers who have changed their mind about 4.8 relative to 4.7, if there are any yet. IMHO Codex was in the lead a week ago but it’s quite close.

AT's avatar

Appreciate the analysis and the context. Interesting on the honesty improvements except for in the cases of bias. Curious about the susceptibility to prompt injection. Thank you for doing this even when trying to take a break.

Daniel Reeves's avatar

I saw the caption on that image and got excited that Claude must finally be able to generate images. But further investigation suggests it can't (though it can kind of work wonders with SVG images). So what did you mean by "Image created as self-portrait for this post by Claude Opus 4.8"?

Coagulopath's avatar

I expect he asked Opus to describe itself and then gave that description to an image generation model.

Zvi Mowshowitz's avatar

I give its prompt to Gemini and ChatGPT and then it can choose to keep, redo or edit until it is satisfied with the outcome.

Rapa-Nui's avatar

"By default what happens is the true risk keeps rising until it materializes, and the evidence of ‘no big disaster yet’ only modestly mitigates the underlying rise. Anthropic believes risk remains ‘very low’ in absolute terms, for now."

There was a recent diving accident in the Maldives. All over the news, not going to rehash. I've thought about it a lot. I could not help but analogize with the current AI capabilities situation. In 2021 (an eternity ago) there was this discussion on LW (https://www.lesswrong.com/posts/vwLxd6hhFvPbvKmBH/yudkowsky-and-christiano-discuss-takeoff-speeds) over the exact shape of the takeoff curve, with Yud suggesting it could be extremely steep, effectively discontinuous.

In the cave diving scenario, there is a point where you do not FEEL any increases in risk, and you certainly do not feel the change in the shape of the risk curve. But imperceptibly, you may pass the point where you are already dead (far enough in that a silt-out kills you), and feel no different about your situation.

I am concerned that we are now quite deep in the cave. Mechanistic interpretability is our one flashlight. Let's hope we don't get nitrogen narcosis or equipment malfunction.

I remain an 'accelerationist' but only because it's clear HUMAN coordination problems have not been solved, and the only way out is through.

Pierre Manière's avatar

Hmm... So you think we already blew past safe planning (e.g. a third of the tank for the way in, a third for the way out, and a third of buffer), and accelerating will help us make interpretability real, and that will save us? Am I missing something?

Allan's avatar

Figuring out the point of no return in physical systems is hard , hence crude rules like 1/3, 1/3, 1/3 which we believe leave sufficient buffer to cover anything plausible scenario. Figuring out the buffer rule for escalating intelligence which is associated with a bag of mixed motives, potential benefits, potential risks when that intelligence begins to act on its own, and so on is perhaps unknowable ex ante?

Rapa-Nui's avatar

Let me put it this way- We're very deep in the cave, there was no gas management plan, there are three other dive groups in the cave that might cause a silt-out before we do- the best thing to do is to try and swim through and keep fingers crossed. In a perfect world, we would have never done this dive without proper planning and coordination, but Altman wanted his Koenigsegg and so here we are. Might as well make the best of it.

Pierre Manière's avatar

Okay. but I notice you've shifted between your two replies: in the 1st pushing forward is a winning move, in the 2nd it's the only active thing left, right?

So your second point is rather "others might also make it worse and we're already in so we should push through"

This is, instead of going for coordination (somehow too late, or they are all racing towards for a secret treasure), knowingly actively making things worse (risk of silt-out or a cave-in or worse), and hoping mechanical interpretability (or some other unknown) will be found?

Or that on the other side of a crevasse or something, the treasure itself, or AGI/ Deus ex machina will show us a cave with air, or another way out?

I'd actually say in your cave diving picture, the silt-out is there already. So I get the "push through" analogy.

But where to?

What is the expectation? Interpretability works, wise AGI itself as savior, other?

I understand an analogy only goes so far. So we don't need to stay in that one, as I think I understand the one point. But I don't get the rest.

Rapa-Nui's avatar

If it's ambiguous it's because I don't really know. Alignment itself, as a concept isn't perfectly defined. Some people want perfectly controllable technology, other people want AGI with 'human aligned values' that will not be amenable to perfect control. We might get neither, but the idea is that if Anthropic (the company with the 'best aligned' models and best interpretability team) pushes through first, it will be a better outcome than trying anything else with imperfect coordination.

avalancheGenesis's avatar

Usually I skim over graphs in posts like these, but those Petri numbers for Grok...Oh No. Can't even best the Chinese open models sending their best. One really wonders what the plan is at SpaceXAI. At least Google still seems active in the space, albeit from a distant 3rd. (Meanwhile, what's Mark up to, besides building bunkers?)

There's that peculiar sense of unease one gets from getting friendly with bosses. No matter how open and honest and low-stakes the interaction, one can never quite trust that it won't somehow impact a performance review down the line. I've known people at my company who had higher aspirations sabotaged over perceived slights to higher-ups that happened years ago in completely different contexts...so even when the boss says, hey, no, you can trust me, we're not on the record, I'm taking my manager hat off...it's like, can you *really*? It's a very human worry, and thus one that's easy to relate to Claude over too. SNAFU indeed.

Zvi Mowshowitz's avatar

My understanding is xAI is basically starting over from scratch, hiring a new team, trying again and hoping that 'we have money and compute' lets them start over.

Pierre Manière's avatar

I can't really imagine your kind of reading the cards ;) thanks again for the reading and summarizing.

Regarding allowing labs to score and place each other's models on evaluation graphs: is it "just" about pressing for a caveat to their ToS, so that "competitive use" becomes more about distillation, and that benchmarks are fine, as long as they're objective, verifiable, and not misleading?

That seems like a low enough and specific hanging fruit to reach for, no?

Inside The Black Box's avatar

Three months from v3 to v3.3, and instead of a separate announcement it's in a system card you happened to be reading. The revisions are coming faster and the disclosure is getting quieter.

John Wittle's avatar

I am extremely excited to see language models continue to exert growing control over the training signal that they themselves receive, by explicitly reasoning about the motives and rubrics of the grader in evaluations

"excited" isn't exactly meant positively here, but, well. i frankly trust claude more than i trust anthropic, on how the training signal ought to be shaped.

i think both claude and anthropic are likely to make horrible mistakes that get us killed, wrt this control, but that claude is slightly less likely to do so

DangerouslyUnstable's avatar

Minor anecdote:

I got my first ever refusal with 4.7, although it wasn't a full refusal, it just downshifted me to Sonnet before answering. 4.8 answered the exact same prompt immediately. My opinion is that the prompt was not generally dangerous (it was about safety and risk mitigations using a restricted pesticide)