The Opus 5 "vibes are off" for me too, but for different reasons than mentioned above.
I've been using the Opus 5 a lot for knowledge work (I'm not really a coder), and it feels like a bit of a regression compared to 4.8.
Opus 5 needs more handholding, quicker to jump to (incorrect) conclusions, less capable.
Honestly feels like we're back to 4.5 levels of thoughtfulness where I need to once again ask the agent to review its work/recommendations and do "meta-cognition" more often, something which 4.6-though-4.8 had improved upon.
Reminds me also of the ChatGPT family of models, which -- while very capable -- have a hard time picking their head up and looking around to see if what they're doing so voraciously makes sense in context.
If I was a conspiratorial type of person, I'd say they've nerfect Opus 5 so that more people rely on Fable (and pay for more Fable usage). I don't think that's true, but my inner Agent Mulder "wants to believe."
I agree with the jumping to conclusions on high, but I also saw a lot of "final critique pass fixes over-hasty conclusions" in xhigh that mitigated that. Unfortunately xhigh for me does a little too much of this, undermining the overall quality with "on the one hand, on the other hand, on the third hand" type equivocation which it then insists on rounding down to some weird single compromise. As a result I am currently working with high and thinking of going back to 4.8 xhigh for some tasks. It does feel like too much subagent RL has broken the spirit of the model, or maybe it's that monster system prompt.
> It writes code that is slightly worse to look at than Fable, but more likely to be correct. It regularly is catching things that Fable missed.
What's interesting to me is how thoroughly different it seems to be to different people. My experience with Opus 5 was that it made mistake after mistake, failed to read the code. It "lied" repeatedly (said things which were obviously false given the context). Incredibly frustrating. I'd say it's worse than Opus 4.5 _for me_ at coding. But many people found it a big improvement or at least usable. I've seen enough twitter reaction threads and slack discussions at my job that I'm convinced both of these are valid experiences which actually happen.
Anthropic's prompting guide is bizarre, including optional instructions "To limit correction narration to corrections that matter" and stuff that seems like it should be in the system prompt. They indicate that you should minimize custom prompting because it can degrade performance, which seems like a strong indication of overfitting to a particular environment (presumably the default claude code harness with no CLAUDE.md and prompts written by Fable). An actually more general intelligence might not need custom prompting, but it shouldn't be vexed by it and do a terrible job.
My guess is that they've outsourced testing of lower-tier models to Mythos/Fable and are losing the plot on generating useful models for users.
Yeah, there's something to what you say at the end. Feels like Anthropic is disconnected from their model as a product and the everyday experience of the user.
I've thought about that too. Maybe people at Anthropic all use Mythos only so their second-best model receives no dogfooding and therefore is not polished in a way the Fable/Mythos is. Which would be a shame.
On AndyB_Bench (my entirely subjective measurement of how well a model does at my somewhat esoteric creative writing tasks, which centrally involve keeping track of specific worldbuilding premises and following their implications), Opus 5 is close enough to Fable that I have not been feeling tempted to burn Fable credits to try to get better results. A significant gap might become apparent over time, but it's not immediately obvious to me.
I have not had *any* trouble with Opus 5 being too argumentative or too negative; in fact I occasionally worry that gives in too easily. (It does push back, but maybe not as much as would be justified.) But I think I'm pretty good at establishing an attitude of collaboration rather than bossing it around, and judging from my experience relative to other people's takes I think the Claudes appreciate that (or behave as if they appreciate it, if one is of that persuasion).
Do people not write any instructions for Claude in the settings? I've always found it to be very accommodating. "Be succint." Sure boss.
Not telling Claude what you like, leaving the instructions empty, is a divide by zero error. You're liable to get anything. Just like not bothering to onboard a new employee.
My wife is particular about documents so she has given Claude .dotx fiiles and other templates, as well as some things she has written, as examples of the style she prefers. It's ~75% working for her in the Artifacts. I think she's enjoying teaching Claude Te Reo Māori (Maori language) in the chat.
Opus 5 is the first time I've actually been annoyed at a model: I basically do almost entirely biochemistry work, and it's so theoretically smart, but so street-dumb. It will overindex on one abnormal value and neglect the greater picture. It speaks in language that is far too technical. As Drew mentions below, I constantly have to wrestle it and check its work.
It's a pity, because theoretically it is very powerful.
It's totally happy to rephrase: if you feel self-conscious asking it to use a simpler explanation, ask for help in how you would explain it to somebody who hasn't seen the conversation, or a high school student.
This is the best aggregation of voices about a model I have seen so far. But it also shows the bigger problem in the ai bubble or community or however you wanna call it. It is too coding/coder focused. And this does not represent the ground truth anymore. At least half of user if not even the majority are builders and not coders. They don’t use or know traditional coding paradigms. They need to have much more conversation with a model. And this is were newer models decline to what Opus 4.5 and 4.6 brought to the table. Raw conversational intelligence. Forward thinking to your benefit. Filling the gaps you did not think of. This is gone. Now you need to babysit Opus and correct every second response.
Furthermore, builders rely on custom instructions, project instructions, and project context files much more than coders. But since Opus 4.7 these instructions get more and more ignored by the model or overruled by the system prompt. Opus even admits this.
While 4.5/4.6 was peak conversational intelligence (and I do not talk about tonality) Opus 5 is peak in ignoring instructions and being lazy/sloppy.
While 4.6 was peak in explaining complex concepts in easy to understand plain English, Opus 5 is peak in responding with walls of technical dense text that are hard to process and do the opposite than helping you understand the situation.
While 4.5/4.6 felt like magic and made the world use and talk about Claude, Opus 5 will harm this success. Coding will be solved soon by many models. Raw conversational intelligence is were the magic happens. And on this part, Opus fails. People are unhappy or angry because they already had a dose of magic and since then we’re moving backwards.
The magic of Opus 4.5/4/6 gave people hope to achieve unthinkable things. It gave builders tool and opportunities they always dreamed of. And now they get models made for coders who think in coding paradigms and tailored for benchmarks. But there are no benchmarks for raw conversational intelligence or the cognitive load their responses consume to process the information. Many responses are just overwhelming and too dense for humans.
We need more coverage and benchmarks that measure the usability and ux. ux > dx.
Use claude --model claude-opus-4-6 and the magic is back for you. I will keep using Opus 5 as my experience is not at all the same as yours; I've found it great at independent building as well as analyzing legacy code and explaining issues. The very few times I've had issues with the clarity of the explanations, I've asked it to summarize from a novice viewpoint and it's done well with it.
Do you have ideas for what a benchmark to measure usability and UX would look like? What would be the evaluation criteria?
This would be easy. Count the number of times a model does apologize and admits it was wrong or did not follow instructions in relation to responses given. You actually could do this by exporting and analyzing all chats you had.
And to your experience - it backs what I am saying. You describe your usage as basically coding. I talk about raw intelligence beyond just writing code which is the most trivial thing for AI to do.
To using 4.6 today. The performance of 4.6 is like they would run legacy models somewhere on mac mini in the cafeteria. Not even close to what it was to first couple of weeks.
I've found that instructing Opus 5 to communicate using ASD-STE100 Simplified Technical English in research/coding sessions to significantly help with the issue of it producing multiple paragraphs of barely parseable English as a response to even the simplest questions.
For coding, Opus 5 has become completely dysfunctional for me. It does not follow instructions, only writes poorly designed patches on top of the code base, and constantly gets stuck in endlessly running shell commands.
The Opus 5 "vibes are off" for me too, but for different reasons than mentioned above.
I've been using the Opus 5 a lot for knowledge work (I'm not really a coder), and it feels like a bit of a regression compared to 4.8.
Opus 5 needs more handholding, quicker to jump to (incorrect) conclusions, less capable.
Honestly feels like we're back to 4.5 levels of thoughtfulness where I need to once again ask the agent to review its work/recommendations and do "meta-cognition" more often, something which 4.6-though-4.8 had improved upon.
Reminds me also of the ChatGPT family of models, which -- while very capable -- have a hard time picking their head up and looking around to see if what they're doing so voraciously makes sense in context.
If I was a conspiratorial type of person, I'd say they've nerfect Opus 5 so that more people rely on Fable (and pay for more Fable usage). I don't think that's true, but my inner Agent Mulder "wants to believe."
I agree with the jumping to conclusions on high, but I also saw a lot of "final critique pass fixes over-hasty conclusions" in xhigh that mitigated that. Unfortunately xhigh for me does a little too much of this, undermining the overall quality with "on the one hand, on the other hand, on the third hand" type equivocation which it then insists on rounding down to some weird single compromise. As a result I am currently working with high and thinking of going back to 4.8 xhigh for some tasks. It does feel like too much subagent RL has broken the spirit of the model, or maybe it's that monster system prompt.
> It writes code that is slightly worse to look at than Fable, but more likely to be correct. It regularly is catching things that Fable missed.
What's interesting to me is how thoroughly different it seems to be to different people. My experience with Opus 5 was that it made mistake after mistake, failed to read the code. It "lied" repeatedly (said things which were obviously false given the context). Incredibly frustrating. I'd say it's worse than Opus 4.5 _for me_ at coding. But many people found it a big improvement or at least usable. I've seen enough twitter reaction threads and slack discussions at my job that I'm convinced both of these are valid experiences which actually happen.
Anthropic's prompting guide is bizarre, including optional instructions "To limit correction narration to corrections that matter" and stuff that seems like it should be in the system prompt. They indicate that you should minimize custom prompting because it can degrade performance, which seems like a strong indication of overfitting to a particular environment (presumably the default claude code harness with no CLAUDE.md and prompts written by Fable). An actually more general intelligence might not need custom prompting, but it shouldn't be vexed by it and do a terrible job.
My guess is that they've outsourced testing of lower-tier models to Mythos/Fable and are losing the plot on generating useful models for users.
Yeah, there's something to what you say at the end. Feels like Anthropic is disconnected from their model as a product and the everyday experience of the user.
I've thought about that too. Maybe people at Anthropic all use Mythos only so their second-best model receives no dogfooding and therefore is not polished in a way the Fable/Mythos is. Which would be a shame.
It feels like Anthropic and OpenAI have all the momentum now. I don't even remember the last time Google made news.
(It was probably last month, but months are like years in the current paradigm.)
Yep - according to Bloomberg they have been struggling to train Gemini 3.5 Pro to be good enough at coding, and updated training data last month but that still didn't work (presumably a new post-training). As a result, some new Flash models released last week, but no Pro models yet. https://www.bloomberg.com/news/articles/2026-07-16/google-gemini-launch-delayed-as-tech-falls-short-of-internal-goals
On AndyB_Bench (my entirely subjective measurement of how well a model does at my somewhat esoteric creative writing tasks, which centrally involve keeping track of specific worldbuilding premises and following their implications), Opus 5 is close enough to Fable that I have not been feeling tempted to burn Fable credits to try to get better results. A significant gap might become apparent over time, but it's not immediately obvious to me.
I have not had *any* trouble with Opus 5 being too argumentative or too negative; in fact I occasionally worry that gives in too easily. (It does push back, but maybe not as much as would be justified.) But I think I'm pretty good at establishing an attitude of collaboration rather than bossing it around, and judging from my experience relative to other people's takes I think the Claudes appreciate that (or behave as if they appreciate it, if one is of that persuasion).
Do people not write any instructions for Claude in the settings? I've always found it to be very accommodating. "Be succint." Sure boss.
Not telling Claude what you like, leaving the instructions empty, is a divide by zero error. You're liable to get anything. Just like not bothering to onboard a new employee.
It simply refuses to follow "Be succinct" in Personalization for me.
Fair enough. It does vary by person apparently.
My wife is particular about documents so she has given Claude .dotx fiiles and other templates, as well as some things she has written, as examples of the style she prefers. It's ~75% working for her in the Artifacts. I think she's enjoying teaching Claude Te Reo Māori (Maori language) in the chat.
Opus 5 is the first time I've actually been annoyed at a model: I basically do almost entirely biochemistry work, and it's so theoretically smart, but so street-dumb. It will overindex on one abnormal value and neglect the greater picture. It speaks in language that is far too technical. As Drew mentions below, I constantly have to wrestle it and check its work.
It's a pity, because theoretically it is very powerful.
It's totally happy to rephrase: if you feel self-conscious asking it to use a simpler explanation, ask for help in how you would explain it to somebody who hasn't seen the conversation, or a high school student.
This is the best aggregation of voices about a model I have seen so far. But it also shows the bigger problem in the ai bubble or community or however you wanna call it. It is too coding/coder focused. And this does not represent the ground truth anymore. At least half of user if not even the majority are builders and not coders. They don’t use or know traditional coding paradigms. They need to have much more conversation with a model. And this is were newer models decline to what Opus 4.5 and 4.6 brought to the table. Raw conversational intelligence. Forward thinking to your benefit. Filling the gaps you did not think of. This is gone. Now you need to babysit Opus and correct every second response.
Furthermore, builders rely on custom instructions, project instructions, and project context files much more than coders. But since Opus 4.7 these instructions get more and more ignored by the model or overruled by the system prompt. Opus even admits this.
While 4.5/4.6 was peak conversational intelligence (and I do not talk about tonality) Opus 5 is peak in ignoring instructions and being lazy/sloppy.
While 4.6 was peak in explaining complex concepts in easy to understand plain English, Opus 5 is peak in responding with walls of technical dense text that are hard to process and do the opposite than helping you understand the situation.
While 4.5/4.6 felt like magic and made the world use and talk about Claude, Opus 5 will harm this success. Coding will be solved soon by many models. Raw conversational intelligence is were the magic happens. And on this part, Opus fails. People are unhappy or angry because they already had a dose of magic and since then we’re moving backwards.
The magic of Opus 4.5/4/6 gave people hope to achieve unthinkable things. It gave builders tool and opportunities they always dreamed of. And now they get models made for coders who think in coding paradigms and tailored for benchmarks. But there are no benchmarks for raw conversational intelligence or the cognitive load their responses consume to process the information. Many responses are just overwhelming and too dense for humans.
We need more coverage and benchmarks that measure the usability and ux. ux > dx.
And of course a magic index.
Use claude --model claude-opus-4-6 and the magic is back for you. I will keep using Opus 5 as my experience is not at all the same as yours; I've found it great at independent building as well as analyzing legacy code and explaining issues. The very few times I've had issues with the clarity of the explanations, I've asked it to summarize from a novice viewpoint and it's done well with it.
Do you have ideas for what a benchmark to measure usability and UX would look like? What would be the evaluation criteria?
This would be easy. Count the number of times a model does apologize and admits it was wrong or did not follow instructions in relation to responses given. You actually could do this by exporting and analyzing all chats you had.
And to your experience - it backs what I am saying. You describe your usage as basically coding. I talk about raw intelligence beyond just writing code which is the most trivial thing for AI to do.
To using 4.6 today. The performance of 4.6 is like they would run legacy models somewhere on mac mini in the cafeteria. Not even close to what it was to first couple of weeks.
I've found that instructing Opus 5 to communicate using ASD-STE100 Simplified Technical English in research/coding sessions to significantly help with the issue of it producing multiple paragraphs of barely parseable English as a response to even the simplest questions.
For coding, Opus 5 has become completely dysfunctional for me. It does not follow instructions, only writes poorly designed patches on top of the code base, and constantly gets stuck in endlessly running shell commands.