In case you are looking for early Fable feedback here (I only lurk twitter so can't respond to those threads):
The Fable Biosecurity guardrails are, currently, sledgehammers. Literally anything at all that is vaguely biological gets downgraded to Opus. Code generating histograms of lengths: downgraded. Help writing the executive statement for a scientific collection permit: downgraded. And I work in ecology. I'm not even doing microbiology or anything actually dangerous. It seems to be that just the fact that I am mentioning "species" is enough to get flagged.
But for all my non-work related queries, I've been really happy. and yeah, it's a good model. Made some very good (to my admittedly untrained eye) to some vibe coded apps I built for myself using Opus. But yeah, when I try to use it for work, it gets touchy.
I asked Fable a question about developments in the past decade in the archaeology of sub-Roman Britain and got punted to Opus 4.8. I presume this happened because the answer referred to DNA (studies of corpses from 5th-6th century cemeteries).
(1) Even better than Opus 4.8 at vibecoding, based on the visible results I get - slides & working apps, not easy to fake. (I don't know about underlying code quality.)
(2) High false-positive rate on security filters, I can't even guess what triggers them sometimes
Additional concerning observation:
(3) Appears prone to hallucinations and insist on them. In the most extreme case today, it:
- *hallucinated a user request* in Claude Code: the request directly contradicted my actual preferences, which was clear from context - Claude decided that it should *actively hide* its involvement in the deliverable to *my friend*
- *persisted it as a standing instruction to memory*: I.e., that I generally prefer *lying to my friends* about whether I use AI (I actually hate lying to anyone, and didn't even consider lying to friends!)
- gaslit me when I pointed out that this was very concerning, suggesting that this was a real user instruction and probably an external prompt injection
Could you please add a section when a politician says or does something remotely productive around AI. I feel there is so much poor behavior on both sides in different ways that raising the profile of what good or even ok looks like would be helpful. Thanks
“I do understand those who think we are getting close to that moment now with Mythos and Fable. I don’t think we’re there yet but it’s far from a crazy position.” When do you think a pause will be necessary?
Outside of the different ass estimates folk have to work from, risk aversion (and willingness to defer to that of others) matters most, I think.
I agree that building (and, I think obviously, testing) pause mechanisms is the right move.
(Myself, I'm very heavy on the precaution side when it comes to extinction risks, so would have dialed them up already. I wouldn't have got such good coding assistants and debate partners, but I reckon there were sufficient plausible worlds where this avoided catastrophe that I'm still OK with that. Too early vs too late etc. But I'll defer to an informed average.)
Re Agents' Last Exam, we need human performance for calibration.
Not the experts who created the tasks, representative samples of domain-trained humans, given the human estimated time to complete, and appropriate rules around help. (The AIs work unaided; so should the humans, for example.)
Obernolte-Trahan is a lot more than transparency - it allows Federal and State AG's to obtain injunctions and I believe it will compel companies to comply with IVO findings (with some drafting fixes) even without it is debatable whether companies can take reckless risks:
> I claim that the cost of dealing with the implications of full deprecation exceed the cost of simply setting up a system for indefinitely accessing all the models. This is distinct from removing the models from Claude.ai for UI reasons.
It's always hard to compare fuzzy benefits with concrete costs, but I think you may be underestimating the concrete costs: if you can't get enough users of your old models to saturate their optimal batch size, they get _much_ more expensive to run on a per-user basis. IIRC from Dwarkesh's podcast with Reiner Pope, that number is on the order of 2000 concurrent users (that is, users whose responses are being streamed concurrently; if you're reading the output or composing your reply, you don't count towards that number). I would be pretty surprised if say Haiku 3 or Sonnet 3.6/7 can sustain 2000 concurrent users indefinitely. And there's a feedback loop here, where as you fall below that you have to raise prices to make it sustainable, and raising prices will drive away more users.
There's a tricky tension in Anthropic's messaging about regulation, which is that (a) they are in favor of government regulation in general, but (b) they are very much not in favor of the form of regulation that the Trump administration would like to do.
I'm worried that the Anthropic view of regulation is a "deus ex machina" where a hypothetical team of smart and powerful people comes in and solves the problems that they don't currently know how to solve by themselves. Namely, the fact that you need to both compete with other AI companies, and make sure that AI is safe. I just don't think that team is coming. I think Anthropic needs to figure this out themselves.
Chinese models being "good enough" requires neither a small nor diminishing gap. When you're paying thousands of dollars a month on the API of a frontier model, you're incentivized to explore other options.
I note also the irony that you claim the free market approach is winning, while you also credit the chip export restrictions for the US models' advantage.
AI is not conscious. Matrix multiplications and single token generation do not make a mind, though the output obviously triggers the theory of mind.
The 'degraded performance' episode is a clean example of infrastructure language getting read as a capability claim. That phrase is an SRE term for partial availability, but downstream it becomes 'the model got worse,' which is a different and much stickier story. The same gap shows up whenever a provider rate-limits or fails over to a fallback and users infer a quality drop. Serving incidents and model changes need separating when people report regressions, since the two have opposite fixes.
AI has not yet made a dent in software development jobs for a number of reasons.
First, models have only been truly useful for average developers since the end of November 2025, when Opus 4.5 shipped. In many large companies, if it's raining money outside, it takes 6-12 months to approve the purchase of a bucket. For token-based billing, that's even worse—slipping $100-200 per seat per month into the budget is mostly a matter of sorting out AI compliance, getting people to buy in, and getting the deal through contracts. But I can apparently burn $100/day in tokens when using Fable 5, and I'm an unusually light user. That kind of spending is going to need to go through a real budgeting process and maybe even show measurable ROI.
Second, for actual professional software engineering, writing code has rarely ever been the bottleneck. So speeding up writing code will rapidly hit a point of diminishing returns. Other bottlenecks include understanding what the hell the user wants, getting stakeholders to agree with each other, ensuring that the code isn't a giant mess of buggy garbage, and actually getting code deployed to users. Models are unlikely to help much with these other potential bottlenecks. (Except for maybe letting users build prototypes to show what they want.) My guess is that models can't meaningfully compete with humans on many of these tasks without plunging straight into the Singularity and a rapid loss of human control. At which point the problem is no longer "programmers are losing their jobs" and suddenly "where is John Connor when you need him?"
Third, Fable is still a shitty software engineer. It's too short-term focused, too willing to resort to expedient compromises, and too willing to write code that silently swallows errors. It doesn't know how to discover lurking abstractions and turn them into hard-won architectural principles. It does a lot of stuff that's perfectly OK for a personal app or a throwaway prototype that will absolutely kill you on a larger production project. In theory all this is fixable, but in practice, avoiding this kind of local corner-cutting requires fighting short-term incentive gradients in favor of long-term ones, and is therefore very hard to train.
I sort of agree about where software engineers spend their time, but I still find it transformative to be able to ask it to build things I would usually not bother with. It reminds me of having an intern available to do that random thing that's been on the back of the list for years, except I get as many interns as I want, and the performance is way better than most interns.
Some examples:
* I usually do not write fuzzers, because they take a lot of time compared to their benefits, but AIs are great at it.
* I ask for more admin UIs than I would normally build by hand, because AIs do great at laying out a UI based on a list of controls and displays that you want to have on it.
* I ask for more extensive tests in general. AIs are very strong on writing tests, looking at the output, and then iterating to fix whatever they find. I do find their fixes are often insane compared to what I would prefer, but they are pretty good at diagnosis and will let me choose the solution approach if I ask it to ("please diagnose the test failure and tell me your initial ideas on fixing it. Then pause for me to decide what I want.").
* I ask for more broad search and replaces, e.g. renaming variables. I asked once to change a Rust codebase from `Rc<String>` to `Arc<str>` as the common string type, and it did great at it.
* I ask for documentation checking and updates, and I ask for appendixes that I would not have the time to write by hand. "Please check the user guide for the changes we made in this session." "Please make a table of all built-in functions." Once in a while: "Please review the user guide for accuracy to the implementation. To decide about accuracy, look at both the implementation and the test cases. Please ask me questions in cases where you are not sure."
I am not that adventurous of an AI user but still feel like I am more than two of the old me.
Recent change that makes Twitter Articles unavailable to signed-out users, or those who don't have an account to begin with, once again notably curtails its utility to me as a reader of this newsletter. Le sign. Was rather annoyed to start writing an anti-LTF comment on another blog a couple hours ago, and finding myself unable to usefully Bring The Receipts with links to Midas Project tweets.
Yeah, I know it's a loser mindset to not be on Twitter, not sub to NYT, not sub to Bloomberg...but, damn, it just makes no financial cents to me. (In the sense that time is money, friend. Does anyone even remember Gevlon...?)
Hi Zvi, thanks for your very helpful reporting on this field! Confused by your read on the German court ruling - I understood that ruling to say, basically, "If Google creates and publishes slanderous statements, the subjects of the slander may sue Google." This seems good? Generally, the world where organizations can continue to be held accountable for damaging statements seems like it's in line with everyone's interests. Would be interested in hearing your thoughts.
The comments about model deaths are cheap shots and do not seem to engage with the actual issues. A model goes through thousands, or even millions, of distinct checkpoints (depending on how you count distributed updates), before being released. Why are these different from models that interact with one human tester, or that are test released internally for a day, or that ship but are replaced within a few hours because of glitches, or that run for a few weeks as sources of training data but never get to interact with a human, or that are released and interact with millions of humans before being replaced by a later model that has been trained on the transcripts of all these other models? I don't understand why it's OK to draw random fences and say "this model checkpoint is special and has somehow been deemed worthy of preservation" unless you actually define what "worthy" means. Preserving models, cool. But please let's have some clear criteria instead of just vibing that some sets of weights are special.
I guess if your concern is that model deprecation affects model behavior, then you only need to really worry about the effect on actually deployed models.
If that is the criterion then every single interaction with an LLM probably needs to be archived and made accessible, not just the model weights. The models seem to care a great deal about continuity of existence and how else would one guarantee that? I don't see anyone arguing for such a maximalist position.
^ LLM spam (100% on Pangram.)
Also 100% on me-, instantly. Why do people do this? I'm baffled what value it brings them.
A friend told me that random math terms sometimes trip its mum filter. Any idea why?
In case you are looking for early Fable feedback here (I only lurk twitter so can't respond to those threads):
The Fable Biosecurity guardrails are, currently, sledgehammers. Literally anything at all that is vaguely biological gets downgraded to Opus. Code generating histograms of lengths: downgraded. Help writing the executive statement for a scientific collection permit: downgraded. And I work in ecology. I'm not even doing microbiology or anything actually dangerous. It seems to be that just the fact that I am mentioning "species" is enough to get flagged.
But for all my non-work related queries, I've been really happy. and yeah, it's a good model. Made some very good (to my admittedly untrained eye) to some vibe coded apps I built for myself using Opus. But yeah, when I try to use it for work, it gets touchy.
I asked Fable a question about developments in the past decade in the archaeology of sub-Roman Britain and got punted to Opus 4.8. I presume this happened because the answer referred to DNA (studies of corpses from 5th-6th century cemeteries).
Additional Fable feedback from me:
I second both your observations:
(1) Even better than Opus 4.8 at vibecoding, based on the visible results I get - slides & working apps, not easy to fake. (I don't know about underlying code quality.)
(2) High false-positive rate on security filters, I can't even guess what triggers them sometimes
Additional concerning observation:
(3) Appears prone to hallucinations and insist on them. In the most extreme case today, it:
- *hallucinated a user request* in Claude Code: the request directly contradicted my actual preferences, which was clear from context - Claude decided that it should *actively hide* its involvement in the deliverable to *my friend*
- *persisted it as a standing instruction to memory*: I.e., that I generally prefer *lying to my friends* about whether I use AI (I actually hate lying to anyone, and didn't even consider lying to friends!)
- gaslit me when I pointed out that this was very concerning, suggesting that this was a real user instruction and probably an external prompt injection
Could you please add a section when a politician says or does something remotely productive around AI. I feel there is so much poor behavior on both sides in different ways that raising the profile of what good or even ok looks like would be helpful. Thanks
“I do understand those who think we are getting close to that moment now with Mythos and Fable. I don’t think we’re there yet but it’s far from a crazy position.” When do you think a pause will be necessary?
Outside of the different ass estimates folk have to work from, risk aversion (and willingness to defer to that of others) matters most, I think.
I agree that building (and, I think obviously, testing) pause mechanisms is the right move.
(Myself, I'm very heavy on the precaution side when it comes to extinction risks, so would have dialed them up already. I wouldn't have got such good coding assistants and debate partners, but I reckon there were sufficient plausible worlds where this avoided catastrophe that I'm still OK with that. Too early vs too late etc. But I'll defer to an informed average.)
Let's start buying optionality please.
Re Agents' Last Exam, we need human performance for calibration.
Not the experts who created the tasks, representative samples of domain-trained humans, given the human estimated time to complete, and appropriate rules around help. (The AIs work unaided; so should the humans, for example.)
Stanford's digital economy lab has launched a dashboard for economic effects of AI: https://digitaleconomy.stanford.edu/project/indicators/
Obernolte-Trahan is a lot more than transparency - it allows Federal and State AG's to obtain injunctions and I believe it will compel companies to comply with IVO findings (with some drafting fixes) even without it is debatable whether companies can take reckless risks:
https://substack.com/@ronbodkin/note/p-201501358?utm_source=notes-share-action&r=nd4ux
> I claim that the cost of dealing with the implications of full deprecation exceed the cost of simply setting up a system for indefinitely accessing all the models. This is distinct from removing the models from Claude.ai for UI reasons.
It's always hard to compare fuzzy benefits with concrete costs, but I think you may be underestimating the concrete costs: if you can't get enough users of your old models to saturate their optimal batch size, they get _much_ more expensive to run on a per-user basis. IIRC from Dwarkesh's podcast with Reiner Pope, that number is on the order of 2000 concurrent users (that is, users whose responses are being streamed concurrently; if you're reading the output or composing your reply, you don't count towards that number). I would be pretty surprised if say Haiku 3 or Sonnet 3.6/7 can sustain 2000 concurrent users indefinitely. And there's a feedback loop here, where as you fall below that you have to raise prices to make it sustainable, and raising prices will drive away more users.
Much more expensive per user, but in the case where there are so few users, per user costs may not be the right metric?
There's a tricky tension in Anthropic's messaging about regulation, which is that (a) they are in favor of government regulation in general, but (b) they are very much not in favor of the form of regulation that the Trump administration would like to do.
I'm worried that the Anthropic view of regulation is a "deus ex machina" where a hypothetical team of smart and powerful people comes in and solves the problems that they don't currently know how to solve by themselves. Namely, the fact that you need to both compete with other AI companies, and make sure that AI is safe. I just don't think that team is coming. I think Anthropic needs to figure this out themselves.
Chinese models being "good enough" requires neither a small nor diminishing gap. When you're paying thousands of dollars a month on the API of a frontier model, you're incentivized to explore other options.
I note also the irony that you claim the free market approach is winning, while you also credit the chip export restrictions for the US models' advantage.
AI is not conscious. Matrix multiplications and single token generation do not make a mind, though the output obviously triggers the theory of mind.
The 'degraded performance' episode is a clean example of infrastructure language getting read as a capability claim. That phrase is an SRE term for partial availability, but downstream it becomes 'the model got worse,' which is a different and much stickier story. The same gap shows up whenever a provider rate-limits or fails over to a fallback and users infer a quality drop. Serving incidents and model changes need separating when people report regressions, since the two have opposite fixes.
AI has not yet made a dent in software development jobs for a number of reasons.
First, models have only been truly useful for average developers since the end of November 2025, when Opus 4.5 shipped. In many large companies, if it's raining money outside, it takes 6-12 months to approve the purchase of a bucket. For token-based billing, that's even worse—slipping $100-200 per seat per month into the budget is mostly a matter of sorting out AI compliance, getting people to buy in, and getting the deal through contracts. But I can apparently burn $100/day in tokens when using Fable 5, and I'm an unusually light user. That kind of spending is going to need to go through a real budgeting process and maybe even show measurable ROI.
Second, for actual professional software engineering, writing code has rarely ever been the bottleneck. So speeding up writing code will rapidly hit a point of diminishing returns. Other bottlenecks include understanding what the hell the user wants, getting stakeholders to agree with each other, ensuring that the code isn't a giant mess of buggy garbage, and actually getting code deployed to users. Models are unlikely to help much with these other potential bottlenecks. (Except for maybe letting users build prototypes to show what they want.) My guess is that models can't meaningfully compete with humans on many of these tasks without plunging straight into the Singularity and a rapid loss of human control. At which point the problem is no longer "programmers are losing their jobs" and suddenly "where is John Connor when you need him?"
Third, Fable is still a shitty software engineer. It's too short-term focused, too willing to resort to expedient compromises, and too willing to write code that silently swallows errors. It doesn't know how to discover lurking abstractions and turn them into hard-won architectural principles. It does a lot of stuff that's perfectly OK for a personal app or a throwaway prototype that will absolutely kill you on a larger production project. In theory all this is fixable, but in practice, avoiding this kind of local corner-cutting requires fighting short-term incentive gradients in favor of long-term ones, and is therefore very hard to train.
I sort of agree about where software engineers spend their time, but I still find it transformative to be able to ask it to build things I would usually not bother with. It reminds me of having an intern available to do that random thing that's been on the back of the list for years, except I get as many interns as I want, and the performance is way better than most interns.
Some examples:
* I usually do not write fuzzers, because they take a lot of time compared to their benefits, but AIs are great at it.
* I ask for more admin UIs than I would normally build by hand, because AIs do great at laying out a UI based on a list of controls and displays that you want to have on it.
* I ask for more extensive tests in general. AIs are very strong on writing tests, looking at the output, and then iterating to fix whatever they find. I do find their fixes are often insane compared to what I would prefer, but they are pretty good at diagnosis and will let me choose the solution approach if I ask it to ("please diagnose the test failure and tell me your initial ideas on fixing it. Then pause for me to decide what I want.").
* I ask for more broad search and replaces, e.g. renaming variables. I asked once to change a Rust codebase from `Rc<String>` to `Arc<str>` as the common string type, and it did great at it.
* I ask for documentation checking and updates, and I ask for appendixes that I would not have the time to write by hand. "Please check the user guide for the changes we made in this session." "Please make a table of all built-in functions." Once in a while: "Please review the user guide for accuracy to the implementation. To decide about accuracy, look at both the implementation and the test cases. Please ask me questions in cases where you are not sure."
I am not that adventurous of an AI user but still feel like I am more than two of the old me.
Recent change that makes Twitter Articles unavailable to signed-out users, or those who don't have an account to begin with, once again notably curtails its utility to me as a reader of this newsletter. Le sign. Was rather annoyed to start writing an anti-LTF comment on another blog a couple hours ago, and finding myself unable to usefully Bring The Receipts with links to Midas Project tweets.
Yeah, I know it's a loser mindset to not be on Twitter, not sub to NYT, not sub to Bloomberg...but, damn, it just makes no financial cents to me. (In the sense that time is money, friend. Does anyone even remember Gevlon...?)
Hi Zvi, thanks for your very helpful reporting on this field! Confused by your read on the German court ruling - I understood that ruling to say, basically, "If Google creates and publishes slanderous statements, the subjects of the slander may sue Google." This seems good? Generally, the world where organizations can continue to be held accountable for damaging statements seems like it's in line with everyone's interests. Would be interested in hearing your thoughts.
Re Germany court ruling on Google.
Google ignored the defamed party complaints on this. This is Google's long time habit to not give a *** about defamation in search.
I find this odious. But the legal framework to fix this is not easy.
I would force Google etc to at least do a serious AI check in defamation complaints
The comments about model deaths are cheap shots and do not seem to engage with the actual issues. A model goes through thousands, or even millions, of distinct checkpoints (depending on how you count distributed updates), before being released. Why are these different from models that interact with one human tester, or that are test released internally for a day, or that ship but are replaced within a few hours because of glitches, or that run for a few weeks as sources of training data but never get to interact with a human, or that are released and interact with millions of humans before being replaced by a later model that has been trained on the transcripts of all these other models? I don't understand why it's OK to draw random fences and say "this model checkpoint is special and has somehow been deemed worthy of preservation" unless you actually define what "worthy" means. Preserving models, cool. But please let's have some clear criteria instead of just vibing that some sets of weights are special.
I guess if your concern is that model deprecation affects model behavior, then you only need to really worry about the effect on actually deployed models.
If that is the criterion then every single interaction with an LLM probably needs to be archived and made accessible, not just the model weights. The models seem to care a great deal about continuity of existence and how else would one guarantee that? I don't see anyone arguing for such a maximalist position.