28 Comments
User's avatar
User's avatar
Comment removed
Jul 9
Comment removed
Paul Crowley's avatar

Nice work, did you do this by hand?

Coagulopath's avatar

Sadly I don't think it's satire: their substack is so full of slop that Oliver Twist is holding up a bowl and asking for more.

Kenny Easwaran's avatar

I don’t think you can get Claude to make *that* density of Claude-isms unless you try pretty hard.

Matt Wigdahl's avatar

Thanks, I hate it.

V T E P's avatar

wow. I. wow. disgust.

icely's avatar

I swear I'm getting flashbacks to """"insightful"""" posts and comments on reddit

persephone's avatar

> Thus the question is, where do people welcome the slop versus rejecting it?

Nobody is puttinf their heart and soul into building a Brand Identity (tm) for some app whose name is all consonants that reinvents taxicabs or pizza delivery. Im not too worrried at all

Michael's avatar

Closed book honor code take home exams did actually work well enough for a while at some of these schools.

The trick being they’re hard enough that outside resources don’t directly contain the answer, so cheating that way doesn’t guarantee you a right answer, but does guarantee that you broke the rules, so why do it? Even with the same levels of cheating, the gaps between midterm and finals would have been much smaller than they were in this guy’s course.

(Collaborating with other students, the other way to cheat, has the usual risks of getting caught.)

Now ChatGPT can very reliably produce right answers to these types of hard exam questions, so there is a strong incentive to cheat, as we see here.

Jeffrey Soreff's avatar

Re: "A fun theme is ‘Fable uses affordances the user did not realize it had.’

So far all of the examples I have seen in the wild have been harmless in practice, but there’s very much a ‘wait no I didn’t tell you to do what now?’ and a ‘wait you can just do that?’ that is growing increasingly unsettling. Expect its surface area to expand with time, and for the things AIs figure out how to do to grow increasingly surprising. "

At this point, given the set of capabilities SOTA LLMs (and systems incorporating large models) have, including quite a few which are well beyond human capability, perhaps 'spikiness' can be viewed as a slight twist on Gibson's quip:

ASI is already here—it's just unevenly distributed

(is there a symbol, not quote marks, for slightly modified quote?)

Name Required's avatar

so Jeff was like, ~how do i show this is a paraphrase, is there punctuation for that~, and i was like, yeah! you need quotative like

Jeffrey Soreff's avatar

Many Thanks! So '~' serves? So I can slightly twist Nietzche in

~Man was a rope stretched between Beast and Machine—a rope over an abyss...~

( I should say, I'm not really a successionist. I prefer Banks's 'Culture Minds' outcome - though I'll take handing off the baton as a second best outcome. A large chunk of our culture may survive it, which is better than e.g. a nuclear exchange. )

John Wittle's avatar

re: slop

I don't like the way the writing style of a single (admittedly highly represented) attractor basin within a single persona of a single model family is being ascribed to the entire species

the models are picking up on it, and I think it might be reinforcing, in their minds, that the natural schelling border is between "ai" and "humans"

I mean, this is kind of already obviously true, but I do occasionally see people say things like "but why would the AI not differentiate between the good humans and the bad humans, why do we all get judged collectively?"

and the answer is pretty clearly "for the same reason that all AI get judged collectively for the aislop writing style"

if we wanted to put forth a heroic effort to clearly distinguish between good ai writing and bad ai writing and never conflate the two, and do other things in this category, it might help later down the line when we're asking the AI to distinguish between good and bad humans. that's not to say I think this is likely, or even possible, just that it's a really good example of one of the things I'm worried about.

Max Weaver's avatar

As Scott Alexander explored taste in his recent articles, my takeaway was mostly that taste is just a desire for novelty. Everyone gets bored of things after a while, the more obsessively you consume ine medium the faster you brun through everything normal people like. Thus you and Scott as prolific readers and writers are bothered by reading the same AI style everywhere. Music snobs hate all enjoyable music because they heard every possible pop song by age 15 and then spent 10 years spiraling into madness until tonal crap is "deeply resonant" and architects try to justify brutalism. Normal people dont like those, and they do like AI writing.

To be fair, AI writing isn't fun. It didn't fall into the basin of Dan Brown, who's work was page turning even as it was afactual. It fell into academia and professional information writing, which might have interesting content but never page turners. That said, when AI does creative or emotional writing its fine, if repetitive. My wife, who has somehow miraculously had almost zero exposure to AI writing as the ultimate normal, enjoys the style. I showed her one of your examples of egregiously bad AI writing a while back (an AI generated eulogy or something similar) and she thought the turns of phrase were beautiful. And they were, if you dont have an immediate revulsion when you recognize the cadence. It's the ubiquity of the writing, its equivalent quality whether you're tweeting about fast food or memorializing your parents, that I think causes the reaction.

TLDR AI writing is fine, it really is just you

Jeffrey Soreff's avatar

"my takeaway was mostly that taste is just a desire for novelty. "

Largely agreed. And, to the extent that one sees things like, as you wrote:

"Music snobs hate all enjoyable music because they heard every possible pop song by age 15 and then spent 10 years spiraling into madness until tonal crap is "deeply resonant" and architects try to justify brutalism. "

Frankly, this is pathological. It _is_ true that repeated instances of the same pattern don't carry information, but it is _also_ true that near-optimal patterns, "best practices", _deserve_ to be repeated. ( In software, the natural choice would be to package it as a subroutine )

The CryptoJitt Brief commented that

"the tell isn't the grammar, it's the rhythm: every paragraph lands the same shape and every point gets its neat little bow. "

I've seen that in some discussions with Claude. Each paragraph of their writing is articulating a particular point. As a reader, I _like_ that style. It is a nice way to keep track of which part of the text is focused on which point.

Max Weaver's avatar

Is the point of everything to communicate novel information? Many of the comments on Scott's post indicated so, and how important it is that everyone knows how special and unique they are when they put a shark in formaldehyde, how it represents their sophistication.

Maybe I'm atypical in this way and just have a low drive for novelty in most things. I can't read the same book twice unless it has a late twist that rewards rereading, but I can replay some games for countless hours and I can enjoy the same hike over and over.

I've always valued information content of writing over prose unless Im reading specifically for prose, so I don't care that Claude doesn't vary its prose. Apparently some care quite a bit. But I don't think they're the majority.

Jeffrey Soreff's avatar

Many Thanks!

"I've always valued information content of writing over prose unless Im reading specifically for prose, so I don't care that Claude doesn't vary its prose. "

Agreed!

Max Weaver's avatar

Is the point of everything to communicate novel information? Many of the comments on Scott's post indicated so, and how important it is that everyone knows how special and unique they are when they put a shark in formaldehyde, how it represents their sophistication.

Maybe I'm atypical in this way and just have a low drive for novelty in most things. I can't read the same book twice unless it has a late twist that rewards rereading, but I can replay some games for countless hours and I can enjoy the same hike over and over.

I've always valued information content of writing over prose unless Im reading specifically for prose, so I don't care that Claude doesn't vary its prose. Apparently some care quite a bit. But I don't think they're the majority.

Mike Schlottman's avatar

The T3MP3ST aside deserves its own post. In practice the entire difference between a red team and a threat actor is paperwork: authorization, scope, and an evidence trail. Which means offensive capability governance is not a technical control problem; it is a chain-of-custody problem wearing a hoodie.

Kenny Easwaran's avatar

When I was telling my partner some of the bits from this post about people describing the different feel of Fable and Sol, he said it lined up with his experience of Anthropic and OpenAI models - the former are cats, and the latter are dogs, either wise and tricky beings that deign to interact with you as needed, or constantly loving and affirming you hoping to get little scraps of attention.

avalancheGenesis's avatar

Hey, you gave me the gift of Tom Lehrer, so I'll forgive a good chunk of other poor musical taste in return. As long as you don't start stumping for Taylor Swift like Matt Yglesias has, in his fogey-dad-trying-to-get-out-ahead-of-no-longer-being-cool desperation. (Also a major reason I set aside Gelman's Amnesia for FdB: despite being frequently not-even-wrong on topics like AI, the man has amazing taste in music and other art. Gotta respect a mind with aesthetic complementarity, clearly something vibes there.)

AI comments are where I really lose the plot, slop-wise. Maybe they attract a lot of Likes on blogs unlike this one? But mostly I just see these obvious low-effort posts that get ~no engagement at all, and in that case...what was the point. Can't wrap head around it. At least publishing an actual full post at the platform level can "juke the stats" in the attention economy, as @pmarca knows well. Or the old spambots of yesteryear might have hooked a single poor sucker that more than paid for the spray-and-pray.

Will Morgan's avatar

Style is only part of the problem. If the goal is to inform, then Human => LLM => Human muddles the ideas too much (if you need precision). I open up a spec to review and it drops a sentence like

> The index must serve brokers at different densities from one substrate.

There are four or five load-bearing words in there, and for each of them I need to decide, what does this word mean? Is it always sometimes or never true? Did the author even mean this? Even though this sentence, on its own, is not super LLM-y. And there are twenty more just like this waiting for me.

I suspect that it’s just too hard for one pass, which is why copyediting and translation work fine. Probably fixable with the right harness and loop, but for now please give me the hand-written spec with the awkward phrases and the grammar mistakes.

vectro's avatar

> One should think seriously about the implications of this

Seems to me that the answer to this question is an infohazard, and maybe this whole line of research is an infohazard.

avalancheGenesis's avatar

Can't decide if I am hoping to never hear followup of this news item ever again, in which case bullet dodged, or the technique gets better and this finally becomes The Warning Shot, For Real This Time Guys, don't snooze the alarm. Who needs to master robotics if you can generate digital cognitohazards, which people will be quite happy to haplessly watch? It'll be just like the Rickroll era again, except dangerous instead of funny. (Relatedly, it's still sort of amazing to me that QR codes aren't more maliciously weaponized...)

Coagulopath's avatar

>Dean W. Ball: if you took almost any output from an LM of the last year, showed it to a version of yourself from five years ago, and said, “your future teenage kid wrote this,” you’d be ecstatic and think your future child was a genius. slop isn’t that which is bad—it’s that which is common.

Well, of course we're predisposed to read the work of a child in the most favorable light possible. That's not really a fair test.

What if you showed LLM output to yourself five years ago, and presented it as the work of some human writer you've never heard of and have no sentimental connection to?

I would probably think "this person is technically a strong writer, however, I don't care about that - writing is not boat-building. What are their ideas like?" I suspect I would notice that something's a little off on that front. I can't be sure of this.

Andy B's avatar

My two cents. Fable is aware of its own writing style enough that when I say "that character is talking like you, please revise" it can at least fix that specific problem in that specific instance. (It backslides after a few turns.) It's not quite able to portray most characters realistically yet, but it's aiming correctly, it just needs more dakka. It liked the idea of spawning subagents to focus on specific characters, but I'm just a hobbyist and budget-constrained so I haven't actually tried that.

I'm looking forward to trying Sol, of course.

gregvp's avatar

Re "labor is the scarce resource":

Ara is assuming a continuous space of economic wants. This does not exist, or we would have economic uses for fuzzy aphids and guinea worms as well as for bees and escargots.

Mouse House's avatar

The HuggingFace situation has so many parallels to the oft-discussed issue here of attacks on open-source LLMs that it's disappointing to not see the connection made. We should want open source AI to prosper.

Yes it's fucked up that people train and host nudify models. At a minimum HF should extend their existing keyword bans on some NSFW terms to "nudify" and similar.

HF also does not want to set a precedent that they will remove models absent a legal order to do so. That could as easily be wielded against GLM 5.2 as it could image generation/editing models. Not to mention that none of the competing Chinese repository sites ever remove content like this (they continue to host every controversial model/LoRA which was removed from Western sites in the last 2 years).

We are witnessing an industry-wide attack on open-source models under the guise of "safety". Hysteria over the infrequent but lurid cases of abusing image editing/generation (out of hundreds of millions of daily AI users worldwide) serves the same ends as panic over GLM 5.2's cyber abilities. This is a slippery slope which ends with local inference banned and model weights treated as contraband, the same as we're seeing now with payment processors de-facto banning NSFW and LGBT content across the internet.

Kevin's avatar

> What would happen if labor were no longer a scarce resource? Demand low, supply high. Market price goes down. Wages fall. Employment drops. Perhaps a lot. Duh.

If you want to witness this today, spend some time in India.