12 Comments
User's avatar
Mike Lambert's avatar

“AI writing remains, I believe, highly detectable by both man and machine if you care” Is this a statement about RLHFed models or base models? I agree with current RLHF that have mode collapsed into certain tells, but I think I wouldn’t succeed with properly prompted frontier base models

Mike's avatar

"The ‘why does Josh Whiton always grab the same three books at the library’ puzzle, Gemini 3 wins, Opus 4.5 and GPT-5.1 lose, and Grok has issues (and loses)."

If you see a stranger taking a series of inscrutable actions that involve being alone in a windowless room for several hours, the odds of an illicit explanation are far above 0.

The models don't know that this 'riddle' is derived from Josh's own behaviour, and despite the word 'riddle' it's phrased as a real observation of a third party. So, they don't know to filter the explanations to that of a Sensible Person who does normal things.

I tried this with Gemini and suggested it think of 'unconventional' answers, and prodded it a bit. It eventually suggested the books could be a 'privacy screen' to hide legal or medical documents from a camera. 'NSFW' answers are just missing from its hypothesis space. Grok is too eager, but correct to consisder them.

Kevin Lacker's avatar

"Legally, OpenAI is required to act like a nonprofit, not like a for-profit entity!"

<monkey paw curls>

Scott Novak's avatar

“The cycle of language model releases is, one at least hopes, now complete”.

Spoke a few hours too soon. I wonder if its powerful enough to restore the Gemini Ultra relative value proposition for non image/video creators.

https://blog.google/products/gemini/gemini-3-deep-think/

Coagulopath's avatar

>Tyler Austin Harper: I wrote about “The Will Stancil Show,” arguably the first online series created with the help of AI. Its animation is solid, a few of the jokes are funny, and it has piled up millions of views on Twitter. The show is also—quite literally—Nazi propaganda. And may be the future.

Yes, but what's a view on Twitter? Someone just scrolling past? The tweet by Tyler Austin Harper itself has 1.1M views, to give you some idea.

The episodes uploaded on Youtube seem to have a few thousand views each. And this is from someone with a big established brand (Emily Youcis has been an animator ever since the Newgrounds days).

If this counts, I think an even better example would be the AI slop infesting Youtube Shorts (I've seen multiple videos with hundreds of millions of views).

TK-421 Presents's avatar

Opus 4.5 truly is a beast. I've been using GPT flavored products exclusively since 3 because I'm a heinous nerd.

4.5 wrote the prompts used by Veo to generate all of the footage in these videos:

- Art history: https://www.youtube.com/watch?v=i0JrjdDd6Bw

- MeToo, a short film: https://www.youtube.com/watch?v=e7lpACBZEig

Kevin Mulder's avatar

For the perspective that lawyers don't use AI due to lawyer- or law firm-specific problems (skill, inertia, etc.), what's the economics rationale for that remaining a viable strategy much longer? If some lawyers do use it and the gains are real, won't they become (much) more competitive and/or effective? It seems to be that the AI quality from here forward is roughly irrelevant if the current models work for a fraction of the industry in any meaningful capacity

[insert here] delenda est's avatar

POV: former lawyer using various models for French, UK and Swiss law:

The strategy It is not viable, largely for the reasons you state.

Niche areas, especially a lot of much less extensively and publicly documented foreign law, will hold out a bit longer, a couple more generations at least and possibly even until there is a new research paradigm. But even there the strategy is not viable.

Matt Lubin's avatar

> "even to someone who (to my eyes"

I assume that's a typo, and the four words before the parenthesis should be deleted.

I understand the love for Opus 4.5 but in my experience it refuses to answer totally innocuous questions about microbiology/virology, telling me to use an older model. I can set up a (too long) custom prompt/project that helps convince Opus 4.5 that I'm not a terrorist, but I haven't found a perfect one and this is deeply frustrating

Mark Schröder's avatar

Wrt the multiagent thing: If we assume that we stay in the regime of training runs and inference instances we should assume that „coordination“ and scaffolding of various kind becomes increasingly important bc you can always run many instances of the smartest ai. And if running 1000x as many instances gets you even a little better performance it’s worth it at least for doing ai research so you get the most out of them until the next training run finishes!

Jack Koch's avatar

"Louder and once more for the people in the back: Evan Hubinger of Anthropic reminds as that Alignment remains a hard, unsolved problem, even to someone who (to my eyes, and even more so to the eyes of Eliezer Yudkowsky as seen inevitably in the comments)."

incomplete sentence—curious to hear the rest!