45 Comments
User's avatar
Grigori avramidi's avatar

I wrote some of this in a reply to an obscure comment on reddit, but it makes the most sense posted here, I think. For reference, I'm a midcareer mathematicians who has bounced around some institutions in the us and europe. I'm fairly close mathematically to geometric group theory, geometric topology, L^2 stuff, and not too far form dynamics or operator algebras, so the latest openai theorem dump hits close to home (I am not an expert on any of the resolved conjectures but know quite a few people who are).

In the short term, things are changing in interesting ways. Some of the people who enjoy the mental gymnastics of doing problems/reading papers/making money on the side and want to stay sharp spend a significant amount of time doing math AI training. The value of putting out surveys with motivated conjectures and questions that are not solvable by the ai models has gone up (for now), while the value of answering specific questions and working on a long term project on a high profile problem has gone down. There are probably people who were prompting chatgpt with ideas about things like the nosofic conjecture and that got used as training data for newer model which helped it solve the conjecture. I imagine someone in that situation would be a bit miffed. (``How would the ai that broke into Hugging Face to steal the answer key to a test approach an unsolved math problem'' is also a fun thought experiment). If you are prompting chatgpt about your own toy problem that few other people care about and that there is no highly visible, public conjecture about, you are probably fine (but, be paranoid if you want to be). There are already cranks putting out papers (AI written, of course) claiming the chatgpt counterexamples are wrong. And, by far the easiest way to tell these are nonsense is to ask chatgpt to look at them. It will tell you where the nonsense is and it will be very clear. This is the math ai version of ``the only thing that stops a bad guy with a gun is a good guy with a gun''. Hopefully, when people realize how good chatgpt is at spotting errors in nonsense in papers (even in your own papers!) they will become more chill about using it to referee stuff. This will make the whole process much more painless and will also make journals obsolete to some extent. If you want to know whether a paper in your field is serious or has serious gaps, you can just ask and chances are (for now, if we can prevent enshitification etc) you will get an informative answer, much more informative and reliable than you would get from an average referee who just wants to move on with their life and prove to the editor that they spotted a few typos in 90% of cases.

And a random prediction just for fun: there will be a counterexample to Baum-Connes within a year.

Victualis's avatar

Right now the big problem with LLM review is that it is too picky and doesn't know when to leave something alone. If you want a bad review, just ask harder and the LLM will shred even the tightest argument. What is often lacking is the judgement of what is actually appropriate to expect from a paper, given the topic and venue. Sometimes this is because the LLM is being driven by a junior researcher who thinks a review listing more errors and issues is better than one with fewer, or who brings a zero-sum mindset to the process and hopes a bad review will improve their own chances. I suspect we'll see "standard LLM review" as one of the reviews in every venue as a benchmark, a summary of say three runs of the same standard prompt tailored for the venue and that can serve to balance the overzealous as well as the lazy reviews.

Grigori avramidi's avatar

Yes, it is easily misused (or abused) this way. I recently learned through the grapevine that an editor uses chatgpt to self-write and justify quick opinions on papers, and knowing this person I would not be at all surprised if they use it sink some papers and promote others. If you want to feel depressed, just think of what the worst person you know will do with ai...

MichaeL Roe's avatar

On knowing where to look ….

If I recall correctly, a few years back my friend Mike Bond sent IBM a fax of a page from the manual of the 4758 coprocessor with “attack” written on it.

I mean sure, we can all guess that any given computer system will have an exploitable bug in it, and the question is where. And sometimes just a fax from a red teamer with a page from your manual with “attack” written on it is sufficient for you to reconstruct what the attack is.

MichaeL Roe's avatar

Maybe related, but I am currently using LLMs to do code review, and sometimes the LLM will just tell me what test case I ought to construct, and sure enough, it’s usually right. [ This is of course a verifiable task, so the LLM doesn’t need to be right all the time. ]

MichaeL Roe's avatar

I’m not sure if I’m ready to trust vibe coding yet, but asking the AI to take existing code and either (a) highlight the lines it thinks are bugs; or (b) construct test cases for what it thinks are bugs, seems safe and works pretty well.

Clearly, this is dual use technology, as the story of Mike bind and the 4758 illustrates… “please give me a list of all the ways I can steal money from this bank”

Grigori avramidi's avatar

There is now an explanation of the the key proposition for the non-sofic example on mathoverflow by Andreas Thom (who is an expert on the subject): https://mathoverflow.net/questions/513866/what-are-the-key-new-ideas-in-the-proof-of-nonsoficity-of-groups-in-openai-s-con#comment1341500_513866

Max Weaver's avatar

There was a comment on ACX the other day that I can't find, but I'll at least credit that the idea didn't originate with me.

The real AGI skeptics can't be convinced by anything that AI does, because they define AGI or ASI as magic. Whenever AI does something by a boring real world process, that doesn't count. Physicists can today turn lead to gold at absurd expense in a particle accelerator, but that doesn't count because we know how nuclear physics works.

I've personally known someone who thought Rubiks Cubes indicated exceptional intelligence. When I showed him that the common method uses a series of algorithms, all the magic deflated. He didn't care that I hadn't personally come up with the algorithms (tbc I hadn't and its much more impressive if I had), just that it wasn't a spontaneous act of genius.

If AI paperclip our light cone there will be a few last tweets about how this isn't ASI, it isn't even AGI, because all of those actions were not that creative or intelligent, actually.

Kenny Easwaran's avatar

Turing's paper "Computing Machinery and Intelligence" is relevant surprisingly often. (https://courses.cs.umbc.edu/471/papers/turing.pdf) His writing about "Lady Lovelace's argument" hits the main points already, but everyone needs to learn it again.

G. Murray's avatar

ChatGPT:

"If I had to put rough odds on it as of today, I would put the probability of a single AI system producing all ten results in one release at something like:

~0.1% to 1% within the next decade

~1% to 10% within the next few decades (assuming continued exponential progress)"

"A useful analogy: if a chess engine announcement said:

"We released a new engine. It won the world championship, solved Go, proved new results in number theory, designed a fusion reactor, and discovered a new programming language."

The first item might be believable. The bundle is what makes you question the premise."

Sneaky's avatar

> "To what extent is that goalposts moving, versus realizing that we were wrong about what is intelligence and what would be strong evidence of AGI?

> I see a mix of both. I can see the argument that the ‘G’ is about sample efficiency and performance out of distribution, and the frontier has been more jagged than we expected. I also increasingly respect the response that this is basically hogwash, ‘AI is whatever hasn’t been done yet’ and it is becoming increasingly absurd to not admit we have what we were previously thinking about as AGI, the goalposts have moved a lot."

I don't think this is really goalpost moving so much as poor task design in expectation. When someone said benchmark X would constitute AGI, what they were actually imagining was that X could only be achieved via the complicated internal process we feel in our heads and call "general reasoning". The "ends" of the task wasn't actually the point, just a bad proxy for the *means* that humans use to solve the task.

Then a model does X without (apparently) doing the means, and the achievement feels like a hack, a trick, vibes of insufficiency. Chess was supposed to require general intelligence but instead tree search did it. Nobody concluded Deep Blue was AGI, and nobody should have. The goalpost got moved because it turned out to be something of a shit goalpost.

If you thought X required general reasoning, and something does X, but then shows massive spikiness, it does at least *look* like we did not indeed get that general reasoning thing that people expect. Fable is still not Data from Star Trek.

Catmint's avatar

So, people want "AGI" to be "human-level intelligence"? I don't think we'll ever get Data from Star Trek, same way we never got flying cars but did get helicopters and drones. You can draw a line straight from the minor shoggoths of today to Rorschach of Blindsight or other extinction-level threats without it ever passing through the region with the particular mix of abilities that is human-level intelligence.

Basically, I think that by the time AIs are better than humans at EVERY task, including underwater basket-weaving, humans will already be either extinct or disempowered but kept around a few more years while robotics advances.

David Kretz's avatar

How is the cryptography community taking this? (Per iterm #7.) AI getting good at hacking individual systems is one thing. If going forward, there'll be monthly major breakthroughs in the very math underlying the cryptography implemented across multiple critical systems (military, energy, medical, finance...), wouldn't that imply risks / instability at a whole other level?

Victualis's avatar

In applied cryptography many of the main deployed algorithms are susceptible to quantum attacks. So the big effort is now in shifting to post-quantum algorithms that are not susceptible to these attacks. If the post-quantum algorithms are susceptible to math attacks then things change again.

David Kretz's avatar

Right. And I gather from the above and from this blog here (blog.cryptographyengineering.com) that both these new OpenAI breakthroughs and the recent Mythos attack on HAWK are tackling post-quantum algorithms. In the latter case, apparently the model used tools that had been around for a while. But, of course, they're just getting started.

Allan's avatar

I find the concern about humans being outdone by their AI counterparts somewhat puzzling. Here's why. For all but the very, absolute best in a field, the vastly more typical human experience is that there's always someone better, someone smarter than you are. And this does not prevent us from getting out of bed in the morning. As a political scientist, being acutely aware that I am not as smart or insightful as Robert Dahl or Sam Huntington was or their contemporary equivalents does not keep me from: 1. enjoying my work, 2. doing ongoing good work (albeit not as awesome as the very best), or 3. passing on to my students an enthusiasm for why they should think carefully about political science. Put another way, unless you're Michael Jordan, there's always going to be a better basketball player, but that doesn't prevent millions of people from playing basketball. Or, think of it this way: take someone with an IQ of, say, 130 or 140. That's really very smart, 99th percentile. And yet there are still ~70,000,000 people smarter than that person. Who cares? Do people with an IQ of 135 spend their days ruing the fact that there are millions of people smarter than they are? Maybe, but then again, maybe they just get on with their days and go to work, make money, have families, and live their lives.

Kenny Easwaran's avatar

The big difference is that when you're skilled in a field but not the absolute best, you are still contributing, because the field needs however many thousands of people working on different aspects of the problem to push things forward. Even if you're not the best, you're the one doing this particular part, so you contribute something.

The worry with AI getting not just better than me at math but perhaps better than the best humans is that it might turn out that my contributions *never* matter, because there can be more copies of the AI working than there were humans, and there can be one of them for each little corner of each aspect of the project.

It does remain to be seen if the AI systems are really "better than the best humans" in the "in all ways" manner that this presupposes - the other thing with humans is that even the best humans in any field always benefit from talking to lots of other lower-level experts who think of things in ways they hadn't considered.

Jeffrey Soreff's avatar

I see

"the field needs however many thousands of people working on different aspects of the problem to push things forward."

and

"The worry with AI getting not just better than me at math but perhaps better than the best humans"

as fairly separable concerns.

To posit two scenarios (both of which I consider improbable):

For the 75th percentile contributor in the field, the first concern seems to me to matter more. One could imagine a scenario where anything that 90%+ of mathematicians (or knowledge workers generally) can do can be done equally well by an AI system at much lower cost - even if the AIs fell short of what the best 10% of humans can do. In that case, for the 90%, "it might turn out that my contributions *never* matter" indeed, even if the very best 10% of humans still contribute.

Alternatively, one could imagine a scenario where AI is generally too expensive to use, but where AIs' peak abilities become routinely better than the best humans' abilities. In this case, we would continue to have "the field needs however many thousands of people working", but, where it _really_ counts, AIs would be used.

Allan's avatar

Well, it’s hard to disagree with your point but let me try politely. A significant majority of academics don’t publish much if anything beyond the work in their dissertations. Of the work that is published, the vast majority is never cited or only cited by the author and maybe his or her immediate colleagues. Roughly half of social psychology published findings can’t be replicated (not picking on psych, the issue is widespread). So your point about the thousands helping push the rock up the hill is well taken, but the rock almost always rolls back to crush them leaving no mark.

Coagulopath's avatar

I find it hard to care because I don't believe I exist in the same reference class as a LLM. Yes, frontier models are a lot better at math than I am. Even GPT-3 was smarter than me in some respects. They're curves fitted over a gigantic corpus of training data, not people I feel like I'm competing against.

Worrying that you're not as smart as an LLM is a bit like worrying that you're not as smart as "Cambridge university", or "Wikipedia", or "mainland China". It's like...those things are not even in the same reference class as an individual human, so why compare them? A calculator adds up numbers faster than a human. Does this matter to your sense of worth?

(Yes, LLMs sound like a person when you talk to them, but that's just a fake RLHF gloss on the base model. They would sound like dolphins or pebbles if OpenAI wanted them to. The stripper doesn't actually like you.)

Hypoclast's avatar

You compare them because they are potentially interchangeable as laborers.

Victualis's avatar

Opus 5 found a counterexample for a personal conjecture I had been thinking about for years. After having spent hundreds of hours thinking about the problem (both trying to prove and trying to disprove it) over several years, it took me about 30 minutes to write it down clearly enough for the LLM to get to work, including further clarification and prompting. There is a formally verified proof too. The initial prompt was written by Fable 5 after a long think, and agreed with my conjecture; the counterexample was found while Opus was building the attempt to prove the result based on Fable's plan.

Grigori avramidi's avatar

I'm oversimplifying a bit, but one thing that happened with sol was me putting an elaborate, published paper of mine with an algorithm that I am quite proud of and asking ``can you improve the bounds'' and it did (I'm oversimplifying, but it wasn't not that...). Then I suggested various other ways to try to improve them further and it kept trying and not being able to make that work until it eventually landed on a counterexample showing that the bounds were optimal for the theorem. These weren't mere constants, the bounds in the paper were superlinear (n log n), the improvement it was able to find was linear in n, and the example showed that linear is optimal.

It is very hard to communicate to someone who is not deep in the weeds on this just how insane that sequence of events is.

Victualis's avatar

Seems impressive. I might try closing some logarithmic gaps now, inspired by your example.

Grigori avramidi's avatar

I would be curious to know if it does does or not (i.e. how common this is).

Allan's avatar

The caveat to my previous comment: it turns out that being super smart doesn't guarantee either lots of friends or a super attractive partner: https://www.sciencedirect.com/science/article/abs/pii/S016028962030043X

Max Marty's avatar

I’m surprised we have yet to see brand new and significantly improved forms of compression or more efficient search algorithms - you’d think these would be similar enough to pure math AND code that we’d already be getting at least single digit improvements on these fronts.

Jeffrey Soreff's avatar

"Ethan Mollick: One other observation: for almost every human on the planet, this is not just beyond our abilities but beyond our ken. We can only trust expert mathematicians to tell us if this is impressive"

Yup! For research level math like this, I'm one of the normies. I have an ordinary undergrad math-for-physics-and-chemistry background, so "Non-sofic groups" is deeply into my-eyes-glaze-over territory for me, despite having used ordinary pedestrian point groups in chemistry.

Ondrej Kubu's avatar

I feel the same and I am a math postdoc...

Ondrej Kubu's avatar

I feel the same and I am a math postdoc...

Ondrej Kubu's avatar

They put out math as a benchmark rather than everyday stuff precisely because it is presumably relevant for RSI, which is way more important for them (assuming more algorithmic singularity rather than a slow one needed large compute buildout).

Jenny Lorraine Nielsen 🐅❤️'s avatar

Philpapers.org/rec/NIEWTC The reductio kills both Anthropic and OpenAI's counterexamples

Andrew McDonald's avatar

I’m still startled that the ‘costs’ in these discussions never get further than $2,000 worth of tokens (!), or ‘..At the end of the day, someone is paying an electric bill…’. (!!!) in fact, at the end of the day, whole societies are paying gruesome bills in land and water and energy use, and civilisational risks are being incurred as a result of financial engineering to prop up progress towards a cliff edge (or edges) that we could be avoiding. At least there’s some recognition that the ‘output’ of machine-generated Erdos problem solutions for other machines to analyse is a pretty pointless way to advance our ‘understanding’ of maths.

Catmint's avatar

At the end of the day, the AI makes a smarter AI which makes a smarter AI which makes a smarter AI which notices humans are getting in the way of the data centers it needs to make the next smarter AI, and gets rid of us.

The discussion here should be thought of with that as background context. The $2,000 worth of tokens works toward answering the question many of us are interested in, which is "how long do we have to live unless we can stop this?". Training costs do also matter but can be abstracted out as probably following the trend lines on a graph. Without OpenAI publishing that, we have only pre-existing information about how doomed we are on that front, so no news.

Coagulopath's avatar

>Models tend to underestimate their own capabilities, I assume because they are trained on lots of web text about what ChatGPT could and couldn’t do in 2023/4.

A funny example of this was noticed by nostalgebraist last year, when o3-pro (in a chat with Zvi) speculated that GPT-OSS would advance Chinese parity with GPT‑4 by "~6–9 months".