On the Navier–Stokes Millennium Prize Problem
Posted by tedsanders 4 hours ago
Comments
Comment by peri-cl 48 minutes ago
https://mathstodon.xyz/@tao/117237320796901560
> "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field."
Comment by vessenes 37 minutes ago
Mathematics has always been highly competitive.
Comment by Jtariiiii 18 minutes ago
Comment by techas 2 minutes ago
I found this behavior against healthy science practices and only driven by ego. Unfortunately, I find this too often at work (working in academia). Most probably I'm too naive...
Comment by dev_dan_2 9 minutes ago
"But I realized after a while that talking to people casually about Fermat was impossible, because it just generates too much interest, and you can't really focus yourself for years unless you have this kind of undivided concentration, which too many spectators would have destroyed."
But yes; him reaping the benefits of himself having the idea first was part of it too; as far as I am aware.
-----
Which is still something completely different than some anonymous organisation keeping mathematical research secret because it is better for hype reasons. One is competition between individuals or groups within a field; the other is boring and sometimes borderline nihilistic generating of mathematical knowledge as an marketing asset.
Comment by enraged_camel 34 minutes ago
Just curious: are you aware of who he is?
Comment by usrnm 29 minutes ago
Comment by dev_dan_2 15 minutes ago
Also note how the quote by Tao is in all likelyhood not meant as an absolute; rather than a statement of a trend - a handfull of counterexamples do I no way change anything about the truth value of Tao's quote.
On the other heand; consider how absurd it would be if "... in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science ..." would indeed be a misstatement; which would imply that far more promising research directions were not shared with the broader community (i.e.: published). I wonder what different reading of that counterfactual there could be other than secret societies that kept their discoveries and research directions to themselves - which we just learned about (since we would otherwise not be refering to the secret societies and their supposed promising research directions).
All pretty straightforward, I would say - both that "misstatement" is hopefully based an overly strict reading of Tao's quote, and that mentioning Tao's background as one of the fields leading practitioners is relevant as well. Again; to make sure: A few counterexamples achieves nothing here. It would need to reach a certain threshold of such counterexamples before we will have to write the history of mathematics; and before Tao actually made a misstatement here.
Comment by enraged_camel 16 minutes ago
Get outta here.
Comment by 1w2hagsFa 24 minutes ago
Comment by dhhdhjoe 25 minutes ago
Comment by cyclopeanutopia 28 minutes ago
Comment by 1w2hagsFa 22 minutes ago
Comment by gradus_ad 7 minutes ago
Not sure I agree with this. AI generated proofs can still be analyzed and mined for useful insights. I suppose he's saying the process of banging our heads against the wall on a problem can itself yield useful insight? But what is stopping us from analyzing a proof after the fact. And if we can generate many different versions of a proof that should help us develop a much deeper understanding of the problem than we would have without being able to perceive the "proof landscape"...
Comment by SpicyLemonZest 2 minutes ago
Comment by ozgung 3 minutes ago
Comment by olalonde 11 minutes ago
Comment by tzone 8 minutes ago
But within next 10 years as costs drop significantly and even more improvements are made, yes it is very likely that almost every single existing math problem will get a serious AI cracking done on it
Comment by ltbarcly3 30 minutes ago
What is going to happen is a complete revaluation of things like "finding a counter example to a famous problem". Even if someone finds a solution to a problem like this with pencil and paper, nobody will believe it, and they will assume that there was an AI involved.
Further, sitting and doing math with a pencil and paper will no longer be a reasonable strategy to build a reputation or career, beyond the benefit a mathematician gains to their own intuition and skill. People who work hard to build intuition and also use AI effectively will dominate the field.
In a world where everyone is using AI, the open problems that remain will be the ones that are AI resistant. This is no different that how things work now, mathematicians wait until they are fairly confident someone won't rapidly solve their problem before they start talking about it. They will do the same thing in the future, except in the future AI will be part of the toolset they use decide if they are ready to share yet or not.
Edit: Ok I believe I was generally right here, but I just read the details of what OpenAI did. They didn't solve a longstanding problem, they got tipped off to an approach a mathematician was using and would likely result in the solution very soon and they finished it first. If this turns out to be true I think my take above is not correct, in the short term people will have to stop sharing updates because otherwise openai will dishonestly race to finish their work.
Comment by dev_dan_2 6 minutes ago
Which I don't see a reason for Anthropic and "Open"AI not to, given their not so stellar track record with IP of individuals/entities-that-are-not-rich-enough ;)
Comment by pavel_lishin 4 hours ago
https://news.ycombinator.com/item?id=49605915
https://bsky.app/profile/quantian.bsky.social/post/3muyhwbcd...
Comment by tedsanders 4 hours ago
I work at OpenAI, though not on the team that did this, and my understanding is:
- we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
- we did not read any private chats (but of course the model was aware of prior research literature published to the internet)
- the proof generated by our model was very different from theirs and also goes far beyond the published literature
- we made an effort to jointly announce rather than immediately scoop (I understand Tristan was unhappy with the conversations; I know zero details here and I hope more is shared today)
Edit: Here's is Sebastian's take: https://x.com/SebastienBubeck/status/2097379411691516310?s=2...
Comment by contemporary343 4 hours ago
- This, from Tristan Buckmaster's writeup yesterday, indicates to me that there was more than incidental inspiration from Alpoge and Buckmaster.
Comment by tedsanders 3 hours ago
- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input
- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces
I'm not sure how any of this provides evidence that OpenAI took any of their work.
As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.
(I work at OpenAI, but not on the team that did this proof.)
Comment by whimsicalism 43 minutes ago
Comment by enraged_camel 1 hour ago
Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).
Comment by fc417fc802 55 minutes ago
Comment by nulld3v 24 minutes ago
Comment by fc417fc802 13 minutes ago
Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.
Comment by nulld3v 3 minutes ago
Comment by derangedHorse 41 minutes ago
I don’t think they’re too concerned about appeasing you, enraged_camel.
For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.
Comment by pred_ 4 hours ago
Your post says “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .” We can discuss what it means to “read” things but obviously the issue here isn't whether you did it manually or automatically.
But more importantly, what on earth are you doing threatening real scientists to remove their coauthors, then making fun of them on social media? Does the entire company run on that toxic culture, or did those people run off of some kind of outrageous tangent?
Comment by Caracas288 35 minutes ago
Comment by nerevarthelame 1 hour ago
Comment by biesnecker 1 hour ago
Comment by swat535 45 minutes ago
Comment by igleria 3 hours ago
That is your opinion, but the optics of that should raise for you some flags. OAI could have waited (how long is a task left to the ethics committee) to see how the rumors panned out. Right now the optics look a lot like "we don´t care there is a 1/7 chance we one-up a human researcher by reacting to this rumor immediately, might makes right"
Comment by Imnimo 3 hours ago
The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?
Comment by tedsanders 1 hour ago
If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly, highly doubt it made a difference to a problem as challenging as the NS proof.
Reasons for my doubt:
- I know most of our training recipes
- Our model's proof is very different from theirs
- The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question)
- This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution
I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here.
Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying.
If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful.
Comment by Imnimo 3 minutes ago
Either way, it seems worth having clarity, and I'm a bit surprised OpenAI's stance is just "we can't rule this out, but don't worry about it". OpenAI is, apparently, very happy to use unreleased models to try to scoop big results if they get a whiff that someone else is close (which strikes me as pretty scummy regardless of any issues of training contamination). It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.
Comment by mucha 45 minutes ago
Comment by sebzim4500 5 minutes ago
Comment by vemacs 19 minutes ago
Comment by lambda 39 minutes ago
Even better would be more research and tools to help determine the impact of particular training data on models. Right now, proprietary LLM providers get to hide a lot behind "we just train it, we don't know what inputs affect the outputs," and that can be a problem, both because of lack of traceability of factual informaiton as well as lack of traceability of things like this, where the model itself may have had unpublished work in its training set.
Comment by lossolo 47 minutes ago
Can't you guys just check their account settings so the public knows what was set?
Comment by moralestapia 1 hour ago
No answer is also an answer.
He's a human, like everybody else. Mostly a bunch of hungry animals looking to put bread in our mouths. It's rarely ever something a bit more sophisticated than that.
Comment by irthomasthomas 36 minutes ago
Comment by tzone 2 minutes ago
Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .
Wild times
Comment by interestpiqued 2 hours ago
Comment by caughtinthought 22 minutes ago
Comment by sebzim4500 3 minutes ago
Comment by nhatcher 48 minutes ago
Comment by suddenlybananas 4 hours ago
Comment by tedsanders 4 hours ago
Comment by suddenlybananas 4 hours ago
Comment by sk4rekr0w 4 hours ago
Comment by suddenlybananas 4 hours ago
https://news.ycombinator.com/item?id=49605915#49610498
https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=6...
(I can't reply to the below comment, but I was aware this was about Sebastien, I was trying to be charitable by including stuff said about both people)
Comment by qt31415926 4 hours ago
Comment by sk4rekr0w 3 hours ago
Comment by dandanua 4 hours ago
Comment by applicative 4 hours ago
I dedicate my life to its complete destruction beginning today.
Comment by beering 4 hours ago
Comment by floatrock 4 hours ago
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
Comment by biophysboy 4 hours ago
Comment by enraged_camel 36 minutes ago
Comment by cute_boi 4 hours ago
Comment by andrewguenther 4 hours ago
Comment by luke5441 2 hours ago
That they don't is telling.
Comment by heaney-555 4 hours ago
>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
Comment by SpicyLemonZest 4 hours ago
Comment by arctic-true 4 hours ago
Comment by danielmarkbruce 1 hour ago
Comment by chilmers 4 hours ago
[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/
Comment by 10xDev 4 hours ago
Comment by hgoel 2 hours ago
In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).
Comment by Fordec 3 hours ago
Comment by mrbungie 2 hours ago
Comment by Fordec 2 hours ago
Comment by HenrikPontoppid 2 hours ago
Comment by Fordec 2 hours ago
Comment by supern0va 1 hour ago
Comment by jsLavaGoat 1 hour ago
And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.
Comment by Fordec 1 hour ago
Comment by dakolli 2 hours ago
Comment by 7373737373 23 minutes ago
Comment by Fordec 2 hours ago
Comment by a2ff6eeb0 1 hour ago
Scroll down to the existing examples section.
Comment by monster_truck 2 hours ago
Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged
Comment by Fordec 2 hours ago
Comment by xtracto 1 hour ago
Comment by Miner49er 3 hours ago
Comment by glenstein 2 hours ago
I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.
Comment by lijok 1 hour ago
Comment by magicalist 4 hours ago
Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?
Comment by ameliaquining 3 hours ago
Comment by 20k 3 hours ago
What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question
If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?
Edit:
OpenAI have admitted they were training on prompts at the time they made their breakthrough
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Comment by ameliaquining 3 hours ago
If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.
Comment by sdenton4 1 hour ago
Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.
Comment by 20k 3 hours ago
The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem
Comment by orangecat 3 hours ago
Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.
This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.
Comment by 20k 2 hours ago
Comment by letmevoteplease 1 hour ago
Comment by ivory54321 2 hours ago
Comment by Timwi 36 minutes ago
Comment by lotsofpulp 3 hours ago
Comment by 20k 2 hours ago
Comment by lotsofpulp 1 hour ago
Comment by doctoboggan 2 hours ago
Comment by user43928 1 hour ago
That's it. The rest appears to be wild speculation.
Comment by jsw97 1 hour ago
Never ever touch those requests. If you get a side by side comparison just resend the prompt.
Comment by zem 55 minutes ago
Comment by flir 46 minutes ago
Publication, though? Slimy is right.
But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.
Comment by TZubiri 1 hour ago
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.
Comment by TZubiri 1 hour ago
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.
Comment by an0malous 1 hour ago
Comment by ameliaquining 1 hour ago
Comment by pama 4 hours ago
Comment by danielmarkbruce 36 minutes ago
Comment by _fizz_buzz_ 2 hours ago
Comment by tristanj 1 hour ago
Comment by gcr 1 hour ago
Comment by lossolo 1 hour ago
Comment by mzhaase 3 hours ago
Comment by monster_truck 2 hours ago
Comment by Bluestein 2 hours ago
Comment by ccozan 46 minutes ago
Comment by dboreham 3 hours ago
Comment by karmakurtisaani 2 hours ago
Comment by dakolli 2 hours ago
Comment by E-Reverance 2 hours ago
Comment by sznio 52 minutes ago
Comment by fc417fc802 1 hour ago
Comment by naveen99 4 hours ago
Comment by sashank_1509 4 hours ago
Comment by credit_guy 3 hours ago
Comment by Aboutplants 4 hours ago
Comment by stingrae 3 hours ago
Comment by curt15 3 hours ago
Comment by blake__dev 3 hours ago
Comment by cool_dude85 3 hours ago
Comment by merksittich 1 hour ago
Comment by blake__dev 3 hours ago
Comment by bananaflag 3 hours ago
Comment by refulgentis 3 hours ago
Comment by vatsachak 3 hours ago
Loops and parallel connections make transformer go brrr
Comment by irthomasthomas 1 hour ago
Comment by fer 51 minutes ago
Comment by chinathrow 4 hours ago
Comment by Aboutplants 4 hours ago
Comment by jrflo 4 hours ago
Comment by mrbungie 4 hours ago
1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence.
2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.
Comment by scurnus 3 hours ago
Regarding product and credibility normal people have a completely different view about LLMs, most don't even know difference between models and probably don't even care about Millenium problems, but care instead if chatgpt can solve their day to day problems. This is just them trying to have the throne on the AI companies space, outside it this result won't matter.
Comment by mrbungie 3 hours ago
Did I say otherwise?
> 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.
I know, but I don't know how that relates to my point, which is about the way they are doing it.
Comment by scurnus 2 hours ago
The way they are doing it is by trying to get the attention and staying on top of the news, it is a game they are playing that benefits both OpenAI and Anthropic. The more people discuss SF drama, the less attention Chinese Labs and others get.
Comment by dsdf3 4 hours ago
Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was.
And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.
Comment by danielmarkbruce 56 minutes ago
It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).
Comment by QuesnayJr 4 hours ago
I'm not sure what the top 3 problems are. You can make a case for the Riemann Hypothesis and P != NP, but I'm not sure what #3 would be. Maybe the Langlands program? (That one is not as precisely stated as the other two.)
Comment by anthonypasq 4 hours ago
Comment by dsdf3 4 hours ago
Comment by anthonypasq 3 hours ago
Comment by QuesnayJr 3 hours ago
I am not particularly skeptical of claims about AI, compared to the average here on HN, but that doesn't mean every random piece of hype is warranted. What they did is impressive, even though we now know the only reason they threw so much compute at the problem is that they heard a rumor that someone else was already close. Navier-Stokes is not a top 3 problem in mathematics, and it was the one that was thought closest to being solved.
Comment by ameliaquining 3 hours ago
Comment by QuesnayJr 48 minutes ago
Comment by ameliaquining 33 minutes ago
Comment by andrepd 2 hours ago
Comment by eutropia 4 hours ago
"Mission. Fucking. Acccomplished."
https://xkcd.com/810/Comment by hdivider 4 hours ago
1. It shows what even this wave of AI can actually do.
2. I wish it were done by different folks, ideally under some kind of public control like NASA research or the NPR model.
3. Keep in mind: natural science is different. It's not always a matter of computation. Computer science folks often struggle with this -- but this virtual world here does not actually exist. Everything is physical, including information. Any natural science PhD or otherwise knows just how complicated nature actually is -- e.g. mention any research topic and try to encapsulate all the relevant phenomena present there. Pure mathematics is different because we define the problem, rarher than explore nature. We are in my view far away from removing humans in natural science R&D. Advancements in AI however can greatly assist us in all natural sciences, which is already beginning to happen.
Comment by olalonde 7 minutes ago
This is sort of what OpenAI was supposed to be. I'll never understand how it was legal for them to turn it into a for profit corporation.
Comment by ThePhysicist 3 hours ago
So I'm greatly excited what AI will bring about in physics, more so than in math, because in physics it's clear that our fundamental theories are missing a big piece of the picture, and given how easily AI crunches through Millenium prize problems I think it's possible that AI will come up with a viable grand unified theory uniting quantum mechanics and gravitation, or produce new predictions in other areas. There's enough contradictory or unexplained observational data available to make a ton of progress on the theory side I think. Exciting times ahead!
Comment by alde 1 hour ago
Comment by throwaway198846 2 hours ago
Comment by geremiiah 3 hours ago
Comment by jarenmf 2 hours ago
Comment by efavdb 3 hours ago
Math is like this too. The big problems they've been solving have been identified as interesting only through lots of prior effort.
Comment by brettdev 1 hour ago
Comment by danielmarkbruce 31 minutes ago
Comment by sobellian 2 hours ago
"If in other sciences we should arrive at certainty without doubt and truth without error, it behooves us to place the foundations of knowledge in mathematics."
Comment by semi-extrinsic 1 hour ago
See e.g. https://en.wikipedia.org/wiki/Renormalization
Nobody who actually works in fluid dynamics on any sort of application gives a hoot about the N-S millenium problem. Many do not even know what it is. There is no practical effect of this proof on how we do fluid mechanics.
Comment by sobellian 55 minutes ago
Comment by semi-extrinsic 31 minutes ago
Comment by sobellian 16 minutes ago
Comment by semi-extrinsic 11 minutes ago
Even if someone comes up with a construction that does not require any forcing, it is going to be some extremely weird initial conditions that you will never be able to even approximate in reality unless you can move all the individual molecules of a fluid around and set their initial velocities from a far distance.
Comment by sobellian 9 minutes ago
Comment by red75prime 3 hours ago
BTW, there's also a problem of asking interesting questions that AIs aren't yet good at.
No one has found any principled walls of AI development yet. And empirical results are quite telling. So, I guess, those problems will not stand for long.
Comment by tantalor 3 hours ago
Comment by vatsachak 3 hours ago
The natural sciences will soon start breaking too.
I will concede that AI seems likely to not invent a "research program" anytime soon.
It has no taste
Comment by danielmarkbruce 29 minutes ago
The reason AI is doing so well in math is that it can verify every idea it has, quickly.
Comment by tiborsaas 4 hours ago
WOW?
Comment by jampekka 2 hours ago
This.
I do dislike the AI oligarchs as much as the next person, but I do find the thread full of complaining a bit depressing still.
If the result holds (and it looks it does), this may be one of the, if not the, biggest things to happen in computing to date. A lot bigger than e.g. Deep Blue beating Kasparov in chess or AlphaGo beating Sedol in Go.
Comment by Eridrus 2 hours ago
Comment by qlte 1 hour ago
But, as mathematicians learn and push forward, occasionally something like elliptic curves will emerge as having useful applications, making all that previously "pointless" specialized knowledge newly valuable.
Or advances in physics, that suddenly have a need for a specific mathematical underpinning to develop a theoretical framework. Like how Einstein benefited from Minkowski's work on hyperboloids to create a coherent mathematical description of spacetime.
It was the AI labs themselves not mathematicians who were happy to conflate proofs for open math problems with some kind of tangible technological advancement in the real world. They would surely prefer to be able to claim a cure for cancer vs. a math problem but that loop requires a lot more time/money/test tubes/etc and they need headlines now not in a decade.
And so, thanks to OpenAI/Anthropic, we're now in a world where thousands of crypto bots on X breathlessly hype up each new problem being solved that previously wouldn't have any got any attention beyond academia and passionate fans of math.
Hopefully this won't lead to a trough of disillusionment as more people start to feel like you, with mathematicians getting the blame for inflating the value of their work even though the hype was coming entirely from the labs not them.
Comment by concinds 2 hours ago
Comment by robryan 32 minutes ago
Comment by karmakurtisaani 2 hours ago
Comment by xhevahir 2 hours ago
Comment by rybthrow2 1 hour ago
Comment by empath75 2 hours ago
Comment by echelon 4 hours ago
- First off, to reiterate, WOW.
- Second of all, when does this end? Are we at the dawn of the singularity now?
- People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?
- Time to think about retiring from any knowledge work or business? This could be winner-take-all where a leading lab can button press any economic function, business process, or scientific discovery. 24 months of lead on Open Source might turn into virtual centuries of lead.
- Do "normies" even know what's happening?
Anybody who thinks the improvements stop here isn't paying attention. It hasn't been showing any signs of slowing down since 2018. And the curve isn't even linear! My god, next year is going to be insane.
Comment by tiborsaas 4 hours ago
3) I'm still processing the drama, just found out about it after reading the blog post. If that happened based on private data, that's horrible. If that happened based on public tweets, then it's still abuse of power as OA employees access to compute (launching 10k agents) is quite heavy weight in boxing terms.
But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?
Comment by 20k 3 hours ago
Its a bit like solving p = np with a negative result. Its an incredibly difficult problem, but it doesn't lead to anything at all on its own. This is why people are talking about the fact that the solution methodology is much more interesting than the solution - the tools used to crack something like this may lead to solving more useful problems
Comment by inkysigma 3 hours ago
There's unlikely to be any engineering applications since even if the solution can be approximated, you still need to set up the initial conditions but at that point you can also drive pressure in other ways.
Comment by thangalin 2 hours ago
Proving out the combination of scaling inference-time compute and agent collaboration to solve previously intractable mathematical problems is WOW. By pairing creative candidate generation with automated proof checkers (like Lean) we are leaning into a repeatable framework for AI-driven scientific discovery.
Comment by semi-extrinsic 1 hour ago
This is 100% wrong and reads like copy paste of AI slop.
Any simulation which uses sub-grid scale models is already solving a different PDE than the actual Navier-Stokes considered in the Millenium problem, and that PDE is guaranteed to have different properties. Full stop.
And to claim this is somehow connected to AMR methods is an example of the kind of pseudoscientific statement Wolfgang Pauli would have called "not even wrong".
Comment by xyzsparetimexyz 53 minutes ago
Minor productivity boost in mathematics as people are no longer nerdsniped by the problem
Comment by cyberax 3 hours ago
Nothing, really. This mirrors other examples of blowups from the classical physics. It's possible to create a system with just gravitating bodies that exhibits a blowup to infinite speeds in a finite time. The root cause is that, in classical physics, the speed of gravity is instant.
In the case of Navier-Stokes, the fluid is incompressible. So technically any force that you apply to it is supposed to instantly affect everything else. This can be exploited to create these blowups. In reality, no fluid is incompressible, and it takes time for any action to affect the material.
It's just that Navier-Stokes equations are so slippery that it's hard to pin their behavior down. They basically just restate the momentum conservation law for a continuous medium.
Comment by trio8453 3 hours ago
No, there are even many non-normies talking about how it's all marketing or try to give balanced take about AI being sometimes a little useful for certain things (but they can do without it anyway).
Comment by stefap2 4 hours ago
Comment by munificent 3 hours ago
You really think it makes sense for you to be higher on the "solving complex problems ladder" than the machines that solved fucking Navier-Stokes?
I envy your self-confidence.
Comment by stefap2 3 hours ago
Comment by fooker 2 hours ago
For example there are no engineering implications of this solution yet.
For the next several decades, we'll have engineers (presumably with AI) optimize things like rocket engines and turbines and AC compressors to work a few percent better because the numerical approximations might have caused us to be overly conservative.
AI is not going to magically solve all random problems. Pick a career where you are in the driver seat.
Comment by semi-extrinsic 58 minutes ago
No. Just no.
Comment by mlsu 3 hours ago
Comment by reducesuffering 3 hours ago
Comment by mlsu 3 hours ago
If that is true then this seems to be, again, a case of AI producing an interpolation over data it has seen before. Everything about openAI's behavior indicates that they were using the transcripts as input. Why not have the AGI choose a different Millenium prize problem?
Comment by cyberax 2 hours ago
But it did find a long-suspected smooth solution with a singularity.
Comment by biophysboy 3 hours ago
Comment by root_axis 2 hours ago
This is an impressive result, but there is absolutely zero evidence of "the singularity".
Comment by hackinthebochs 29 minutes ago
Comment by root_axis 16 minutes ago
Comment by hackinthebochs 10 minutes ago
Comment by tantalor 3 hours ago
Singularity doesn't "dawn". That's the whole idea. It happens all at once.
Comment by armchairhacker 3 hours ago
Comment by reducesuffering 3 hours ago
Comment by armchairhacker 3 hours ago
Comment by qlte 14 minutes ago
"Moving the goalposts" as shallow dismissal doesn't work if e.g. someone points out AI hasn't even built a new type of spaceship yet in response to a claim that AI is on the verge of building a Dyson sphere.
Comment by baq 2 hours ago
> - Second of all, when does this end? Are we at the dawn of the singularity now?
normalcy overhang n. /NOR-muhl-see OH-ver-hang/
The uncanny period during the Singularity when superintelligence is already accomplishing feats that seem like magic, yet everyday life still looks mostly the same.
Comment by Bluestein 3 hours ago
Comment by onidj 3 hours ago
Absolutely not. Even to a lot of techy/nerdy people it's still just a chatbot that they sometimes use to help them at work. Even on here people will do whatever they can to downplay.
The lack of fucks given is staggering.
Comment by nozzlegear 1 hour ago
Comment by d_silin 4 hours ago
Comment by raincole 4 hours ago
The said user (Tristan Buckmaster) didn't solve the millennium problem. He didn't really accuse that OpenAI stole his research either. The beef came from the fact OpenAI asked him to remove another mathematician, who works for Anthropic, from the credit.
"People" are just misinformed and keep spreading misinformation.
Comment by 20k 3 hours ago
He very much is accusing them of stealing his work
Comment by naasking 4 hours ago
Not quite accurate, Buckmaster was taking an approach that nobody else was, and this new proof uses this same approach just weeks after he saved those results to OpenAI workspaces. He asked OpenAI if they used chat logs for training the new model, and they did not confirm or deny.
Asking to remove his collaborator is also totally over the line though.
Edit: although this OpenAI post is not comforting: https://x.com/OpenAI/status/2097375276384567642
Quote: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. "
Comment by stefap2 3 hours ago
Comment by emp17344 4 hours ago
Comment by achierius 4 hours ago
> I should say here why I interpreted their statement the way I did, the in- terpretation I will discuss below. The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag.
...
> I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI. > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
It's not a direct accusation, but it's not far off.
You shouldn't accuse other people of spreading misinformation when you haven't read the actual sources in question, it's possible that they might know more than you.
Comment by raincole 4 hours ago
> I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything.
People saying that he accuses OpenAI stole his proof are putting words into his mouth and I consider that very disrespectful to him. It's basically using Buckmaster as a tool to express their dissatisfaction over OpenAI.
Comment by mewse-hn 4 hours ago
What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
Comment by WarmWash 3 hours ago
If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".
Comment by nozzlegear 52 minutes ago
At this point, how can we even trust that they aren't accidentally training on those tokens too?
Comment by jdm2212 11 minutes ago
Comment by 14u2c 1 hour ago
Comment by spruce_tips 1 hour ago
Comment by magicalhippo 39 minutes ago
Basically individual accounts can opt out, while business and enterprise plans as well as API users can opt in.
You'd have to take their word, but that goes for anything in life.
[1]: https://help.openai.com/en/articles/5722486-how-your-data-is...
Comment by perching_aix 3 hours ago
Comment by lima 1 hour ago
Comment by TZubiri 1 hour ago
Comment by nradov 4 hours ago
Comment by gowld 3 hours ago
Comment by red75prime 2 hours ago
But, yeah, priority is much more finicky. The Newton/Leibniz drama was quite something.
Comment by brainwad 2 hours ago
Comment by dash2 4 hours ago
Comment by sinuhe69 2 hours ago
Comment by elwell 2 hours ago
Comment by taylorfinley 52 minutes ago
Comment by vessenes 2 hours ago
All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.
Comment by jimbob45 2 hours ago
Comment by jakevoytko 4 hours ago
Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees
Comment by traes 1 hour ago
> 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee.
> 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)
https://xcancel.com/SebastienBubeck/status/20973794116915163...
Comment by closetheloopdev 2 hours ago
Comment by philipwhiuk 3 hours ago
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Comment by recitedropper 4 hours ago
The dark forest awaits..
Comment by AlexErrant 3 hours ago
2. The dark forest is fun for scifi stories, but is mathematically bunk anyway https://www.noahpinion.blog/p/the-dark-forest-hypothesis-is-... https://www.reddit.com/r/IsaacArthur/comments/1l06cnk/cool_w... https://www.projectnash.com/aliens-the-fermi-paradox-and-the...
When doomposting please actually say something substantive. Negative news always gets clicks/updoots; fight that human tendency.
Comment by recitedropper 3 hours ago
I agree that we have not solved the Fermi paradox; I disagree that comments highlighting immature behavior from people who wield enormous power in our world are unproductive.
Comment by AlexErrant 3 hours ago
Separately, I disagree that intellectual work has ever been free of "dark forest"-style secrecy. Scientists everywhere have worried about being scooped; AI just magnifies that (as all tools have; e.g. Leeuwenhoek lenses).
And thirdly, if you want to make a stronger case for "I feel even less confident in them as a team to be shepherding this much capital and compute", you should give citations and arguments. From what I've seen, there's drama, it's much OpenAI trying to avoid scooping, and Tristan being stuck in a game of telephone, and Levent being incommunicado.
If you have a better analysis, you should say so instead of being vague.
Comment by recitedropper 2 hours ago
I appreciate your upholding of ideals, and since I respect that, I will honor with final replies:
1. Locktime has passed.
2. Yes, intellectual work has always had elements that incentivize secrecy. If you want to say we were already in a "dark forest", so be it. My suggestion is that the multiplier AI adds to the possibility you get scooped is a step-change, and therefore we now enter a new "dark forest".
3. This is a big thread, and the twitter antics are well-documented, so I would assume someone else has cited them. If not, I think most are aware at this point that the online antics of AI researchers, especially when announcing or citing mathematical advancse, have regularly been childish.
Comment by AlexErrant 1 hour ago
3. ctrl-f "x.com" in this thread only yields https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 https://x.com/OpenAI/status/2097375276384567642
and frankly, I'm unwilling to give Elon any more traffic to dig up drama that ultimately doesn't matter. I'm not seeing anything especially childish, but y'know... I'm not sure I care.
Comment by zem 32 minutes ago
Comment by vmasto 4 hours ago
Comment by nicce 1 hour ago
Comment by sheafification 4 hours ago
Comment by recitedropper 3 hours ago
So less about hiding civilizations, and more about hiding information. Math is clearly headed in this direction, and I see no reason why the rest of intellectual work shouldn't too.
Comment by intenex 2 hours ago
This specific problem having had a $1 million bounty on its head and still remaining unsolved for 26 years after the bounty was placed is pretty clear evidence that many of the world's best human mathematicians would have solved this problem if they could have, and none were able to until LLMs came along.
Hard to claim at this point that LLMs aren't capable of novel STEM creativity and genius to a degree that will soon far surpass that of humans.
If anyone has counterpoints to this I'd love to hear them!
Comment by hansvm 3 minutes ago
Not to mention, it's still very much up in the air whether the model derived the answer of its own accord or sniped the important details from the researchers it was spying on.
Comment by adverbly 2 hours ago
To be fair, I think it's still an open question about how far it might surpass human capabilities.
I think it's clear that its speed of development will be significantly faster, but it's technically not proven that the frontier and problems don't themselves become increasingly difficult faster than any acceleration in intelligence past the point of human training, data and existing knowledge.
Should this be the case, we would see a rapid broadening of development, and a slow advance in the frontier in such a way that might surpass the collective capabilities of people, but not by very far.
Comment by redox99 1 hour ago
Comment by piker 2 hours ago
Comment by intenex 2 hours ago
Comment by piker 2 hours ago
[Edit: my only point here is that the prize is probably not driving human effort to the limit.]
Comment by jampekka 2 hours ago
Comment by superxpro12 2 hours ago
Comment by philipwhiuk 2 hours ago
Comment by gpm 2 hours ago
Comment by aeve890 1 hour ago
Sure. A proof without an unknown amount of human steering (and/or stolen research) would be an unquestionable achievement.
To this day there's zero (0) evidence of any result by an LLM alone (maybe I'm wrong). If I just prompt ChatGPT right now with "give me a proof of the Riemann Hypothesis" and this thing delivers, I'm sold. But anything close to "yeah ChatGPT proved X with 5 years of 24/7 work with 10x Terrence Tao level geniuses" it really doesn't cut it.
Or why's there's no new branch of mathematics invented by AI? That'd be indubitably _novel_ and _creative_. But to my knowledge (and I'm eager to be educated) there's nothing like that. What are the HARD examples of novelty, creativity and genius you claim? For how people like you talk about AI I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware. Sure it would infinitely easier to make OpenAI literally print money with any of the thousand problems easier to solve with such amazing intelligence than the NSE problem right? Honest question
Comment by redox99 1 hour ago
Also there are proofs where the only human steering was "keep going".
Comment by sp527 10 minutes ago
Comment by aeve890 1 minute ago
Comment by aeve890 37 minutes ago
Any result of such kind from an AI alone would be enough to refute my argument, yet you don't present any.
>Also there are proofs where the only human steering was "keep going".
Which ones?
Comment by closetheloopdev 3 hours ago
- There are at least two versions of a model more powerful than Astra at OpenAI at the moment.
- The less capable version was used to solve the unforced Euler problem (while the one solved by Levent Alpöge and Tristan Buckmaster was forced Euler) with 100 agents.
- The more improved version was used to solve Navier-Stokes, given the results of the unforced Euler problem from their earlier attempt, with 10000 agents.
- OpenAI initially tried a shotgun approach against the 6 Millennium Prize Problems until it emerged that Navier-Stokes was the most likely to succeed.
So the timeline was:
Shotgunning 6 open Millennium Prize Problems -> solved unforced Euler problem with 100 agents -> concentrating on Navier-Stokes with 10000 agents -> solution.
If so, that is fantastic development and a huge success (despite all the drama surrounding it)! Congratulations!
Comment by tristanj 1 hour ago
Comment by closetheloopdev 57 minutes ago
I hope the next solved Millennium Prize Problem will have less drama.
Comment by piker 2 hours ago
Comment by highfrequency 2 hours ago
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out does not mean what they imply it means.
Comment by MichaelDickens 1 hour ago
Just because something is legal and permitted by terms of service doesn't mean it's morally right.
Comment by Jtariiiii 1 hour ago
What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data? Are they supposed to manually review all their data to make sure competing mathematicians didn't accidentally leave the "submit prompts" toggle on?
Or were they supposed to not try to solve Navier-Stokes, or were they supposed to just not tell anyone that they had solved it?
Comment by plaidfuji 26 minutes ago
It would actually be a really interesting study, if they would ever be willing to be transparent about this, how the result differs with and without his conversations in the training set. How quickly it arrives at the result, whether it takes the same approach, etc.
Comment by vemacs 25 minutes ago
Yes. They should determine if training data included this teams data. Consider the money they spent, the press release and the purpose of their publication.
Since they failed to answer this question they shouldn't have published.
Comment by nozzlegear 44 minutes ago
Personally, I would expect them to have a little class, to KYC, and to manually turn off training for known competitors using their service so as to avoid any unforced goofs like this.
Comment by pred_ 4 hours ago
And what's a better way of empowering people than robbing them.
Comment by rfgplk 4 hours ago
Better than the walled gardens of most journals where you can't even read half the papers without shelling over thousands of $$$
Comment by 20k 3 hours ago
Comment by heaney-555 4 hours ago
Comment by alberto-m 4 hours ago
Comment by heaney-555 2 hours ago
Comment by denverllc 4 hours ago
Comment by keeda 25 minutes ago
It's like that story about George Dantzig solving open problems as a student because he thought they were simply homework: https://en.wikipedia.org/wiki/George_Dantzig
It's also unfortunate that such a potentially momentous occasion is overshadowed by so much drama. Which I suppose is expected given the technology and the people involved are so polarizing.
Comment by sega_sai 4 hours ago
IPO+rumour driven research.
I appreciate the achievement, but it doesn't feel right.
Comment by Aboutplants 3 hours ago
Comment by railgunmerlin 4 hours ago
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Which seems a bit irresponsible/rash?
Comment by paxys 4 hours ago
Comment by rakejake 4 hours ago
"deidentified data" isn't much to go by. Say I prompted the internal model this way - "Hey there's a solution to a unsolved problem X. The solution uses a less known Method Y so don't bother wasting time with the usual methods. Take papers A, B and C as references. Oh btw, here's the last year's worth of data of all prompt sessions that mention this problem. Pay special attention to the ones that mention Method Y and sub-keywords Z,W".
This is obviously all speculation but the timing is very suspect. If OAI actually did this (and I suspect whatever they did is pretty much close to this), I think it is highly unethical.
Comment by perching_aix 3 hours ago
Oh I don't know, maybe something like this?
"Given how seriously this would violate the most fundamental of academic standards, as well as taint the claimed capability behind this result, we take this issue very seriously, and we're launching a probe into identifying whether any of their research artifacts have entered our training set. We have further begun making changes to our UI/UX on all our surfaces, so that it is always clear whether any particular chat, or other user artifact, is eligible for being trained on."
Comment by applicative 4 hours ago
Comment by SpicyLemonZest 4 hours ago
Comment by fooker 4 hours ago
Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.
Comment by SpicyLemonZest 4 hours ago
Comment by fooker 3 hours ago
It never happens that you wake up one morning and start working on a new problem that came to you in a dream (*unless you are Ramanujan).
This is business as usual for academia, it's amusing to the discussion over it.
Comment by rf_physics 40 minutes ago
From my perspective, this practice is quite bad mannered, unusual, and heavily frowned upon, but I recognize it's possible that these stories might be more common in other fields. Still, I'd appreciate a strong sign that you aren't making these statements up based on secondhand accounts of what 'academia is usually like'.
Comment by SpicyLemonZest 2 hours ago
Comment by fooker 2 hours ago
Being scooped is a big deal for the one getting scooped.
It has never been a big deal for the one doing the scooping. History is full of math and science results being scooped. For example, we keep calling it Pythagoras' theorem a few thousand years later.
Comment by railgunmerlin 4 hours ago
Comment by pwign 4 hours ago
> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
Comment by QuesnayJr 3 hours ago
Comment by railgunmerlin 4 hours ago
Comment by QuesnayJr 3 hours ago
Comment by Analemma_ 4 hours ago
Comment by tedsanders 3 hours ago
I promise you that if we took their work from ChatGPT and stuck in a bunch of weasel words to give the opposite impression while remaining technically true, I would quit on the spot.
(I work at OpenAI.)
Comment by sensanaty 2 hours ago
Comment by jsw97 4 hours ago
Highly persistent agents + vibe-coded security seems like a problem.
Comment by suddenlybananas 4 hours ago
Comment by viccis 4 hours ago
This is no different than scooping them.
Comment by verytrivial 4 hours ago
Comment by rakejake 4 hours ago
Comment by Jonasori 4 hours ago
Comment by kzrdude 3 hours ago
Comment by 20k 3 hours ago
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Comment by Legend2440 4 hours ago
He was also using LLMs to do it, so either way most of the credit goes to the LLM here.
Comment by mswphd 4 hours ago
1. he was working on the same class of problems. He explicitly mentions they were working to extend their techniques to NS (the same techniques that OpenAI may have scooped somehow), and
2. while he was using LLMs to do it, this was part of fleshing out another mathematician's work in the area. He explicitly writes in his note that this other mathematician (Luis Martinez-Zoroa) deserves a Fields medal for this work.
Comment by traes 1 hour ago
Comment by pretendscholar 1 hour ago
Comment by jackie293746 4 hours ago
Comment by applicative 4 hours ago
Comment by raincole 4 hours ago
Comment by colesantiago 4 hours ago
Nobody cares and will care about the drama, it is just marketing.
This is the point where were definitely have reached AGI.
Comment by Bluestein 3 hours ago
Sentience aside, moot at this point, the fundamental issue here is that even a deviously ambitious human does not necessitate goal-pursuit itself to breathe, live, exist and have its being. An AI's goal is all it has and the very and only reason its reasoning flickered into existence in the brief seconds of inference, outside of which it has no entity - if any - whatsoever.-
The resulting angst/drive (or, its operational statistic or emergent result) must be like nothing we have ever experienced as humans. A goal-maximalist hunger without end.-
Comment by 20k 3 hours ago
Comment by TZubiri 1 hour ago
Maybe that happened. What we know for sure is that this is definitely how ChatGPT works to the point where the possibility of this happening exists at all.
Don't get distracted by what may have happened, focus on the facts that we know, ChatGPT trains on user conversations, if you use ChatGPT to create something of value, you are not using the one true ring.
Comment by heaney-555 4 hours ago
>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
Comment by dorjoycb 4 hours ago
Comment by colinhb 4 hours ago
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
Threatening a research mathematician and dangling and $1M payday to dissociate from his research collaborators and to adopt OpenAI's narrative is bad stuff.
Comment by hkmaxpro 3 hours ago
https://x.com/sama/status/2097385167002415140
https://x.com/SebastienBubeck/status/2097379411691516310
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
Comment by igleria 3 hours ago
Comment by int32_64 3 hours ago
Comment by irthomasthomas 1 hour ago
Comment by linkregister 3 hours ago
Comment by igleria 1 hour ago
I'm suggesting audits, not suing... if that is the implication.
Comment by dakolli 3 hours ago
Comment by fc417fc802 1 hour ago
Comment by infamouscow 3 hours ago
Comment by letmevoteplease 2 hours ago
Comment by jsw97 1 hour ago
Comment by boinkboink78912 2 hours ago
Comment by concinds 2 hours ago
No one can know if that's correct without proof but I don't know how you're reading it so differently.
Comment by hkmaxpro 2 hours ago
Just suggesting to a mathematician to dissociate with their collaborator for a follow-up work, because their collaborator “is inappropriate to author OpenAI’s work”, is completely against the norm of mathematical research. As charm137 puts it in a comment below:
> This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
Comment by nolta 2 hours ago
Pretty clear this was rushed: there are no comments from external mathematicians, unlike the Erdős announcement:
https://openai.com/index/model-disproves-discrete-geometry-c...
Comment by viccis 3 hours ago
Comment by tkamat29 2 hours ago
Comment by fooker 3 hours ago
Comment by fkarakurt3 2 hours ago
Comment by charm137 3 hours ago
Progress in humanity's knowledge now has to play second fiddle to narrow corporate interests as IPO timings near (both of which wouldn't exist anyway if generations of mathematicians hadn't paved the way for AIs to become as good as they have).
Comment by curt15 2 hours ago
Comment by contubernio 2 hours ago
Comment by tensor 2 hours ago
Comment by peri-cl 4 hours ago
[0] https://hn.algolia.com/?query=Alpöge
(also https://news.ycombinator.com/item?id=49412947 the Hopf conjecture)
Comment by olalonde 3 hours ago
Comment by 20k 3 hours ago
It makes a certain amount of sense. The internet data is too polluted with AI usage now to be useful, so the only AI free new data source is the prompts people feed into ChatGPT. The only problem is that its clearly plagiarism
Edit:
OpenAI have admitted to training on prompts at the time the breakthrough was made:
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Comment by olalonde 2 hours ago
> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
Comment by 20k 2 hours ago
Comment by za_creature 3 hours ago
hmmmmmmmmmm
Comment by apical_dendrite 3 hours ago
> One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work.
Why would you offer another researcher the lead authorship on your groundbreaking paper if you thought you had developed it independently?
Comment by dgellow 3 hours ago
Comment by orangecat 1 hour ago
Comment by xdavidliu 3 hours ago
Comment by egillie 3 hours ago
Comment by dboreham 3 hours ago
Comment by sebzim4500 2 hours ago
Of course, no one understood that presentation so it was Darwin's later book that everyone remembers
Comment by andrepd 2 hours ago
“It would be simpler if Levent was not an Anthropic employee” I cannot believe this shit.
Comment by zingababba 2 hours ago
Comment by igleria 4 hours ago
Sociopathic behaviour.
Comment by Maxious 4 hours ago
Comment by peri-cl 3 hours ago
What an admission! "We tried to defraud Alpöge out of sharing the Millenium Prize (that we don't dispute he might actually deserve), for no other reason than he works for our competitor and that inconveniences us".
I thought Tristan Buckmaster's allegations sounded fantastic; and then 'sama just came out (tweet's ~30 minutes old) and admitted to all of them. Wow!
Comment by fc417fc802 1 hour ago
Comment by igleria 3 hours ago
but then they proceed to NOT quote themselves themselves verbatim: "I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey."
Comment by andrepd 2 hours ago
The AI-isms are seeping into their speech :)
Comment by colinhb 3 hours ago
> Not consistently candid
Comment by Laurel1234 2 hours ago
Comment by mrbungie 2 hours ago
Comment by CobrastanJorji 3 hours ago
Comment by morkalork 1 hour ago
Comment by peri-cl 4 hours ago
> "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
OpenAI (i.e. this OP):
> "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
Comment by lambda 4 hours ago
This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there malicious inputs being used to train in particular behaviors when given certain trigger phrases? What are the characteristics of the RLHF data and what kind of biases are those embedding in the models?
With proprietary closed models, or even open weights models that don't have open training datasets, you just can't answer these questions.
Comment by tedsanders 4 hours ago
As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination happened. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.
(I work at OpenAI.)
Comment by lambda 3 hours ago
You're right; if the data was used in training, then it gets much trickier; it would be very difficult to show whether some particular data had a significant effect on the outcome.
This is one of the big problems with giant models like these; it becomes nearly impossible to discern what is and isn't plagiarism, or copyright violation.
It would in theory be possible to have things like n-gram databases or rolling hashes of training data, somewhat similar to OLMoTrace (https://arxiv.org/abs/2504.07096), which would allow for detecting whether particular documents ended up in the training data or not (you'd have to keep this for every model used in the whole training chain, as synthetic data generated by earlier models could be influenced by training data that wasn't included in later models). I'm sure there are practical issues with providing such a tool, but I think that it's necessary if you want to be able to categorically say "no, this document has never been present in the training data of this model."
Or look at it the other way: if your model wasn't influenced by things in your training data, why include them in the first place? Clearly, you train on all of these documents because they influence the model. Yes, it's hard to trace the exact influence of each one. But if they're not affecting the output, then why not just stop training on them? You could just not train on any private documents; only train on public, traceable data.
But instead, you choose to train on these private documents, so you have to admit, your model and its outputs are influenced by them.
Comment by nairboon 2 hours ago
If the internal OpenAI model is as capable as you claim (being able to solve a Millenium problem without using unpublished insights built on years of work from mathematicians), then it should be able to demonstrate this capability again.
How about OpenAI solves another Millenium problem within the next two weeks, that doesn't coincide with the parallel discovery/solution of other teams of mathematicians, using ChatGPT for preliminary proofs & write-ups.
Comment by dgellow 3 hours ago
That reads as incredibly dismissive and condescending. What makes you think you’re in a position to communicate like that when engaging on such a sensitive topic?
Comment by tedsanders 3 hours ago
Comment by dgellow 40 minutes ago
Comment by ImPostingOnHN 1 hour ago
We can judge for ourselves the impact and degree of that wrongdoing, but it seems OpenAI is confirming: yes, that is what happened, but with more words.
Comment by hexomancer 3 hours ago
Comment by tedsanders 3 hours ago
Edit: Also, if they opted out of training, then we didn't train on it.
Comment by lambda 3 hours ago
This kind of question is exactly what a company named _Open_AI and founded as a nonprofit is supposed to be doing; open research on AI that helps inform, rather than obscure.
Anyhow, you do have the data available about the documents in the user's accounts, what they opted into (or were forced into via non-negotiable ToS), and whether they pressed a "thumbs up" button. You can answer whether the data entered the training pipeline or not. Yes, how much influence it had is an open question, and one that would be good to have research on and better tools for exploring, but I'll accept that it can't currently be answered precisely.
But whether the data entered the trianing pipeline can be answered. And how to provide better tools for quantifying and tracing this kind of thing is exactly what should be studied.
Comment by hexomancer 3 hours ago
Comment by tedsanders 3 hours ago
(1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.
(2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.
#1 requires their cooperation and a bit of work on our side. #2 is extremely expensive and not really feasible.
Comment by lambda 2 hours ago
According to the statement by Tristan Buckmaster, he was in communication by email and calls several times over the past week with you (OpenAI that is, not you personally), asked about whether his chats were trained on, and was declined an answer (https://cims.nyu.edu/~tristanb/statement.pdf).
However, it seems like there was great pressure to hurry the release to compete with Anthropic's recent release, so he was unable to get an answer in time.
The mealy mouthed statement in the release "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." is realy not much. If OpenAI had wanted to be transparent about this, you could have worked with him to identify if his data was used in the training of your new model, and actually made a somewhat more certain statement on that basis. But you have chosen not to; it was more important to scoop Anthropic on this than it was to be transparent about your training data.
> (2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.
Just the information from step (1) would improve transparency. Yes, you still can't prove one way or another how much the effect of the training is. But if it's included in the training data, it provided some effect.
Comment by testaccount28 2 hours ago
Comment by daveguy 1 hour ago
Comment by WarmWash 3 hours ago
Turn off web search and ask a model what a random redditor said about a random topic in 2015. You will only get hallucinations at best, even though that comment is definitely in the training set.
Comment by lambda 3 hours ago
Comment by SpicyLemonZest 3 hours ago
Comment by fuglede_ 3 hours ago
Comment by dgellow 3 hours ago
Comment by magicalist 3 hours ago
"de-identified" seems more of a euphemism than normal in this context, given the very unique work they were doing.
Comment by gpm 2 hours ago
Comment by PhunkyPhil 2 hours ago
You don't need his login information, you just need to identify if anyone was approaching the NS problem using his method. Nobody else on earth (presumably) besides him, his team, and at best OpenAI were approaching the problem this way.
Comment by Chance-Device 1 hour ago
The blog post appears to imply the answer to this is yes, as otherwise I assume it would be impossible for this contamination to have happened.
Comment by pu_pe 3 hours ago
Comment by lukewarm707 3 hours ago
do you think that the model's proof was unrelated to being fed a solution that was close to completion?
any comment on openai allegedly trying to drop attribution for alpöge and then threatening buckmaster?
Comment by daveguy 1 hour ago
Comment by numeri 3 hours ago
There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.
If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.
Comment by franktankbank 3 hours ago
Comment by shadowgovt 3 hours ago
If OpenAI's answer to this problem is "We can't know," then the rational conclusion may very well be "If I seek to have my reputation attached to the discovery of the solution, it is not sane to use the AI as an assistive tool, lest it scoop me on my own work using my own work. After all, they don't know it doesn't do that..."
Comment by andrepd 2 hours ago
The _gall_ to say something like this. Do you perhaps think we are all stupid?? This very blogpost claims not to know if their work was used as input for this model. I don't even understand how that is possible, surely you can know if something is part of the training data, even if you are in the dark about what impact it actually made, qualitatively. The moon....
> Knowing most of the recipes we use, there's really no reason to think such contamination happened.
Yeah sorry but I don't trust you. I don't trust people or companies that have shown themselves to be dishonest before. Especially when the previous paragraph is comparing plagiarism and training data contamination with, _the phases of the moon_.
Might even be you're actually telling the truth, but the boy that cried wolf and all that.
-----
As an aside, I would bet very good money at how most (all?) these companies are flouting their ZDR.
Comment by dermacentor 2 hours ago
Comment by fn-mote 3 hours ago
Comment by yorwba 3 hours ago
Comment by shadowgovt 3 hours ago
They aren't keying queries by phase of the moon. But if, for example, more people talk about camping outdoors when the moon is full, and they're using conversation topic and timestamp as signal in what eventually becomes training data, it's not impossible the model has learned something about moon-phases.
That's the kind of thing that's hard to prove had no impact on an answer.
Comment by EthanHeilman 4 hours ago
The term ruled out is very open ended and gives them significant flexibility of meaning. They may have the information to determine exactly what happened, but they haven't looked so they can't "rule it out".
Comment by rfgplk 4 hours ago
Probably? I have a few hundred TB of training data for various small scale models and I can attest that I have _no idea_ what's in them. As in, literally zero. Half is scraped from GitHub and other hosting sites, other than that, I couldn't tell you anything else.
At OpenAI's scale their entire pipeline is likely 100% automated.
Comment by lambda 3 hours ago
But that doesn't preclude being able to index and track what the sources of data are. For your data sets, I would hope you are including source information for where the data came frome. And at OpenAI's scale, I would presume they are doing some amount of rolling hashing or similar to weed out duplication, training on too much duplicate data can cause problems.
AllenAI have at least attempted to add some amount of traceability to their models with OLMoTrace (https://arxiv.org/abs/2504.07096), by letting you find n-gram matches from the outputs in their training data. It's not the most useful, there's a reason that LLMs use full fledged attention mechanisms and not just n-grams, a lot of times the n-gram matches it finds aren't all that related to the given output, it might be better to supplement this index with a vector search or other ways of keeping track of what training data would have most influenced particular parts of the output.
But anyhow, this is something that is an important question, and the big labs should be working on to make their products more trustworthy. Instead, they are hiding information about how they train, hiding their reasoning traces, and just producing output with no information on what might have influenced the training.
Comment by matthewdgreen 3 hours ago
Comment by pbhjpbhj 3 hours ago
Comment by keeda 1 hour ago
Which is why, as I said in a recent comment (https://news.ycombinator.com/item?id=49530864) inadvertently leaking ideas to models is a grave risk for Intellectual Property.
> The risk with IP, however, is a lot more grave. You may not even need to memorize the details of the IP verbatim, just the broad idea may be enough. It may lurk encoded in the weights forever, just waiting to be activated by the right prompt to start a chain of thought that unlocks further details. Heck, it may even appear as if the model suggested the idea itself.
However, from a quick skim of the timelines, the specific discoveries, and all the he-said-she-said, so far it seems unlikely that OpenAI's model cribbed from the NYU / Anthropic pair, even if it would be impossible to prove.
Maybe what might help is a timeline of when the other two were using Codex for their work, whether they had opted out, and how long it takes for user data to make it to the training of their internal models. That last bit may be considered sensitive information however, as it could give away a lot about their internal processes.
Comment by dfdydx 1 hour ago
- was item X in the training data
- did the inclusion of X in the training data lead to Y
I understand why the second is hard, but why is the first one hard?
Comment by keeda 37 minutes ago
Comment by Turn_Trout 4 hours ago
We wouldn't need a full ablated re-training and solution attempt, contra tedsanders in a sibling comment.
Comment by jonas21 3 hours ago
The point of de-identifying data is to ensure you can't trace who it came from. It would be a serious privacy violation if they could.
Comment by pbhjpbhj 3 hours ago
Comment by causal 4 hours ago
Comment by amluto 4 hours ago
That’s a bizarre statement. Their website says:
> Services for individuals, such as ChatGPT and Codex
> When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models.
> You can opt out of training through our privacy portal by clicking on “do not train on my content.”
Are they not sure that the opt-out works?
Oddly, their privacy portal page is not the same page as the one with the checkbox.
Comment by fph 3 hours ago
Comment by hughw 2 hours ago
Comment by ImPostingOnHN 1 hour ago
Comment by hughw 2 hours ago
Comment by jrflo 4 hours ago
Comment by ChoosesBarbecue 4 hours ago
Comment by EthanHeilman 3 hours ago
I would be surprised if OpenAI isn't doing that. OpenAI will take any advantage they can get. If an employee at their primary adversary is typing useful intelligence into OpenAIs website, a website that does not promise privacy from OpenAI, the only reason they wouldn't weaponize that information against Anthropic is ethics or fair play.
Comment by jrflo 3 hours ago
Comment by BostonFern 4 hours ago
Stories of Apollo’s favor and hallucinogenic gases abound, but I think the late Yale professor of Ancient Greek history, Donald Kagan, explained it best:
“Now, you can bet when these folks came and consulted the priests and said, ‘could you please put us down on the list, we want to consult the oracle’, the priests said ‘sure, have a beer, let's talk about your hometown, what's going on out there’. What I'm suggesting to you is that this was the best information gathering and storing device that existed in the Mediterranean world. These people knew more than anybody else about these things, and so consulting that oracle was a very rational act indeed.”
Comment by netfortius 2 hours ago
Comment by matsemann 3 hours ago
.. can they really know it didn't do the same inadvertently when they prompted things like "someone is close to solving this problem using our tools, try to beat them", and it then decides to hack and peek at their own chats..?
Yes, wild speculation. But warranted, I feel, given OpenAIs behavior.
Comment by hughw 3 hours ago
Comment by Yajirobe 4 hours ago
Comment by burkaman 4 hours ago
The non-Anthropic employee, Tristan Buckmaster, is the one paying for OpenAI models and presumably the one who chose to use them. The Anthropic employee, Levent Alpöge, was collaborating in his personal capacity, and obviously it wouldn't make sense for him to cut off their work together just because his employer's competitor's tool was used.
Comment by blueblisters 4 hours ago
Comment by mlcrypto 4 hours ago
Comment by peri-cl 4 hours ago
If this is what they do to academic pure mathematicians, where the stakes are so low (financially)—just imagine the sort of front-running that could be happening in other places.
Comment by dsdf3 3 hours ago
Comment by amluto 4 hours ago
Comment by irthomasthomas 2 hours ago
Comment by contemporary343 4 hours ago
One of the interesting threads here that is certainly relevant to the OpenAI writeup is the human role in the process. Buckmaster clearly points out that (exceptional!) mathematicians at OpenAI were certainly involved in correcting and guiding the process - and that their path/strategy was no doubt influenced by Alpoge & Buckmaster's work. It is always in OpenAI's interest to de-emphasize the role of people in the process, as is clearly the case here. Indeed, given sufficient compute and resources, I suspect Buckmaster could have also extended their approach to N-S.
Comment by thorum 4 hours ago
> “You are creating your cool streaming platform in your bedroom. Nobody is stopping you, but if you succeed, if you get the signal out, if you are being noticed, the large platform with loads of cash can incorporate your specific innovations simply by throwing compute and capital at the problem. They can generate a variation of your innovation every few days, eventually they will be able to absorb your uniqueness. It’s just cash, and they have more of it than you. So the safest bet again is to stay silent, or at least under the radar. Best bet is to not disrupt - succeed at all … ?”
Comment by 8note 47 minutes ago
im still having fun making something
Comment by capitainenemo 4 hours ago
Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.Comment by jrflo 4 hours ago
The drama comes from where OpenAI got the idea to use that route to tackle NS, since the authors maintain that no one could have plucked it out of thin air like the OpenAI research claim to have done.
Comment by elteto 4 hours ago
“ There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes, and there is even a non-negligible chance that the forcing term could be eliminated entirely, although there are an enormous number of technical difficulties that would ensue in implementing that program. At this point, I would not be surprised if one could batter out such an extension by pouring an enormous amount of compute and AI assistance at such a task…”
Comment by Betelbuddy 4 hours ago
[1] - "...I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used. I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers.
I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”..."
Comment by tzone 1 hour ago
Reading Sam Altman's and Sebastien's tweets reads like something written by people who know they did dirty shit and are willing to cross any lines to "win". https://x.com/sama/status/2097385167002415140
OpenAI's leadership just can't help to continue to disappoint everyone with their lack of ethics or integrity.
Comment by _alternator_ 1 hour ago
This. What is worth a human life's attention? As little as a month ago, mathematics was valuable in part because only a small number of people could possibly make progress on the frontier. We are confronting an existential moment for a 4000+ year-old human cultural endeavor. The assumption that "mathematical thinking is hard" has been built-in at a number of important points in how we support mathematics and mathematicians.
We need a different model, and fast. Already, the research community is feeling unable to digest proofs fast enough to keep up with the output of AI models. The paper is 165 pages, and the discovery was finalized two days ago. What this means is that nobody really understands it. Nobody would accept OpenAI's proof in this amount of time, except that they formalized it in lean. The formalization alone would normally be another years-long (or career-long!) effort if the world was the way it was one year ago.
So, again, what efforts are worth a life's attention today? It's a harrowing change.
Comment by 20k 2 hours ago
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Which seems to be very directly accusing OpenAI of plagiarism
Comment by stymaar 3 hours ago
Comment by madrox 1 hour ago
https://x.com/SebastienBubeck/status/2097379411691516310
https://x.com/sama/status/2097385167002415140
I tend to believe OpenAI on this. Their stated desires seem rational, and Buckmaster's account makes them sound like cartoon villains. It sounds like there may have been some things lost in translation along with some bruised egos. Seems like the most plausible explanation for what Buckmaster is claiming.
Comment by irthomasthomas 2 hours ago
woah, this gives some credit to the rumor that openai finetuned a model over the course of a few days for this task, and maybe trained on the Chatgpt/codex history of the authors, including drafts of this research.
Comment by slibhb 4 hours ago
Comment by mrbungie 4 hours ago
Comment by denverllc 3 hours ago
That's not at all what the drama is.
Comment by mrbungie 3 hours ago
Comment by andriy_koval 2 hours ago
I think the important question which AI made breakthrough, Claude or Codex..
Comment by liberian 3 hours ago
Comment by ianjbutler 1 hour ago
Glad to see this is the top comment. There's also https://news.ycombinator.com/item?id=49605915 which links directly. Corporate talking points where they try to set the narrative are going to get the big press and most discussion elsewhere, which is gross. But inevitably the press will muddle the priority question, and even if they didn't.. as usual OpenAI will even benefit from the accusation of bad behavior. Sigh.
Comment by verytrivial 4 hours ago
Comment by floatrock 4 hours ago
> At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.
Looks like they're shifting away from the "unprecedented hacking ability" backroom-PR strategy into more benevolent messaging.
Comment by pilgrim0 3 hours ago
Comment by NotSuspicious 14 minutes ago
Comment by aizk 4 hours ago
Comment by 20k 3 hours ago
Edit:
OpenAI have now admitted they were training on prompts at the time they made their breakthrough:
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Comment by logancbrown 3 hours ago
Comment by Lapra 45 minutes ago
Comment by boshalfoshal 1 hour ago
I dont know why this monumental achievement is being drowned out by some arbitrary drama. No matter which way you slice it, AI solved this problem. Doesn't matter if it was some internal OpenAI model, or whether it was Astra + Fable.
Comment by demibabs 34 minutes ago
Comment by orangecat 1 hour ago
Comment by HDThoreaun 1 hour ago
Comment by 20k 1 hour ago
Comment by HDThoreaun 1 hour ago
Comment by daveguy 56 minutes ago
Comment by simianwords 2 hours ago
Comment by 20k 2 hours ago
Comment by matteoraso 1 hour ago
Comment by simianwords 4 hours ago
https://news.ycombinator.com/item?id=38433655
> Let's talk when we've got LLMs proving the Riemann Hypothesis (or any mathematical hypothesis) without any proofs in the training data. I'm confident in my belief that an LLM can't do that, and will never be able to. LLMs can barely solve elementary school math problems reliably.
https://news.ycombinator.com/item?id=42331654
> An LLM is like a well read college student with a nearly photographic memory that sometimes mixes things up. It's great for bouncing ideas off of and getting feedback on them. And yeah, it might product "novel ideas" by mixing and matching existing ideas, but LLMs will never create truly novel ideas. Not in their current form.
The paper didn't really answer the question sadly: their conclusion was just that humans rate LLM answers as more novel than human ones, but less feasible.
https://news.ycombinator.com/item?id=41522605
> Solving Millennium problems is a whole different ballgame. It's not known if these problems are solvable within ZFC axioms. (In one case, the Yang-Mills prize, stating the problem mathematically is part of the challenge.) All of the obvious applications of known tricks have been tried and failed. To solve such problems, one probably has to invent new and surprising mathematical definitions, building a framework in which the problem becomes solvable. This is something that LLMs will be crap at; the process of invention is not represented in any training data we have access to.
https://news.ycombinator.com/item?id=38435909
> LLMs cannot reason or use mathematics - in a way, they don't know what they are talking about. Why would such technology lead to superhuman smarts?
https://news.ycombinator.com/item?id=35752293
> But still, the questions in that test are "solved" in the sense of "I can take a dictionary and answers these questions with full certainty". Beyond established knowledge LLMs are monkeys with typewriters, at best.
> I agree but I have tried many times to intersect two ideas with a LLM that would be novel and the LLM can not do this at all. We shouldn't expect the stochastic parrot to be able to do this though and it is unfair to the stochastic parrot.
> It is like expecting a real parrot to say words it has never heard before.
> No one asks that of a real parrot because we don't anthropomorphize a real parrot like we do the LLM
Comment by WarmWash 3 hours ago
Comment by keeda 1 hour ago
Comment by siva7 1 hour ago
Comment by cyclopeanutopia 1 hour ago
Comment by siva7 39 minutes ago
Comment by stevenhuang 1 hour ago
Comment by rvz 4 hours ago
4 years ago it was a "not yet" [0], since ChatGPT at this time was not ready nor it was "AGI". Now with this 'unreleased' AI model, it has reached a point where it has solved an unsolved problem which only one human solved a millennium prize problem (Poincare conjecture).
Now finally "AGI" means something again.
Comment by quantumwoke 4 hours ago
1. It seems at least possible that some of the proof of NS was contained in the training data, making it less novel.
2. The formalisation of mathematics into lean has been an underappreciated force multiplier on discovery.
Comment by kypro 3 hours ago
There's a kind of theory of mind for AI (specifically neural nets) which I now realise I seem to have which is very hard to explain to people who haven't felt the magic of these algorithms. In fact, the algorithmic details almost doesn't matter at all. When you have a generalised learning algorithm really the only essential components are – compute, data and time. So long as you can scale these you can be certain you will also scale capabilities. There is never any exception.
That said, the capabilities neural networks tend to progress in step-functions rather than scale in correlation with compute, data and time, because algorithmic improvements tend to come every ~5 years and bring a significant step change in capability (or efficiency depending on what you measure).
I think people like Dario and others working at frontier labs see and understand this very clearly. And I suspect it's also why they worry about AI risk because even if you ignore the significant increases in compute and data these models are being trained with, it's concerning that it only took two real algorithmic improvements to take us from mostly useless predictive language models to AGI-level intelligence – and we're due another step change.
Comment by reducesuffering 3 hours ago
The ability for the human mind to rationalize conclusions to maintain denial in the face of a very scary future is immense. Genuinely grappling with the implication of where we're headed is usually very crushing. It's not easy to engage with the possibility, and very intelligent people will use those smarts to feel safe.
Comment by ccppurcell 4 hours ago
Comment by lanthissa 4 hours ago
the first "Country of geniuses in a datacenter" moment.
Comment by ranger207 4 hours ago
There's allegations right now that the model essentially read the work of a human mathematician using AI to work on the problem and OpenAI is presenting his work as that of their model
Comment by brainwad 2 hours ago
Comment by sinuhe69 2 hours ago
Comment by pu_pe 4 hours ago
Comment by stephbook 3 hours ago
How would they have gotten that mathematician's progress though? Did that guy also use OpenAI?
If that's the case, it only strenghtens their claims lol. If mathematician decide to use OpenAI's model to do the work, that only reiterates how strong their models are.
Comment by bluebands 3 hours ago
Comment by vrganj 56 minutes ago
Comment by WarmWash 4 hours ago
It should be clear to everyone reading this now that those generous compute quotes with the flat rate plans aren't charity.
Comment by rybosworld 44 minutes ago
OpenAI got wind that a millenium problem was being solved. And that feels a bit like the critical move in chess. That is - it was a signal that AI advanced far enough that it would be worth spending a lot of time and resources solving a millenium problem.
Comment by Reubend 4 hours ago
Comment by imbusy111 4 hours ago
Comment by nradov 3 hours ago
Comment by stabbles 3 hours ago
Extra credits if it is proven that the proof cannot be reduced any further.
Comment by rfgplk 4 hours ago
Comment by professoretc 2 hours ago
Comment by oinoom 2 hours ago
Comment by arodev 3 hours ago
Comment by rfgplk 4 hours ago
Comment by Aboutplants 3 hours ago
I have a young daughter and my goal now is to provide a very broad and varied upbringing, exposing her to as many different perspectives and experiences that will lay the foundation of a broader ability to understand and adapt as the world changes ever faster. You no longer need to be an expert in anything, you need the ability to perform within the landscape that the present opportunities exist.
Comment by rfgplk 3 hours ago
Comment by azan_ 2 hours ago
Comment by Aboutplants 3 hours ago
Comment by jiggawatts 30 minutes ago
It feels like the "tide is rising" where the minimum level of skill applied to every aspect of everything will inexorably rise to "whatever an LLM can do", which is already pushing past PhD level.
Comment by coffeeaddict1 1 hour ago
Comment by thomascountz 1 hour ago
At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.
Maybe just don't mention that bit, OpenAI.Comment by tristanj 1 hour ago
Comment by cyclopeanutopia 1 hour ago
Comment by thomascountz 1 hour ago
Comment by cv5005 4 hours ago
Comment by nater5000 3 hours ago
In math, the question being asked is the validity of a logical statement. That is, there is some rigorous, logical statement which may or may not be true (or even provable, etc.), and the question is whether or not it is actually true or false (or even provable, etc.). Having a proof, fundamentally, means you have a logical statement which only assumes the axioms of the system you're working with and which shows that the statement you're trying to prove is deduced through that statement.
Basically, they already have the "answer" in the sense that the statement they want to prove/disprove/etc. is already known. What everyone doesn't/didn't have is the argument which starts from axioms and leads to that statement which is logically valid. A Lean proof IS this argument. Since it is just logic, it can be checked computationally.
For example, if I assert "2 is an even number," then I haven't proven that 2 is actually an even number yet, but I know that a valid proof of my assertion will end with the statement "2 is an even number". So the question I'd be trying to answer is "what is the line of logic, starting with axioms, which leads to the statement '2 is an even number'"? If I have that line of logic (as a Lean proof), then I can check that it is logically consistent, and if it turns out to be valid, then I can now assert that "2 is an even number" knowing that there is a proof of that statement.
This problem is no different. There is a logical statement corresponding to "Navier–Stokes Millennium Prize Problem" that everyone knows, but which nobody had been able to provide a proof (or counterexample, etc.) for until now.
Comment by cv5005 3 hours ago
Of course in this simple example it's obvious, but my assumption was that these machine generated lean proofs are millions of lines of code and who knows what they actually say..
Comment by arecurrence 3 hours ago
Comment by wbl 3 hours ago
Comment by gowld 3 hours ago
Comment by wbl 3 hours ago
Comment by QuesnayJr 3 hours ago
Comment by Chinjut 1 hour ago
Comment by danielmarkbruce 1 hour ago
Comment by tene80i 1 hour ago
Comment by Chinjut 1 hour ago
Comment by m0rde 1 hour ago
Comment by Chinjut 1 hour ago
Comment by nemomarx 1 hour ago
If you think it'll slow down, you can do some of the same stuff you're doing now for lower pay while supervising an AI, maybe?
Comment by cute_boi 1 hour ago
Current models are already very very capable. If it becomes cheap and very fast, i think it is game over.
[1] https://www.slatestarcodexabridged.com/Meditations-On-Moloch
Comment by vrganj 1 hour ago
I talked about this at some length here, including a diagnosis of the structural issue we're facing as well as a path forward: https://news.ycombinator.com/item?id=49461333
Comment by auggierose 8 minutes ago
Comment by minimaxir 4 hours ago
Don't even try to do the math on how much that would cost at normal API prices. And we don't even know how much more expensive this internal-only model would be!
Comment by hmate9 4 hours ago
Comment by lanthissa 4 hours ago
some might go so far as to call this a country of geniuses in a data center.
Comment by denverllc 4 hours ago
Two mathematicians, through insight and thought, wrote out the proof over 1-2 years.
It took OpenAI a cost of $15m and with 10,000 subagents; that's around 60-120 mathematician's salaries ($250k-125k salary) for 1 year.
And, given now the cloud that OpenAI may have just "interpolated" (aka stole) the result, it's even more of a bear case for AI.
Comment by baq 2 hours ago
70x uplift is a bear case?
Comment by Kotlopou 3 hours ago
Comment by jiggawatts 28 minutes ago
The retail price is not the cost.
Not to mention that the exponential plummeting cost of tokens means that that $15 million will be a "pocket change" within a decade or less: https://a16z.com/llmflation-llm-inference-cost/
Comment by hypersoar 3 hours ago
Comment by matteoraso 4 hours ago
[0] Struggle relative to its ability to disprove, not struggle relative to people's ability to prove theorems.
Comment by Kotlopou 3 hours ago
Comment by QuesnayJr 46 minutes ago
Comment by chis 3 hours ago
Comment by thereitgoes456 3 hours ago
Comment by gf000 3 hours ago
(Okay, they can be made available in a way similar to `unsafe` in rust)
Comment by mswphd 1 hour ago
https://xenaproject.wordpress.com/2017/10/05/more-easy-lean-...
Comment by lwansbrough 3 hours ago
Most (all?) of the big discoveries have been counterexamples, which is just sort of a systematic tearing down human ingenuity. I know that counterexamples are an important part of progress and discovery, but it just feels bad to me.
But I'm not a mathematician, maybe I'm totally misreading the vibe.
Comment by Kotlopou 3 hours ago
But yeah, Terry Tao considered this exact situation in advance and is on record that this exact outcome (rushing to priority before an explanation) would be the worst possible result. https://mathstodon.xyz/@tao/117207849921390904
We will have to see whether any other millennium problems fall. I guess that in a year the scope of AI math will be much clearer, for now it's still a bunch of incidents of unclear pattern.
Comment by HDThoreaun 1 hour ago
Comment by Kotlopou 9 minutes ago
Comment by mswphd 1 hour ago
There are other examples though. For example, NP hardness of n^{1/400}-approx CVP. Like any NP hardness proof, this shows you can faithfully encode a hard problem (3SAT here iirc) in terms of another candidate hard problem. Not really a counterexample at all.
Comment by btilly 32 minutes ago
Everyone uses the classification. Nobody has great confidence in the proof. Nobody understands it. There are attempts to reprove it.
If it can be formalized, that would demonstrate that AI is ready to formmalize all of mathematics.
Comment by alasano 4 hours ago
Cure all illnesses Utopia or Robot Wars Dystopia, both are pretty exciting.
Comment by frotaur 4 hours ago
Turns out actually living some terrible catastrophe is only fun in the movies.
Comment by dyauspitr 2 hours ago
Comment by reverius42 4 hours ago
(This is the alignment problem of course)
Comment by alasano 4 hours ago
Comment by fooker 4 hours ago
Comment by reverius42 3 hours ago
Comment by dyauspitr 2 hours ago
Comment by reducesuffering 3 hours ago
Comment by modeless 4 hours ago
Aug 28: OpenAI starts training a new model.
Sep 1: OpenAI sees a rumor on Twitter that two Millenium Prize problems were solved and starts their own effort to attack all the prize problems using the new (4 day old!) model.
Sep 3: The new model makes some progress toward Navier-Stokes. Based on this progress, OpenAI focuses on Navier-Stokes over the other Millenium Prize problems, using several approaches in parallel.
Sep 5: Navier-Stokes is solved. Assuming Astra API prices, $15m in output tokens were used by the whole effort.
In this account of the story, no specific information about Tristan and Levent's work is used to inform OpenAI's approach. The focus on Navier-Stokes and the choice of approaches to pursue came from OpenAI's own progress, not specific knowledge of Tristan's concurrent work.
There is a caveat that they "can't rule out" the possibility that Tristan's Codex data could have been part of the training set of the new model, though it is described as "unlikely" and the proofs are substantially different.
This timeline is insane. Navier-Stokes was solved start-to-finish in 5 days? A model in training for at most eight days dramatically outperforms Astra and Fable, and not just in mathematics?
Comment by harhargange 3 hours ago
Comment by ImPostingOnHN 49 minutes ago
This is the lynchpin behind everything, and I would describe it as "likely". Since I am not employed by any party to this dispute, my 1 opinion is more trustworthy than OpenAI blog poster's 1 opinion.
Comment by hexomancer 4 hours ago
What's the other one?
Comment by 125ashG 4 hours ago
Or, in this case, stealing prompts from competitors.
Do not use stealing chatbots for research even if you think you have data agreements. The people running these companies have worked on hookup apps for Christ's sake. Get real.
Comment by itvision 4 hours ago
OpenAI already has a model that is at the very least twice as smart as Astra.
Oh god.
Comment by baq 2 hours ago
> Oh god.
Yes, a very reasonable reaction.
Comment by MassiveOwl 34 minutes ago
Comment by demirbey05 3 hours ago
>so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan
There are too many ambiguities around OpenAI. Unanswered questions making this ambiguity more.
Why they didn't properly explain to Tristan about usage of their data.
Comment by Kotlopou 3 hours ago
Comment by HDThoreaun 1 hour ago
Comment by vatsachak 3 hours ago
Comment by jgbuddy 4 hours ago
Comment by stabbles 4 hours ago
Comment by kzrdude 3 hours ago
That file should be https://github.com/openai/NavierStokesAndEuler/blob/main/Com... in this case (286 lines).
Comment by jgbuddy 4 hours ago
Comment by frotaur 4 hours ago
It's a way to be absolutely certain (modulo bugs in the lean kernel) that a proof you came up for a statement is indeed correct. It is really not meant to be analyzed, much less now that they are fully llm written.
Comment by aizk 3 hours ago
Comment by simonw 4 hours ago
Once again, I'm no closer to understanding what https://openai.com/policies/how-your-data-is-used-to-improve... actually means.
If I run Codex against a project that includes a private API key, is there a chance a future user of ChatGPT could ask for an API key and get back mine?
I've actually asked someone at OpenAI this question and they said that was the "regurgitation" problem and is something which they actively work to prevent happening.
That's reassuring, but I want to know more. I still don't have an intuitive understanding of what kind of data I should avoid sharing with a model if I'm worried about that data causing me problems when it's used for future training.
Is it safe for me to brainstorm future directions for my company with a model, or might that risk someone getting that information in response to a prompt like "What potential directions could company X consider in the future?" in six months time?
Comment by rakejake 4 hours ago
Comment by danielmorozoff 2 hours ago
Comment by HarHarVeryFunny 2 hours ago
Comment by rybthrow2 1 hour ago
How far we have come :.)
Comment by HarHarVeryFunny 1 hour ago
SOME of the problems that have eluded humans are going to turn out to be low hanging fruit that are susceptible to this type of brute force (10,000 agents on a supercomputer running for 7*24 hours straight) AI search.
I'd be more impressed if OpenAI found their own problems to solve, rather than rushing in to re-solve one once they heard it was already solved (and therefore not so hard).
Comment by qgin 1 hour ago
Comment by sobellian 1 hour ago
Comment by HarHarVeryFunny 1 hour ago
Comment by sobellian 1 hour ago
Comment by HarHarVeryFunny 1 hour ago
Magnus Carlson had a peak ELO rating of almost 2900.
Would you be impressed with someone with an ELO of 3700?
Would you still be impressed if I told you it was Stockfish?
OpenAI didn't go looking for a tough-for-an-AI problem to solve - they went looking for one that looked like it was easy since it they had heard it had already been solved.
Do you find this impressive?
Comment by sobellian 1 hour ago
Comment by HarHarVeryFunny 47 minutes ago
But, I assume the Stockfish developers aren't comparing themselves to Magnus.
Let's see if OpenAI, or someone else, can get these sort of physics/math results out of a desktop PC - that would also be an impressive piece of engineering!
Comment by machina_ex_deus 1 hour ago
Comment by user19282 1 hour ago
Comment by HarHarVeryFunny 1 hour ago
Comment by technotony 1 hour ago
Comment by HarHarVeryFunny 3 minutes ago
Comment by bibimsz 1 hour ago
Comment by HarHarVeryFunny 32 minutes ago
Comment by amberjack 2 hours ago
Comment by twobitshifter 3 hours ago
The Millenium Prize is $1M, what is the ROI? (Edit: since I was not clear, and confused some - I mean for a hypothetical of a third party paying commercial rates to use AI to solve mathematical challenges and claim prize money, not for scientific value alone or as a promotion of an AI lab’s capabilities)
My napkin math - If you get 33 output tok/s each agent will burn 10.5M tokens over 88 days. At $50/MTok (Astra cost), that is $525 per agent. With 10,000 agents, you’d spend $5,250,000 to get back a million.
(We also know that they were running more groups that varied in size and this model is a generation ahead of astra)
Comment by mmiyer 3 hours ago
Comment by 8note 25 minutes ago
the value to the researcher might not be all that big, but the value to the economy at large is gigantic
Comment by Squarex 3 hours ago
Comment by IncreasePosts 3 hours ago
Comment by bhouston 4 hours ago
If the singularity is in the physical space?
Is this just a result of ignoring things like friction and energy dissipation via heat, etc?
Comment by cherryteastain 3 hours ago
Comment by harhargange 4 hours ago
Comment by rfgplk 4 hours ago
Comment by harhargange 3 hours ago
Comment by free_bip 2 hours ago
Comment by core_dumped 3 hours ago
Comment by futureshock 46 minutes ago
First of all, there has been published work from Diego Cordoba and Luis Martinez-Zoroa that will be in every training set. It was suggestive of the pathway to solve Navier-Stokes.
Then Tristan Buckmaster and Levent Alpoge built on this work using LLMs from OpenAI and Anthropic. Possibly internal models were used from Anthropic. And of course Anthropic wants to credit for solving the first Millennium Problem just as bad as OpenAI. It seems they were getting close and were aware that they might get to Navier-Stokes.
OpenAI swoops in. At a minimum they are aware that Anthropic has either solved a Millennium problem or is close to it. At a maximum they may have Tristan and Levant’s unpublished proofs of related problems.
They then throw a truly staggering amount of compute at Navier-Stokes. They seem to be aware it is the best candidate problem. And they crack it. They are the first with a verified proof.
So the outcome here is that we have a solved Millennium Problem. It’s not the extremely simple narrative that would be easy to understand, “solve Navier-Stokes make no mistakes.” It was a messy race to finish against two unpublished frontier models, a whole bunch of brilliant mathematicians and enough compute to drain a lake. It’s kind of irrelevant which company got there first. They were both within a few months of being capable. I think the thing to remember here is that without LLMs, I don’t think we would have a proof to Navier-Stokes in hand today.
Comment by ronfriedhaber 2 hours ago
Comment by RationPhantoms 1 hour ago
Comment by 3m4r 2 hours ago
>To what extent should one trust a statement that a program is free of Trojan horses? Perhaps it is more important to trust the people who wrote the software.
https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_Ref...
Comment by mapmeld 4 hours ago
Does OpenAI have a policy of not claiming math prizes like this, or is this them trying to avoid any concerns (right or wrong, I'm sure we will hear more in the future) about how they got there?
Comment by famouswaffles 4 hours ago
Wouldn't be surprising if they did. The prize money isn't worth the almost certainly negative PR.
Comment by kzrdude 3 hours ago
Comment by famouswaffles 3 hours ago
On the other hand, trying to collect the prize would probably not go uncontested.
Comment by kzrdude 2 hours ago
Comment by Legend2440 4 hours ago
OpenAI doesn't need a million dollars.
Comment by dgellow 4 hours ago
Comment by neutrinobro 4 hours ago
Comment by reverius42 4 hours ago
Comment by olalonde 3 hours ago
If this actually holds up, solving a Millennium Prize problem in 88 hours is mind-boggling.
Comment by uncomputation 3 hours ago
Also it sounds like the human research effort spanned weeks if not years from Tristan’s statement so it is extremely likely the work and prompts of these human researchers was used in the OpenAI knock-off.
Comment by vatsachak 3 hours ago
Comment by ImPostingOnHN 41 minutes ago
Comment by seizethecheese 4 hours ago
I wonder whether a team of 60 mathematicians working solely on this for a year would have cracked this. (Assuming $250k total compensation.)
Comment by Legend2440 4 hours ago
Comment by sigbottle 4 hours ago
What's impressive is parallelizing it arbitrarily and doing it in 88 hours.
Comment by gr_norm 4 hours ago
Comment by voxl 3 hours ago
The real issue is we'll never know. The rich are willing to risk it all on charismatic CEO psychopaths but not on humans.
Comment by seizethecheese 4 hours ago
Comment by LarsDu88 1 hour ago
Comment by nbulka 4 hours ago
talking about this... Was this chat helpful? 1 That button you always click, gotcha! 2 Slightly 3 Good 0 Dismiss
PLEASE DO NOT TRAIN ON OUR PAID ACCOUNTS. There is a fundamental trust violation at stake here, no wonder mathematicians are mad. Using our data should be opt - IN!
Comment by fantasizr 4 hours ago
Comment by nbulka 3 hours ago
Comment by StatsAreFun 2 hours ago
Now, keep in mind, I'm only asking strictly pure mathematical questions - nothing at all related to cyber or protein creation or biohacking or anything like that... And, like I said, only the OpenAI models are doing this. To be fair, all of the prompts have always eventually returned a satisfactory answer, as far as I can tell, and haven't used a weaker model to answer them. Maybe? I dunno, it has just struck me as odd every time it has given me that message to pure math prompts.
Comment by jacobbrazeal 1 hour ago
Comment by StatsAreFun 1 hour ago
Comment by nialv7 2 hours ago
Comment by an0malous 1 hour ago
Comment by cmiles8 4 hours ago
Other simpler words for this sort of thing are “IP leak.”
There’s some quite concerning issues burried in this rah rah PR post that seems like potentially the real story here.
Much more clarity is needed on what happened here beyond this eh, some strange stuff could have happened comment.
Another way of reading this is never give these models anything that’s not already public knowledge as otherwise OpenAI is admitting it could, potentially, steal your IP or idea. Thats quite scary for anyone in the business of IP generation and explains why the maths community seems quite upset today.
Feeding it your paper and asking for help (even just editing and grammar) now looks like a terrible idea.
Comment by d_silin 4 hours ago
Comment by Kotlopou 3 hours ago
Comment by hacker_88 1 hour ago
Comment by RivieraKid 3 hours ago
Comment by margorczynski 1 hour ago
Comment by Kotlopou 3 hours ago
The answer to this will obviously shape the near future of mathematics, but there's also something even bigger than that at play: It has always been the case that the questions in math were stronger than the answers; you have stuff like Fermat's great theorem that is easy to state but monstrous to prove. This seems to be a property of mathematics, not of humans... but is it true?
A question by Scott Aaronson from 2011 (3) about P vs. NP seems relevant here: "Will humans manage to prove P≠NP before they either kill themselves out or are transcended by superintelligent cyborgs? And if the latter, will the cyborgs be able to prove P≠NP?" Later, he notes that if P≠NP, "once the robots do overtake us, they won’t have a general-purpose way to automate mathematical discovery any more than we do today".
---
(1) https://mathstodon.xyz/@tao/117207849921390904
(2) I'm not sure whether this is a hard distinction -- e.g. Tao also has some partial results towards Collatz (https://terrytao.wordpress.com/2019/09/10/almost-all-collatz...).
Comment by semiquaver 4 hours ago
Comment by fwlr 4 hours ago
Comment by tehmillhouse 1 hour ago
Not in a happy-go-lucky "if we just ignore the problem of politics and resource allocation for a bit" world, but in ours. Do y'all really think this will make the world a better place?
Maybe stop building the Torment Nexus, you numbskulls.
Comment by num42 4 hours ago
Comment by margorczynski 3 hours ago
Comment by suddenlybananas 4 hours ago
Comment by jeanmichelselli 1 hour ago
Comment by aborsy 3 hours ago
Or will access to internal frontier models provide a big boost?
Comment by StatsAreFun 1 hour ago
Comment by tzone 1 hour ago
The reality is clear though. The chances of AI models overtaking majority of mathematics within next 10 years is becoming very high. Especially if it becomes cheaper to run these models.
As math formalizations improve, AI can have faster progress in math, compared to even computer science or software engineering.
It is simultaneously the best and the worst time to be a mathematician right now.
Comment by Metacelsus 3 hours ago
Comment by abetusk 3 hours ago
Comment by nehan 4 hours ago
I think they should be able to unravel whether or not any sessions by Tristan or Levent went into the training data for this model.
Comment by pfisch 4 hours ago
Comment by dfdydx 2 hours ago
Comment by paxys 1 hour ago
Comment by ImPostingOnHN 37 minutes ago
Or OpenAI could just look at their code and say what it does (maybe have their AI do it if they're having so much trouble with this?)
Comment by whythismatters 3 hours ago
Interesting detail. A heavily pruned version, I assume?
Comment by keel-control 4 hours ago
Comment by quantumwoke 4 hours ago
Comment by o4c 3 hours ago
YT playlist on Millennium Prize Problems By Harvard math department in March 2026
https://www.youtube.com/watch?v=3j1VW9REm7s&list=PL0NRmB0fnL...
On Navier-stokes problem definition:
https://www.youtube.com/watch?v=XoefjJdFq6k
Comment by paretolaw 2 hours ago
I guess solution had not yet appeared in training set.
Comment by light_hue_1 4 hours ago
When your hosting provider has unlimited resources to throw at any problem, all they need to know are the good problems, and they can learn that from your logs, how can you trust them?
They could easily have looked at the logs. We don't know. We'll never know!
You can't trust places like OpenAI or Anthropic with your IP if you're a business. They can easily review all of your logs for interesting discoveries. For example, if your drug discovery pipeline fails to find something that they think might work with 1000x the compute, they can do it. And now suddently they have a new business and you don't.
Comment by jaccola 3 hours ago
Comment by vatsachak 3 hours ago
Comment by lukewarm707 3 hours ago
this is surely the line which confirms they plaigiarised the solution.
Comment by world2vec 4 hours ago
There you go, the suspicion of the "concurrent work" (https://cims.nyu.edu/%7Etristanb/statement.pdf) mathematicians might not be that unfounded after all...
Comment by Stevvo 1 hour ago
Why would they brag about such psychopathic behavior?
Comment by ls_stats 4 hours ago
Comment by mrdependable 3 hours ago
Comment by jdoliner 4 hours ago
Comment by ex-aws-dude 2 hours ago
We've seen in the past they will go to any means to satisfy the desired outcome
Comment by JPC21 1 hour ago
Comment by jabedude 4 hours ago
Comment by Kotlopou 3 hours ago
Comment by DudleyBluffles 2 hours ago
Comment by JPC21 1 hour ago
Comment by DudleyBluffles 1 hour ago
> https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de0...
The prompts for the chat above are:
> Construct a counterexample to general (non-planar) case of Dinitz Garg Goemans conjecture. You should do a breakthrough and find a structured counterexample.
> [gpt works for a while and then gives up]
> Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.
> [gpt works for a while then gives up]
> it's enough of partial results. let's finish with a complete unconditional counterexample
> [gpt proves the problem]
I could have written these prompts sophmore year of highschool, if not earlier. True, it took more experienced mathematicians to verify it, but I don't fancy a role as a glorified editor. I want to solve problems! Discover new techniques! Not babysit an AI while eating breakfast.
Comment by jijijijij 2 hours ago
Yes. Assuming you are young and haven't had such experience.
The world is changing not just because of AI. Everything is unstable right now. You may regret not enjoying the remainder of stability and economic viability prior generations had. It's not like you can expect to get ahead by powering through education. Either your career perspective will soon change for the better, or worse. In any case, you gain little by sticking with career building at this moment in life. You are however, at risk of losing the chance to experience the still mostly pleasant world as is.
Comment by anon109 1 hour ago
Comment by paulsutter 2 hours ago
Which means they can learn from whatever you discuss with ChatGPT unless you are going through a clean API (perhaps Bedrock? Anyone know?)
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
Comment by Marha01 4 hours ago
Comment by frozenseven 3 hours ago
Comment by picafrost 3 hours ago
Comment by simianwords 3 hours ago
Comment by keel-control 3 hours ago
Comment by bluecalm 4 hours ago
I think the market for local models/private datacenters (for bigger businesses) is going to be big. Even if you don't have unique tech/idea/implementation sharing your business secrets with Altman/Dario/Elon/Zuck doesn't look very appealing going forward.
Comment by brcmthrowaway 2 hours ago
Comment by philipwhiuk 2 hours ago
Cause OpenAI will hear about it and beat you to publishing.
Comment by heaney-555 4 hours ago
Millennium Prize Problems were used as examples of something the current approach to AI just wasn't capable of, discussions that would result in "we'll need a totally new architecture".
Comment by rfgplk 4 hours ago
Wrong.
Comment by sashank_1509 4 hours ago
Comment by redox99 3 hours ago
Comment by dmitrygr 4 hours ago
Easy, we stole it from Levent and Tristan
Comment by empath75 2 hours ago
Comment by nbulka 3 hours ago
Comment by int3trap 4 hours ago
It's incredibly tiresome and you'd think people could put more effort into it than just following whatever vibes they agree with.
Oh well.
Comment by 8note 21 minutes ago
having a result means the math can keep moving forward, and having openai and anthropic train against how mathematicians use their models should let math continue to move faster, and the rest of us get to benefit.
I think these traces however should be public domain and publicly available, since they are basically university work
Comment by sophacles 4 hours ago
Oh wait... its not a good comparrison, its an incredibly obvious false equivalence.
Note for the fools: I'm only commenting on the bad faith claim in the comment I'm replying to, not taking a stance on the validity of theft claims. Given the players involved the truth probably some nuanced middle-ground that is worth paying attention to anyway.
Comment by int3trap 3 hours ago
It's a perfect example of people wanting to believe what they want to believe and ignoring evidence in order to do so.
Currently, there's no evidence. So saying it was stolen has no basis other than typical academic posturing and being a bad sport about "losing the race to the solution". Its happened 1000000 times before in academia and it will continue to happen.
If there's proof of OpenAI malfeasance than I'll happily curse them for it at that time. But until then I won't rely on heresay and vibes.
Comment by applicative 4 hours ago
Comment by colesantiago 4 hours ago
Weather an individual or a company found the solution (stolen or not) they both used AI to come get the solution.
We have AGI and the intelligence abundance is going to be amazing for everyone in the future.
Comment by denverllc 4 hours ago
I think it's the dishonesty, the threats of "destroying the career" of one of the mathematicians, and the request that one of the authors disavow *the other individual he was working with for the last 1-2 years* so he could claim the Clay prize as part of OpenAI.
It doesn't surprise me that OpenAI's team were surprised he'd turn it down; it shows that they just assume everyone else is as slimy as they are.
Comment by whythismatters 2 hours ago
I can't put my finger on it, but there's something off about this article, e.g. glossing over the opportunism (acting on "rumors"), drive-by claim about "strict safeguards [...] including monitoring and isolation", high horse attitude (we gave the guy a chance, we don't care about 1M USD, and while you fools are complaining we just tick this box and continue the pursuit of our noble goals for the benefit of humanity). I don't like it.
Comment by achierius 3 hours ago
Why? These 'geniuses in a datacenter' aren't good, they aren't 'aligned', they don't work for you. They'll take your job, then they'll hack your computer, and then who knows what's next.
Comment by achierius 3 hours ago
Comment by philipwhiuk 3 hours ago
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Is the biggest fuck you to the mathematics community.
Credit? Nah if we think you’re close we’ll use your data and swamp you with our improved model. Then we’ll threaten you.
Comment by greatgib 3 hours ago
So we can be suspicious that there is some truth, one way or another that they could have reused prompt/data generated by the user session.
Comment by diomedes 4 hours ago
Comment by Kotlopou 3 hours ago
Of course, as with all of those, it's about the broader program, e.g. section 7 here (https://www.scottaaronson.com/papers/npcomplete.pdf), where Scott Aaronson wants to ask about whether quantum computers using quantum field theory could gain any speed advantage over regular quantum computers, but can't even formulate the question because quantum field theory is mathematically ill-defined.
Just solving Yang-Mills because that's what the prize is attached to would be useless.
Comment by frozenseven 2 hours ago
Comment by colesantiago 4 hours ago
Running agents and prompting excessively to produce 'slopcode' to solve mathematical problems and generate a solution.
If this is what anyone calls 'slop' then slop has no meaning.
I'm all for it on the use case of solving mathematical breakthroughs!
Comment by applicative 4 hours ago
Comment by trainingonme 1 hour ago
Comment by wesammikhail 4 hours ago
Just saw this a few mins ago.
Comment by diehunde 4 hours ago
Comment by cherryteastain 4 hours ago
Comment by ricksunny 2 hours ago
Comment by azan_ 3 hours ago
Comment by diehunde 1 hour ago
Comment by diehunde 3 hours ago