Do you think it happened? Research stolen from their Codex private chats
Posted by KingOfMyRoom 13 hours ago
Comments
Comment by QuadmasterXLII 10 hours ago
These models will absolutely remember a brilliant insight that appeared a single time in the pretrain corpus, because if they couldn't-- they would get a slightly worse loss. I don't know what this wishy-washing "well maybe we trained on it but we didn't read it" is supposed to mean.
(Example of Fable knowing the content of a deeply unimportant LessWrong post I wrote: https://claude.ai/share/a907b46c-bf7b-4fca-9c71-8582cf8507cc a working Navier Stokes solution would be way more salient)
Comment by LUmBULtERA 9 hours ago
Comment by anon373839 5 hours ago
Comment by QuadmasterXLII 8 hours ago
Comment by mucha 8 hours ago
Comment by drooby 4 hours ago
He doesn't make a claim that implies "don't train" might not matter in the way you might think.
They cannot rule out that the data was trained because they have no per-user provenance tracking through the training pipeline once data is de-identified..
The entire point of de-identification is the inability to know the source of data. If the researcher forgot to hit "do not train" then that's that..
The only thing they have is a coincidence and the fact that the LLM may have used the training data that then researcher technically may have agreed to share.
Whether or not that's smoking gun of anything is hard to say. And the fact may remain that the proofs are significantly different, we do not know.
Comment by Havoc 12 hours ago
Everyone I think assumes training in the general “it affects the distribution” sense. But if it fishes precise novel insights out as alleged then it makes these products dramatically less valuable. At least the non zdr ones
Comment by znnajdla 10 hours ago
Comment by mannanj 11 hours ago
I think people straw man the whole data argument on “opt out of training purposes” and ignore the mandatory analytic purposes. Which probably includes figuring out trends and important research problem data to steal from users and even business ideas.
Comment by Havoc 9 hours ago
Comment by jrflo 12 hours ago
Comment by naniel 9 hours ago
What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit.
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/
Comment by avadodin 6 hours ago
Anyways, the moral of the story is: Not your GPU, not your data.
Comment by willtemperley 12 hours ago
Comment by quicklywilliam 10 hours ago
Comment by khelavastr 12 hours ago
Comment by ChrisArchitect 12 hours ago
And currently: https://news.ycombinator.com/item?id=49613262
Comment by Ydarbleoj 13 hours ago
Comment by cousinbryce 10 hours ago
Comment by josefritzishere 10 hours ago
Comment by londons_explore 8 hours ago
Really looks like they aren't yet training on conversation histories effectively.
Comment by e_l 4 hours ago
They might have determined that your canaries (whether it is a unique string, URL or similar) have a low signal-to-noise ratio and not worth following up.
Unless you provide more details, it's difficult to ascertain whether your conclusion is correct.
Comment by ungreased0675 7 hours ago
Comment by slipperybeluga 8 hours ago