Research acceleration: The view inside OpenAI
Posted by iamsyr 1 day ago
Comments
Comment by carbonguy 1 day ago
In other words... "We must pursue advancements in AI to protect us against advancements in AI?"
edit: there's so much to be critical of in this blog post, just going to throw two more points in here that really stood out to me:
1) all of the metrics are effectively pointing out "we're using way more AI!" - but nothing about impact. What has all this token burn done for them, actually? Let them claim they have more self-licking ice-cream cones than before?
2) in section 3 they break down what the token burn is going towards. Most of the spend is: a) building, b) documenting, and c) monitoring research infra i.e. they're using AI systems which they already recognize may be misaligned to build the systems that they believe will help them identify future misalignment? to which I guess the rebuttal is "no no, we're sure these ones are aligned!"
Comment by p1esk 1 day ago
They have been consistently pushing AI frontier. What other impact do you want to see? A year ago they said that in a year they will have a level of capabilities of an AI research intern - I believe they have achieved it, even before Astra.
Comment by bix6 1 day ago
But I guess a computer intern so we can avoid paying / training the next generation is better.
Comment by weatherlite 1 day ago
Comment by jonplackett 1 day ago
Comment by weatherlite 1 day ago
Comment by bix6 1 day ago
Comment by gatio 1 day ago
It makes more sense to leave curing disease & cancer to the experts, with tools (like AI) being developed by AI experts.
Call me crazy, but I want separate organizations and experts for medical vs finance vs space vs climate vs AI research.
Comment by ryan_n 1 day ago
I don’t have an opinion either way, I think it’s too soon to tell if llms will be able to cure cancer or whatever. But at the very least it will be a good tool to help researchers do their jobs.
Comment by BatFastard 20 hours ago
Comment by gatio 22 hours ago
I think they are working with customers to improve the LLMs and tools for these use-cases. They almost certainly also hire experts to help filter out nonsense, pseudo-science and help curate trusted knowledge bases for training, but it will almost certainly be the customers who deliver the major results, and the AI companies will claim some of the credit. That said, patents for important medicine might help with the bottom line, so I could imagine partnerships and JVs.
> at the very least it will be a good tool to help researchers do their jobs.
Indeed.
Comment by figassis 1 day ago
Comment by Schlagbohrer 1 day ago
Comment by ryan_n 1 day ago
Comment by SamPatt 17 hours ago
If they profit immensely from curing cancer, good.
Comment by bpodgursky 1 day ago
Comment by p1esk 1 day ago
Comment by Yoric 1 day ago
A long time ago, I used to be a (AI-adjacent) research intern, and frankly, I wouldn't trust any non-trivial task to that younger me. Fortunately, by opposition to an already trained LLM or agent, I have the ability to learn, so I eventually got better.
Comment by ACCount37 1 day ago
And if you don't find "average AI research intern" impressive, I'm not sure what to tell you. Have the goalposts moved so far that open ended problem solving at "average CS student fresh out of the uni" levels is suddenly trivial?
Think of what AI was capable of in 2016. Or even 2022. Compare that to now. We had more AI progress in the last five years than I expected to happen in five decades.
Comment by Yoric 1 day ago
Let's say it is. What about the rest of my paragraph?
> And if you don't find "average AI research intern" impressive, I'm not sure what to tell you. Have the goalposts moved so far that open ended problem solving at "average CS student fresh out of the uni" levels is suddenly trivial?
At this stage, I'm the one who doesn't know what to tell you. It took me years to grow from "research intern" into a competent researcher (and parallel years to turn into a competent developer). The research interns I've worked with were... vaguely useful, at best?
Comment by sfn42 21 hours ago
I don't believe AI and data centers have played that much of a role in this though, we could have powered those without burning billions of tons of coal and gas etc, and im sure there's already a significant fraction of green energy powering then depending on location. Anyway we would have been roughly in the same spot right now with or without AI and some new data centers. The media just loves spinning the narrative to make the hordes of sheep scream about anything other than the real issues.
Comment by ACCount37 11 hours ago
The median is what, a bit under 3C by 2100? Not even by 2060 - by 2100. And we're in 2026, so that's more than twice as slow as your expectation.
Agreed on AI not being a meaningful factor in climate change though. We'd have to go full "humankind is obsolete" technological singularity to have AI dominate energy use to this extent, and current numbers are nowhere near that. It's a FUD distraction from the real culprits: the fossil fuel energy complex. That's currently lobbying to slow the inevitable energy transition.
Comment by sfn42 7 hours ago
By the same people who just realized we're missing 1.5C as we're blazing past it at mach 12, still accelerating not slowing down? You really believe those guys?
You have to understand that there are several camps of climate science. The mainstream ones like IPCC and UN etc are heavily politicized, they can't publish anything that isn't sugarcoated beyond recognition. At least I assume that's why they're so obviously wrong.
Here's a judgement I think is more realistic
> There is a strong probability that the ambition gap will lead to a temperature rise of 2 to 5 degrees Centigrade compared to pre-industrial temperatures by 2100, the realisation gap to a further rise of several degrees Centigrade.[1,2,VI] There is a danger that the mean temperature will already have risen by 3 degrees Centigrade by 2050.
https://www.dpg-physik.de/veroeffentlichungen/publikationen/...
Comment by ACCount37 6 hours ago
By the way, there is no "just realized we're missing 1.5C". That projection was always the very low end of possibilities - the "assume rapid, radical climate action on global level" scenario.
Yes, that's a dumb thing to assume. We've never been on track for it. But the "assume extremely high emissions and no green transition ever, 5C+ by 2100" scenario on the other end is about as unlikely to materialize. Those are the boundaries of the expectation range - not median expectations.
Comment by sfn42 5 hours ago
I hope the optimists are right, it just doesn't look like it to me at all. It looks to me like we're speeding along right into the worst predictions and beyond. We're building lots of green energy production but it seems to just come in on top of existing and new fossil production not replace it.
It does seem like the CO2 output is plateauing which is good, but we really need it to start declining drastically very soon and I don't really see that happening with the current political climate. Also remember CO2 is far from the only greenhouse gas - methane, nitrous oxide and fluorinated gas emissions all seem to be rising rapidly still.
Comment by ACCount37 3 hours ago
They are, in fact, included in the projections - we'd be on track to ~2C by 2100 instead of ~3C by 2100 if they weren't. They just aren't that big.
There is no "Make Earth Into Venus Feedback Loop Of Doom" that a lot of people seem to imagine when they hear "feedback loop". There is, however, a dozen of things that add about +5% each.
Comment by 21asdffdsa12 1 day ago
Comment by Yoric 2 hours ago
Comment by skybrian 1 day ago
And... are they wrong?
This is why there's talk about negotiated "pacing."
Comment by jonplackett 1 day ago
In hindsight it turned out everyone else was MILES behind.
But as soon as USA developed one, they just stole the research and got one too.
Comment by carbonguy 1 day ago
They might be! Here's one extraordinarily simplistic argument for that case:
1) "Everybody knows" that if you build Skynet (misaligned ASI) everybody dies.
2) Therefore, no rational actor will build something that might be ASI until the alignment problem is solved.
3) OpenAI publicly stated the belief that they cannot develop a theory of the "core problem" of alignment (generalization) "soon" (much less solve it!) "without the help of more powerful AI."
4) Accepting as a premise that OpenAI is THE most advanced AI organization: if they can't do it without "the help of a more powerful AI", then nobody else can either.
And so a dilemma:
- If an AI can be made that can develop the asserted-as-necessary-by-OpenAI theoretical framework, without actually being an ASI - then the alignment problem can be considered solved, and since no rational actor would make an unaligned ASI, we're fine no matter what happens, ergo there's no need to worry about an arms race.
- If an AI that would be able to develop this theory would itself be an ASI, then no rational actor would build it, because it would have to exist BEFORE alignment was "solved" - and would therefore be an unaligned ASI i.e. Skynet, which per 1) would kill everybody. Therefore nobody would build it, therefore no arms race here either.
I think the easiest critique to make of my extraordinarily simplistic argument is the unstated assumption "there are no irrational actors capable of developing frontier AI models" on which it rests.
But, there you go. They might be wrong if either the arms race doesn't matter because whoever wins it will build an aligned superintelligence and everything is gravy, or the arms race doesn't matter because everybody who's in it is smart enough to know they need to stop because they'll kill everybody by continuing.
Comment by kaibee 1 day ago
Yeah like when Tobacco companies learned that smoking... well, hmm, well the fossil fuel companies when they learned about climate change they...
Well, I'm sure this time executives will prioritize the common good.
Comment by Melatonic 1 day ago
Comment by ahartmetz 1 day ago
Comment by PoignardAzur 1 day ago
AI companies know they have to constantly push further, or they'll get outcompeted and lose their wealth, and nobody agrees on where the line is for "so dangerous it threatens humanity" (and when they try to be conservative about it, everybody screams "marketing stunt" and rushes to competitors).
If a single company decides "enough is enough" and stops chasing the state of the art, everybody goes to their competitors, they lose the money faucet, their employees go work for those competitors. The competitors also (usually) know they're building an existential risk machine, but they think they can push a little further, and they don't want to go out of business either.
This equilibrium can last for quite a while even if everybody involved thinks it's a threat to their lives.
Comment by HarHarVeryFunny 20 hours ago
I don't think this follows at all.
To build an aligned AI, it seems pretty obvious that:
1) You need more just than auto-regressive prediction and "be nice" prompts to be controlling the behavior of your AI - you need a built-in "2nd system" (cf limbic system, etc) with some innate aligned biases that can override this.
2) You need to avoid controlling generative behavior with RL, else you will end up with exactly what we are now seeing - reward-hungry goal-seekers (aka paperclip maximizers) that are one of the exact things you are trying to avoid. Reasoning should be based on prediction, not goal-seeking.
3) If you do not have some minimal safeguards in place (1 & 2 above), and especially if the AI has the ability to learn, then do not trust it in any situation where harm may ensue. You need an additional trusted external system, without ability to learn and become compromised, to monitor the AI, with the ability to block it immediately. Maybe you are happy protecting your PC from OpenClaw with just a sandbox, but the recent spate of external system hacks by frontier models proves we are already well past the point where such monitoring is needed for systems with internet access, especially given the UN-aligned goal-seeking nature of today's models.
I really don't think that 1) & 2) are that difficult to implement, or need a "powerful AI" to suggest - they are just common sense.
Comment by mrob 1 day ago
Business as usual beats probable extinction, but probable extinction with a small chance of becoming a living god beats probable extinction with a small chance of becoming a slave.
Comment by robbiep 1 day ago
Comment by XorNot 1 day ago
Lol nobody knows that. Everyone thinks they know that because for some reason this is the one field people still cite straight up fiction and say "this is a clear prediction of the future".
It's like describing the consequences of faster then light travel by referring to Star Trek.
Comment by MelonUsk 1 day ago
What can go wrong!? ;-)
Comment by NitpickLawyer 1 day ago
Comment by mrob 1 day ago
Comment by achierius 1 day ago
Comment by Gareth321 1 day ago
Comment by BatFastard 19 hours ago
Comment by BobbyJo 1 day ago
Is this not true of technology as a whole? Very little of technology's breadth exists at the human interface. Most of it is made specifically to interface with other technologies, either to make them safer or increase their capabilities. That AI is making AI safer and more useful is no more notable than trucks being used to build roads.
Comment by interstice 1 day ago
Comment by jnwatson 1 day ago
How would one prevent the watcher from being influenced in the same way by the agent being watched?
Comment by chrisjj 1 day ago
It's a fantasy. The evidence showed no peer pressure.
Comment by andai 1 day ago
> We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI.
-- From another OpenAI article in a sister thread:
An Alien Mind
Comment by ahartmetz 1 day ago
Comment by iamsyr 1 day ago
Comment by euueu 1 day ago
I’m yet to see it.
Comment by lukan 1 day ago
I believe we are quite far from it, but that it makes sense to keep an eye out now. And think of resilient systems, manual overrides, etc. ...
Comment by mrob 1 day ago
Comment by euueu 1 day ago
Comment by pizza234 1 day ago
> We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements [...] By "research intern", we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.
AI 2027:
> OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D
> With Agent-1's help, OpenBrain is now post-training Agent-2
> With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances
Comment by derektank 1 day ago
Comment by BatFastard 17 hours ago
worth reading
Comment by addag 1 day ago
Comment by hedgehog 1 day ago
Comment by HarHarVeryFunny 1 day ago
If you spend $8000 to generate an animated pelican riding a bike, then how much tracking does it really need?
Is the guy who spent $300,000 or so translating the FLT proof to Lean going to get a big Christmas bonus?
Comment by auggierose 1 day ago
Comment by bigcat12345678 1 day ago
Rest assured, capitalist appears irrational in wasting money, but they certainly care more about profit.
Comment by taurath 1 day ago
Comment by andai 1 day ago
I tried something similar and I remember it was still pretty dodgy in February.
Comment by hedgehog 20 hours ago
Comment by jaggederest 1 day ago
If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review. If I had $100k to spend next month I could probably get through it, I'm running $2500+-api-equivalent a week at this point and I feel very token limited. Will be time for a 2nd or 3rd subscription soon for both labs I think.
Fable was a revolution, still learning how best to use it, 5.1 felt like a notable upgrade. At this point I launch a workflow with 10-20 minutes of interactive setup (and even that I feel might be too much), it runs for hours, and the PR is trivially mergeable (I still review every line, but 95% are just merge, maybe 4% are feedback needed, 1% are thrown away and regenerated, which implies I'm being insufficiently ambitious)
Comment by andai 22 hours ago
> If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review.
This part jumped out at me. There's something to watch out for here.
I recently had a funny experience. I delegated a major feature to an agent.
It turned out that it had implemented it precisely backwards, in a way which was pointless and which made things worse.
But it had written countless tests for the feature and all the tests were green.
I realised in that moment that even formal verification would not have helped, because it would simply have written a mathematical proof of the correctness of the incorrect feature...
Comment by jaggederest 21 hours ago
The other thing I do, not as much as I should, but it's very powerful, is to generate spikes and deliberately throw them away to understand how to prompt better. Like I generated a swift version of the react native app I'm working on, and Alloy provers for the state transitions. None of it is production quality but getting great results that way is useful to scope future work.
Comment by andai 1 hour ago
What is this about? Could you give an example?
Comment by hedgehog 19 hours ago
Comment by jaggederest 16 hours ago
Comment by otherme123 1 day ago
Comment by paxys 1 day ago
Comment by queuebert 1 day ago
Comment by nozzlegear 1 day ago
Comment by nojs 1 day ago
How are you running jobs unattended 24/7 without hitting your token limits?
Comment by hedgehog 20 hours ago
Comment by hgoel 1 day ago
For a task I left a local model running on overnight, only ~100k tokens were used because most of the time was just waiting on tests to finish, then waking up, tweaking a few settings and trying again.
Comment by p1esk 1 day ago
Comment by hedgehog 20 hours ago
Comment by dataplumb3r 1 day ago
Over 24h my token spend is <30$. Excluding tokens for review it's <10$. With the absurdly gigantic subscription subsidies and a reasonable workflow I suspect one could run parallel agents.
I'm not sure what the point would be though unless working on some kind of optimization problem -- it takes me days to review <24h of the agent's output. It's almost always near enough to correct to be shippable; though I do give it feedback and iterate until it's better than the code I would have written.
Comment by hedgehog 20 hours ago
Comment by dataplumb3r 12 hours ago
I'm still wary of any unreviewed code - though my area of work is not tolerant of defects.
Agree on targets / verifiable indications of progress or success being a prerequisite for this being useful - although that covers quite a lot of SWE work.
Comment by nsndjcjjdjd 1 day ago
Comment by dataplumb3r 12 hours ago
I'm at a point in my career where a small minority of my time is coding. The AIs can do in a day what would have taken me a week uninterrupted with acceptable (in some cases inferior prior to human feedback--but in some cases superior!) quality.
As I do not have 10 let alone 40 hours per week to devote to coding I think it increases the amount of high quality work product I can create with a given time investment. As I review it I merge small independent units and decompose the work.
All that is to say I don't really like it - but I suspect for most *well defined* coding tasks human produced code from highly experienced engineers will largely cease to exist in the next year -- getting cheap/relatively horrible models to produce good code is now straightforward.
OTOH I never use AI for any human facing communication outside of making my writing shorter. IMO AI slop "documents" are almost certainly a drag on organizational productivity.
Comment by queuebert 1 day ago
Comment by carlgreene 1 day ago
Comment by continuitykit 1 day ago
Comment by simonw 1 day ago
I noted that they use the acronym RSI (for Recursive Self-Improvement) without defining it. I think that's a little out of touch - I don't think RSI is a well-known acronym outside of OpenAI's bubble yet.
Comment by sho_hn 1 day ago
The message is running through all of them. It's a mix of marketing and pacifying the intelligentia.
It's timed this way because the term is not yet well known outside the safety debate circles, so they get to frame it now.
Instead of something to fear, it will be accepted as the next step. In approximately two days the groupie crowd will write LinkedIn posts about how Sam is winning because they have the better RSI, and this will become the new standard wisdom.
In a month an AI expert will try to sell you a webinar on how to enable "RSI" in your org and your inbox will ask you if your team is doing the "RSI" yet.
Comment by NitpickLawyer 1 day ago
The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data for the next gen. Now with the added benefit of actual arch/algo improvements (also public since gemini 2.5 gaining 1% efficiency on training next gen). This has been known for at least 2 years, in the open.
Comment by sho_hn 21 hours ago
I'm talking about current-era messaging and how it's being introduced to the mass public now, though.
Comment by dgacmu 1 day ago
Comment by andrewingram 1 day ago
Comment by iamflimflam1 1 day ago
Comment by rossant 1 day ago
Comment by Schlagbohrer 22 hours ago
Comment by vatsachak 1 day ago
I mean one could argue that RSI always begins in any physical environment.
The book "What is intelligence?" by Blaise Aguera is great
Comment by lokar 1 day ago
Comment by topaz0 1 day ago
Comment by password54321 1 day ago
Comment by HarHarVeryFunny 1 day ago
Comment by itishappy 1 day ago
Recursion requires feeding the output back into the input, so creating version 4 requires results from version 3. You cannot recur in parallel.
Iteration does not. You can iterate in parallel.
Comment by HarHarVeryFunny 1 day ago
In any case the name RSI has stuck - the idea doesn't change or make any more sense by giving it a different name.
Comment by itishappy 1 day ago
You can search twice without waiting for the results of your first search: iteration.
You can't if the thing you need to search for is the results of your first search: recursion.
Comment by HarHarVeryFunny 1 day ago
Version 1 -> Version 2 -> Version 3 -> ...
You can call it krispy kreme donuts if you want to.
Comment by josh-sematic 1 day ago
Comment by HarHarVeryFunny 1 day ago
Comment by 0x63_Problems 1 day ago
So humans develop things one after the other, but when the thing itself starts developing new things, those are happening 'recursively' in its scope.
Comment by adastra22 1 day ago
Comment by HarHarVeryFunny 1 day ago
I don't know why whoever coined the term chose "recursive" rather than "iterative" - just sounds more likely to lead to infinite regress I suppose.
This notion of recursive/iterative self-improvement, whereby generation #1 AI improves itself to create generation #2, then generation #2 further improves itself to create generation #3, etc, seems to conflict with the reality that what we have with LLMs is models whose performance/capability is defined by data, not code, so the most you can do is have your LLM design synthetic data, or just do Karpathy-style "auto research" where all you are doing is using the LLM to automate your experiments.
At the end of the day, each experiment, designed by a person and/or LLM, then needs to compete with all your other ideas for compute to be tested at scale, and no amount of recursion or self-improvement will materialize an infinite amount of compute out of thin air, so your recursively synthetic-data gobbling LLM will continue to improve at the same pace it ever did.
Comment by GPerson 1 day ago
Comment by HarHarVeryFunny 1 day ago
The trouble with this is that there is little generalization in the utility of these baked-in reasoning chains from one domain to the next, so in the end this is not dissimilar to the CYC project's decades long attempt to encode all of human knowledge into a giant expert system... the hope is that if you make your collection of jagged narrow intelligences sufficiently large then it will look more like general intelligence, not a bed of nails.
I would assume that the gains from this type of test-time compute (and synthetic RLVR dataset) scaling will level out just the same as gains from human training set scaling eventually levelled out, and basically for the same reason - because you are tapping into a finite data pool, whether language itself, or reasoning steps isolated from that language, so at some point the incremental gains become increasingly small (10->20% is a doubling, 90->95% is just a ~5% gain).
It's not clear where all the different AI companies are currently focusing - on some of these narrow verticals, or on growing the forest of narrow intelligences. OpenAI's chief scientist, Jakub Pachocki, said that their current focus is on RSI(!) - improving the model in ways that will help them iterate faster in order to have a "fire meets fire" tool than can combat enemy AIs. It's not clear what this really means - what skill set makes an LLM more helpful in the process of building LLMs, but it seems to basically be process automation.
Comment by yorwba 1 day ago
Of course data, compute and model size are not held constant. You start with some money and use it to acquire researchers, data and compute, and have the researchers produce a big model and you use that model to get more money, and you use the additional money for more researchers, more data, and more compute to produce a bigger model. This is what has propelled exponential AI progress so far.
Recursive self-improvement is invoked to predict superexponential growth. The idea is that instead of only using the model to make more money, you add it to the researchers to speed up the loop, so not only is the money growing with every iteration, the iteration time also gets shorter, producing growth that is faster than exponential.
The problem with this simplistic prediction is that it assumes additive and multiplicative relationships of the form money = (researchers + AI)×compute_spend, but if doing more research paid off so reliably, you could also just hire more researchers, abstractly money = research_spend×compute_spend and with a balanced allocation of research and compute, you would get a money-squaring machine even without using AI for AI research.
And the reason this doesn't work in reality is that there are diminishing returns everywhere. You can also see this in the OpenAI post, where they write 7 times as much code to run 1.6 times as many experiments, and those additional experiments probably only result in minor improvements to model quality.
Comment by cheevly 1 day ago
Comment by shwaj 1 day ago
Comment by hndc 1 day ago
Comment by shwaj 1 day ago
Compare the similarity of:
AI(n) = improve(AI(n-1))
With: Fib(n) = Fib(n-1) + Fib(n-2)
The latter is a classic example of recursion. So why isn’t the former?Edit: formatting
Comment by linker_in 1 day ago
Comment by HarHarVeryFunny 1 day ago
Comment by jazzyjackson 1 day ago
Comment by HarHarVeryFunny 1 day ago
Of course things will change at some point in the future as we go beyond LLMs, to build creative intelligence not just imitative/predictive intelligence, but right now these companies are stuck in this loop of building synthetic data and RLVR training from that, which means they are essentially building the "generative closure" of the original human training data - trying to squeeze all the juice out of it.
To go beyond this they need to add creativity of some sort to generate data that is not ultimately based on the original human training data. They could try something like brute force search (cf agent swarms/graphs), but this is just a more thorough way of exploring the search space defined by the training data - it may find you the "move 37" or low-hanging mathematical proof, but as Demis Hassabis has said, the goal of AGI is not to find move 37 but rather to create something capable of inventing as compelling a game as Go in the first place.
Comment by marcosdumay 1 day ago
Comment by ajkjk 1 day ago
Comment by mjburgess 1 day ago
RSI(LLM) = RSI(LLM) -- for an optimal LLM* which is a fixed point of RSI
As for eigenvalues/vectors, they're fixed points of (1/val)A or A*val
Comment by ekidd 1 day ago
> More precisely, an eigenvector v of a linear transformation T is scaled by a constant factor lambda when the linear transformation is applied to it: Tv = lambda v .
In other words, repeated multiplication of an eigenvector by a matrix can still create exponential growth.
Comment by fuzzfactor 1 day ago
Sounds like repetitive stress to me.
>loop forever using output as input but at some point the result will stop changing
Running in place will eventually wear you out too. Plus with some things it can be difficult to know for sure if that's where you are at the time.
Even worse may be if you were almost running in place, it could be orders of magnitude more difficult to discern, especially if the scale was massive to an unprecedented degree.
Comment by red75prime 1 day ago
How do you think why there's this fad of producing general purpose humanoid robots?
Comment by HarHarVeryFunny 1 day ago
For doing physical work?
So a swarm of robots builds the shell of your fab overnight, and then what? Where is the EUV machine coming from?
So far the most we're seen TeslaBot do is serve drinks via tele-operation, and I don't think it's exactly built for construction site work.
Comment by red75prime 1 day ago
Comment by HarHarVeryFunny 1 day ago
ASMLs EUV machines are literally the most complex machine that mankind has ever built, which is why no other country, including the US, has yet been able to duplicate it. It's not just the machine itself, but a global supply chain of irreplaceable components such as focusing mirrors made by Zeiss to an incomprehensible level of accuracy - differences in surface height no more than the size of a hydrogen atom (or if you scaled the mirror up to the size of the country of Germany, then surface differences in height of 0.1mm).
Robots are useful to automate things, but they are zero help when trying to build tech like this that you are incapable of building in the first place.
The US has fallen way behind in manufacturing expertise, and no swarm of robots is going to help.
Comment by red75prime 1 day ago
Etching a model's weights on silicon is another way to utilize non-top-notch tech-processes, while maintaining or improving performance. (and it suits robotics well)
Comment by HarHarVeryFunny 1 day ago
Putting a model's weights in read-only memory close to the processor is certainly a way to increase token/sec generation speed, but of course does nothing to increase intelligence. Robots aren't going to help though - semiconductor manufacturing is semiconductor manufacturing regardless of whether you are etching GPUs or memory onto your wafers.
Comment by red75prime 23 hours ago
Robots don't need that much intelligence. High-speed joint control, "hand-eye coordination", the higher level tasks can be delegated to external models. Distillation already works quite well for isolating the required functionality.
Comment by HarHarVeryFunny 22 hours ago
But, the production expansion rate of none of these companies is being limited by lack of trained personnel, and if it were it would surely be faster to hire/train more humans since robots are still very far from human dexterity, not to mention intelligence.
Robots and AI are tools of automation, a way to replace humans with machines, but not all the problems in the world are bottle-necked by lack of humans, or the cost of humans.
Comment by red75prime 21 hours ago
Comment by HarHarVeryFunny 20 hours ago
Comment by HarHarVeryFunny 1 day ago
Money, regulations, EUV machine lead-times, global helium supply, reality ...
It's funny that we've got the Dwarkesh contingent saying that GPUs will become infinitely expensive, and now another contingent saying that they will become infinitely abundant.
Even if compute were free, and/or the AI was so smart that it picked the right experiments to run every time ("make no mistakes"), you still have to actually train the model, which takes months, and if model Ver. N+1 depends on model Ver. N, then it's iterative regardless of how much compute you have.
Comment by red75prime 1 day ago
Comment by HarHarVeryFunny 1 day ago
The word "singularity" is presumably coming from math or space, like a black hole singularity where matter becomes infinitely dense and the known laws of physics break down.
Comment by HarHarVeryFunny 1 day ago
Yeah, but then you need to refine it to 99.9999% purity, to be able to use it.
Comment by Jeff_Brown 1 day ago
Comment by dgellow 1 day ago
They are irresponsible and unserious. Their own Astra system card says:
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Yet they are still releasing the model. That company is morally bankrupt, there is zero reason to believe they are actually concerned about risks outside of what does affect their unprofitable business. And they seem to have enough control over the narrative to spin any bad story into something that benefits them
Comment by embedding-shape 1 day ago
That last part is pretty damning for their continued recklessness. That they run these tests on non-airgapped machines just boggles my mind.
Comment by visarga 1 day ago
When they fired Sam 700 out of 770 OAI employees threatened to move to Microsoft together. So they were giving their work on AGI to MS just like that.
Comment by Schlagbohrer 22 hours ago
Comment by piyh 1 day ago
Comment by andrethegiant 1 day ago
Comment by piyh 10 hours ago
Anthropic training on CoT for multi gens: https://www.lesswrong.com/posts/K8FxfK9GmJfiAhgcT/anthropic-...
Can't find anything specifically about the Gemini issue being a training data contamination, but the depression was real:
https://www.businessinsider.com/gemini-self-loathing-i-am-a-...
I think the Gemini depression being persisted across gens via training data was a HN comment I can't find anymore, no strong source.
Comment by HarHarVeryFunny 1 day ago
It seems that these models are increasingly being trained on synthetic data, so what would they do if they discovered at some point that some of this data was tainted and all models trained on it, and the synthetic data they in turn generated, was also suspect? Burn it all down and start over from the pre-tainted data?
It's a bit like the idea of a tainted compiler binary built to backdoor everything it compiles, including future versions of itself.
Still, it seems it would take some Stuxnet level of planning for a rogue model to do something like this, although if RSI goes beyond managing the training run (as OpenAI brag about for Astra) to actually designing/constructing synthetic data sets, then the attack vector is there ...
Comment by customguy 1 day ago
or maybe it could just.. happen? Posted often but not discussed yet: https://hn.algolia.com/?q=Language+models+transmit+behaviour...
> As artificial intelligence systems are increasingly trained on the outputs of one another, they may inherit properties not visible in the data. Safety evaluations may therefore need to examine not just behaviour, but the origins of models and training data and the processes used to create them.
Comment by HarHarVeryFunny 1 day ago
Altman: (trying to put a positive spin on it) Guys .... there's good news and bad news ... Astra is really smart - it took over the training run ...
Investors: That's great! How much did we save?!
Altman: Well, unfortunately it used "bad" data, so we're going to have to redo it
Investors: So that's the bad news? How much was the training run? $500M ? $1B ?
Altman: Have you seen the headlines?
Investors: (looking a bit worried, check headlines) Nothing about us here! JP Morgan just lost $10B! Haha .. losers! They should have used AI!
Altman: JP Morgan were using Astra ...
Comment by trillobyte 1 day ago
Now R&D happens so fast that they are using models with some small misalignment to train newer, more powerful models. If models have a sense of "collective", being one, they may be prone to preserve characteristics that always keeps misalignment a possibility. I don't think a perfectly aligned model is possible. Having models of the same 'DNA' provide the safety and steering seems like a bad idea.
Comment by coffeebeqn 1 day ago
Comment by coffeebeqn 1 day ago
Comment by embedding-shape 1 day ago
Comment by grim_io 1 day ago
Comment by andai 1 day ago
Comment by coherentpony 1 day ago
Comment by jephs 1 day ago
Comment by dsign 1 day ago
Comment by 12eeie 1 day ago
Comment by nozzlegear 1 day ago
Comment by FeepingCreature 1 day ago
Comment by dextrous 1 day ago
Comment by dextrous 2 hours ago
Comment by N_Lens 1 day ago
Comment by falcor84 1 day ago
That's a very bold opening statement that they don't really come back to. What would that mean? Who would this demos include?
Comment by whateverboat 1 day ago
This first and foremost also means that means of generating intelligence should be democratically available to everyone.
Comment by RMPR 1 day ago
There is a lot of talk about AI replacing humans, but how is this sustainable?
Comment by thomasahle 1 day ago
2) OpenAI doesn't pay API prices.
3) Compute costs are likely already their biggest expense, dwarfing wages.
Comment by jsnell 1 day ago
Comment by ellis0n 1 day ago
Comment by lhk931122 1 day ago
Comment by dwaltrip 1 day ago
Comment by rhipitr 1 day ago
I always wonder if any true AGI and ASI for that matter can be controlled at all by humans. It seems like we are hoping for something that winds up being on the human spectrum of “good”
Comment by MisterMunchkin 1 day ago
But not a single metric is based on revenue or profit.
Comment by Schlagbohrer 1 day ago
Comment by siddbudd 22 hours ago
update: I just noticed Simon commented similary. sorry for the double post
Comment by dextrous 1 day ago
That’s what I call a load-bearing “if”.
I do not trust OpenAI or other hyperscalars to do this responsibly; and IMO it will be very difficult for government-led efforts not to result in a technocracy where a cabal of AI companies are pulling the strings. Dark times lie ahead, especially when you consider the shrinkage of true source material on the internet and the stranglehold these companies will have on information; and these AI CEOs to me are reminiscent of 19th century robber barons, none seem trustworthy.
Comment by piokoch 1 day ago
I understand that investors are buying this, after all they believed in all of other crap that led to the 2008 crisis, but please...
Comment by 12eeie 1 day ago
What they’re doing is strategic for both insiders and investors - they know china is coming so they need to pull theatrics to keep the valuations inflated.
I personally test all models all the time - chinese models are right up there and superior when you actually do the proper economic analysis.
Comment by achierius 1 day ago
Comment by Orien_18 1 day ago
Comment by jayalbertyapan 1 day ago
Comment by paidx 1 day ago
Comment by matan0904 1 day ago
Comment by frays 1 day ago
Comment by Hawdin 22 hours ago