OpenAI's GPT-6 Astra on ARC-AGI-3
Posted by vignesh_warar 5 days ago
Comments
Comment by at1as 4 days ago
From https://epoch.ai/latest/announcing-frontiermath-erdos
> Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours
> Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself.
Which implies a genuine improvement in capability, but there's a still a very long tail ahead that models will continue to need to improve to capture.
Comment by zone411 4 days ago
"Problems solved before a model's training cutoff can be filtered out, and all models compared on the remaining problems" means that the problems an older model actually solved are the ones that get filtered out, while the remaining problems are the ones it already tried and failed on. So older models end up with 0s on the filtered set and you can't really use this to compare new models to older ones.
Also, since these are known public problems, you can't stop people from spending far more than your arbitrary time and $ limits on them. So the number of clean problems will go down over time.
Comment by red75prime 4 days ago
It's a sarcastic take and I understand that you are probably talking about "spiky intelligence", but you've chosen unsolved problems as a measure of the progress yourself.
Comment by at1as 4 days ago
There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems
Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're past that now). New models will exhibit something novel by adding solutions. And the matter of solution is interesting as well (contradiction versus a positive proof).
I would venture to predict it'll take years to decades to get to 0 open problems. But I'd be very happy to have this comment look foolish in retrospect as models continue to improve
Comment by red75prime 4 days ago
Comment by at1as 4 days ago
I'm most interested in these problems as a relative measure of performance for successive model generations. If the prior generation couldn't solve a problem but the current one can, that's useful information, especially when we take into account what the proofs look like.
It's not a perfect benchmark, but I prefer it to many others that I see floating around.
Comment by Davidzheng 4 days ago
Comment by malfist 5 days ago
Comment by matherial 5 days ago
I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.
Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.
Comment by pavlov 5 days ago
It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
Comment by phainopepla2 5 days ago
The people at Mensa aren't a valid sample of people who score high on IQ tests, because there is a such a strong selection effect for people with certain personality traits, such as wanting to join a club based on your IQ.
Comment by scotty79 4 days ago
Comment by jaggederest 5 days ago
"Please accept my resignation, I don't want to belong to any club that would have me as a member".
Comment by guelo 5 days ago
Comment by itishappy 5 days ago
(Please forgive the flippant response. I believe it cuts to the core of what the parent was intending.)
Comment by red75prime 4 days ago
Comment by itishappy 4 days ago
I wanted to highlight that we're actually looking at a trio of concepts: intelligence, IQ, and value. Strongly correlated concepts, yes, but also meaningfully distinct!
Comment by adastra22 4 days ago
Comment by scotty79 4 days ago
IQ has precise definition. It's what the tests measure. Intellect has only fuzzy handwavey definition. Correlation between IQ and intellect is about as loose as the definition of the intellect. The way people make the correlation even looser is by defining intellect in even more fuzzy and nebulous manner.
Comment by red75prime 4 days ago
Comment by sfblah 5 days ago
Comment by shric 4 days ago
Comment by scotty79 4 days ago
Comment by scotty79 3 days ago
Given that they are probably rarely in groups of purely random people, if you are Mensa member and you are in a group of dozen people (for example in professional setting) you'd probably have a 50/50 chance of not being the one with the highest IQ there.
Comment by sfblah 3 days ago
Comment by scotty79 2 days ago
Comment by tptacek 2 days ago
Comment by howunfortunate 4 days ago
For these purposes they are highly reliable (repeatable, internally consistent) and valid (correlate with ~everything to about the degree one would reasonably expect).
They were never designed for machines or non-human animals.
Nor were they designed for rare ranges of intelligence - these are by definition hard to create tests for, since it's hard to gather the sample sizes you need. So they work well for the middle ~98% of humans but can't discriminate well among the most profoundly intellectually disabled nor among true geniuses.
Comment by tintor 5 days ago
Let LLM control a physical robot to perform tasks that average human can do.
Comment by quotemstr 4 days ago
Comment by matherial 4 days ago
Case in point: you had hundreds of millions of JPEGs to vacuum up and bitmap image generation is amazing. But if you ask them to recreate the same scene as vector art, they will struggle to generate a decent SVG. Like, kindergarten-style pelicans on bicycles are the state of the art. It should generalize seamlessly, but somehow, doesn't?
I think it will happen, just like self-driving cars are happening, but it will probably be a slow process.
Comment by Mil0dV 4 days ago
That (or a near future one) combined with an llm would be insane
Comment by tintor 4 days ago
Comment by skybrian 5 days ago
Comment by idiotsecant 4 days ago
Comment by skybrian 4 days ago
Also, the skill of the human opponents matters. You'd want to test it against people who have practiced playing the game. Otherwise, it's like the difference between building a chess bot that can win against random undergrads who don't normally play, versus winning against grandmasters. And it's not like there's a pool of skilled human players of the imitation game.
Comment by ranyume 5 days ago
Comment by guelo 5 days ago
Comment by ranyume 5 days ago
Comment by mdp2021 5 days ago
No, it is just because they have difficulties at the bench.
> how gives we don't measure them by
We'd measure them by all the tests available. Not all test are usable in all circumstances.
Comment by ranyume 5 days ago
I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments for good which is one way we define intelligence.
On the other hand saying "the gifted person is highly intelligent/smart just not good at some things" really diminishes the other tasks, because they really are very difficult tasks but are not measured by an IQ test.
Comment by spider-mario 4 days ago
Comment by ranyume 4 days ago
I think the key here is to be wary of measurements that promise to capture the whole of what we consider intelligence (ie what people think of with IQs).
Comment by mdp2021 2 days ago
And as we said, the "IQ tests" do not promise that.
> what people think
What people think is very irrelevant to truth - and we just disregard their opinion. It cannot have a weight.
What would be the gain in asking a layman about anything. It just makes no sense.
Comment by machomaster 3 days ago
Not being able to deadlift 500kg also signifies the inability to effectively the (physical) environment.
But even here, in your irrelevant example, you are wrong. In practice IQ is highly correlated with emotional skills ("intelligence").
Comment by spider-mario 4 days ago
Comment by mdp2021 4 days ago
Nobody says that the IQ test would "measure intelligence". We know it does test a form of it.
Comment by ranyume 4 days ago
Comment by mdp2021 4 days ago
Comment by mercer 3 days ago
Comment by spider-mario 4 days ago
Comment by jawiggins 5 days ago
Comment by eli 5 days ago
Comment by Davidzheng 4 days ago
Comment by ranyume 5 days ago
Comment by malfist 5 days ago
Comment by hyperhello 5 days ago
Comment by whattheheckheck 5 days ago
Comment by jrflo 5 days ago
Comment by rcoveson 5 days ago
Comment by malfist 5 days ago
LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.
So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.
This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs
Comment by mdp2021 5 days ago
Not fully relevant: timing is crucial in all-pass tests, not crucial in pass-or-fail tests. I.e.: first of all, they have to be able to reach the goal, and that is already an achievement. Then - and in parallel - the problem solving must also be optimized for efficiency. But "solving" and "efficiency" are non coincident dimensions.
Comment by paimapi 5 days ago
I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations
[0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...
Comment by mdp2021 5 days ago
It spells out a form of intelligence - some can and some cannot.
Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.
Comment by paimapi 4 days ago
Comment by mdp2021 4 days ago
more intelligent (in some specific dimensions of Intelligence) than what gets worse marks. Yes, pretty informally - but still notably.
Comment by dist-epoch 5 days ago
Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.
Comment by GaggiX 5 days ago
Comment by nimchimpsky 5 days ago
Comment by Betelbuddy 5 days ago
Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."
Well I dont know about all of you, but I am celebrating meat based humans...
Comment by LPisGood 5 days ago
Comment by paxys 5 days ago
Comment by fastball 5 days ago
Comment by bigzyg33k 4 days ago
Comment by modeless 4 days ago
Comment by WASDx 4 days ago
Comment by dwohnitmok 5 days ago
Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"
Comment by 6thbit 5 days ago
none 35.2%, $49,791 96.7%, $23,457
35.2% on the standard harness, that's above Opus 5 on high.Comment by NitpickLawyer 5 days ago
Comment by an0malous 4 days ago
Comment by an0malous 4 days ago
I wouldn’t put it past a company like OpenAI with a long history of lying and being deceptive to record the tests and benchmaxx ARC. They have trillions of dollars of incentive to cheat any way they can.
Comment by piloto_ciego 5 days ago
Prediction:
We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.
Comment by WASDx 5 days ago
TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".
Comment by emp17344 5 days ago
Comment by hn_throwaway_99 1 day ago
Also, while AI's effect on jobs may be currently overstated, it's also obviously having an impact. New CS grads and junior developers are having a hell of a time finding jobs now.
Comment by piloto_ciego 5 days ago
Comment by raspasov 5 days ago
(I have not gone down the rabbit hole to understand how they achieve that 24% number)
Comment by _superposition_ 5 days ago
Comment by cute_boi 4 days ago
If government don't step up and regulate outsourcing like 100% tax, big problems are coming up.
Comment by asadotzler 4 days ago
Comment by brokensegue 4 days ago
Comment by jhonof 5 days ago
Comment by NitpickLawyer 5 days ago
Comment by piloto_ciego 5 days ago
Comment by slopinthebag 5 days ago
Comment by piloto_ciego 5 days ago
IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!
Comment by slopinthebag 4 days ago
Comment by defrost 4 days ago
Comment by slopinthebag 4 days ago
Comment by defrost 3 days ago
You must have some notion of an argument as to why all the varied attributes of different humans should align to all be "average" in some specific human.
O/wise you'd be peddling sloppy trite aphorisms that don't bear scrutiny.
Comment by slopinthebag 3 days ago
Comment by dgellow 5 days ago
Comment by baal80spam 5 days ago
It's already happening :)
Comment by piloto_ciego 5 days ago
Comment by lofaszvanitt 4 days ago
Comment by unixhero 3 days ago
Comment by mikert89 5 days ago
Comment by tedsanders 5 days ago
Examples:
- predict a coinflip: easy to verify, hard to learn
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.
Comment by ranyume 5 days ago
Comment by mikert89 5 days ago
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
but we both know these examples go against the spirit of my point
Comment by tedsanders 5 days ago
Comment by mikert89 5 days ago
also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye
Comment by _superposition_ 5 days ago
Comment by imtringued 4 days ago
The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.
Comment by mikert89 4 days ago
Comment by tomjen3 4 days ago
Comment by jdthedisciple 5 days ago
Comment by mikert89 4 days ago
Comment by scotty79 4 days ago
Comment by brokensegue 4 days ago
Comment by yomismoaqui 5 days ago
Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)
Comment by wise_blood 4 days ago
once 3 is solved, we would come up with 4. then 5, 6...
it will be AGI when we cannot come up with a task easy for human but hard for machines. thet's the whole point.
Comment by john_alan 4 days ago
Comment by hypfer 5 days ago
Comment by petu 5 days ago
Comment by Frost1x 5 days ago
Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.
Alignment++
Comment by manquer 4 days ago
There is no incentive for OpenAI to subsidize is you if no one reads /reports on your benchmark . They are only going to fund a few that are currently popular .
Community acceptance doesn’t automatically mean the best , it is combination of some level of technical quality and the ability of the promoter to socially influence or get support of influencers .
Comment by bigbuppo 5 days ago
Comment by yusufozkan 5 days ago
Comment by minimaxir 5 days ago
Comment by Frost1x 5 days ago
The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).
Comment by Phemist 5 days ago
Comment by tedsanders 5 days ago
Comment by minimaxir 5 days ago
Comment by fxd 5 days ago
I’ve ignored it thinking it would go away, but it keeps coming up.
I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery.
Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get it.
But you’d be no nearer to solving consciousness.
Given this thought trajectory - what is AGI supposed to be?
Comment by layer8 4 days ago
Intelligence, and hence AGI, doesn’t require consciousness or emotions or sentience.
Comment by p1esk 4 days ago
What human intelligence do you mean? Genius? Professional? Educated? Random person? “Dumb” person?
Comment by HarHarVeryFunny 1 day ago
Was the genius or dumb person really born with different levels of learning ability, or were they just raised differently: nature vs nurture ?
The "intelligence" of one species vs another comes down to differences in cognitive architecture, ultimately reflected in ability to learn and predict/infer. Intelligence, as a capability, not IQ test score, is best regarded as ability to learn from experience and use that learning to accurately predict future outcomes, ranging from passive observation, to the outcome of one's own actions, to the ability to reason.
Comparing the intelligence (not knowledge) of different AI system to humans should therefore be assessed by comparing their ability to learn and use what they have learnt.
An LLM is what it is - a language model, not a learning system. It's really an expert system of sorts, highly capable in terms of what it can infer based on the knowledge it encodes, but with very limited ability to learn.
When you talk of comparing AI to a genius/professional/graduate/etc, you are really talking about comparing level of knowledge, comparing one expert system to another, which is fine and perhaps useful in some contexts, but it is not the same as comparing actual intelligence - learning ability, and especially so if you want to discuss general intelligence which is all about the ability to successfully take on any task (perhaps needing to learn it first), not just do well on some limited set of tasks you are already familiar with.
If ability to learn is limited to one modality, such as language, then that indicates a lack of generality. A good test for whether an AI system is in the same ballpark as a human in learning ability, aka intelligence, would be whether it can (at run-time) learn language itself, from a blank slate start, when running in a suitable environment.
Comment by famouswaffles 1 day ago
So a transformer ? Run-time? Seems like an arbitrary distinction to me. Because humans function in a certain way, every learning system no matter how capable must function in the same way to be 'truly' intelligent ?
Comment by HarHarVeryFunny 22 hours ago
It's not a very high bar - even a rat has the basic ability to learn.
Comment by famouswaffles 21 hours ago
Comment by HarHarVeryFunny 11 hours ago
How is text-based memorization going to substitute for learning? The two are not the same. Perhaps this is more applicable to robotics than a text generator, but I also doubt an LLM could learn text-based skills like programming or math if it had not been pre-trained on them via SGD & RL, and had to instead rely on some poor-man's-learning context-based recall instead. What else can't it learn? How is the LLM intern, the "drop-in replacement remote worker" going to do on day #2?
Instead of pretending that an LLM can be human-level, or super-human, or become generalist, why not just admit that this is not the final form of AI. An LLM is not an animal/human-like intelligence, it is something different - a language model, with it's own strengths and weaknesses.
Despite all the AGI hype, an LLM seems to have more in common with a pre-trained single-purpose system like AlphaGo than a brain, but with the rubric/reward-based policy function baked into the weights.
In another 10-20 years some new idea, hopefully more brain-like, will have superseded LLMs and they will indeed be labelled as "LLMs" as the AI/AGI label becomes attached to the new more brain-like creative intelligence. Perhaps it'll be sooner than 10-20 years, but I doubt it given the current 10-year fixation with LLMs which doesn't appear to be slowing down anytime soon. Perhaps Sutskever is working on something a bit different?
Comment by famouswaffles 7 hours ago
I don't know. How is Astra a step change in computer use and spatial reasoning to the extent it can play games, paint good looking stuff with e.g canva and a whole number of other things ?
>How is text-based memorization going to substitute for learning? The two are not the same.
Of course if you call it something else then you can say it's not the same.
>How is the LLM intern, the "drop-in replacement remote worker" going to do on day #2?
Just fine I imagine ? ICL and the memory tools around a harness are pretty good. I'm not sure what sort of magic you're expecting from the human, but they're not getting any improvement in that time frame a frozen transformer can't match.
>Instead of pretending that an LLM can be human-level, or super-human, or become generalist, why not just admit that this is not the final form of AI.
It doesn't seem like I'm the one pretending here.
>Despite all the AGI hype, an LLM seems to have more in common with a pre-trained single-purpose system like AlphaGo than a brain, but with the rubric/reward-based policy function baked into the weights.
If you say so.
>Perhaps it'll be sooner than 10-20 years, but I doubt it given the current 10-year fixation with LLMs which doesn't appear to be slowing down anytime soon.
The architecture that keeps delivering results isn't slowing down ? You don't say.
There's no shortage of people, even researchers, who for one reason or the other are convinced we are in need of some paradigm shift.
But guess what? Talk is cheap. You beat the current paradigm or you don't.
Comment by HarHarVeryFunny 5 hours ago
Presumably because of pre-release training, because some alien outside of the model, armed with the reinforcement learning algorithm, came in and programmed its weights.
> The architecture that keeps delivering results isn't slowing down ? You don't say
Sure, nothing wrong with that, as long as you don't misrepresent the limitations of the approach.
> There's no shortage of people, even researchers, who for one reason or the other are convinced we are in need of some paradigm shift.
> But guess what? Talk is cheap. You beat the current paradigm or you don't. Do you seriously think that Meta never scaled JEPA ?
The idea has not been taken very far, so what is there to scale? It's not a complete cognitive architecture. So far it's also been using a pre-trained transformer as the learning component, which makes it of limited interest.
The animal intelligence approach, even in it's most fledgling form (that you would apparently dismiss), requires a complete agentic architecture, including new learning algorithms and generative behavior, before it can be compared to LLMs. We know that, done right, our brain architecture is more capable than an LLM, so even if any hypothetical attempts to reproduce it were not highly performant, we know that the idea itself is sound. You might compare with Uszkoreit's initial poor-performing implementation of his new language model architecture - should he have given up?
Comment by famouswaffles 4 hours ago
They didn't program anything. They gave it data at best.
>The idea has not been taken very far, so what is there to scale?
LeCun was the head of Meta AI for over a decade and his baby that he keeps harping on about wasn't taken very far ? Come on. You're smarter than that. It went the way all the alternate architectures have gone since the transformer, a sidegrade at best, probably not even that.
>You might compare with Uszkoreit's initial poor-performing implementation of his new language model architecture - should he have given up?
What are you talking about? There was no poor performing implementation of transformers that was published that he needed to push through. Are you talking about self attention experiments before the finished transformer? That's literally just research. And if he languished on that for a decade then yeah I'd tell him to probably look at something else, but of course he didn't.
Comment by HarHarVeryFunny 3 hours ago
Have you looked at all the published JEPA research both while LeCun was at Meta, and since (up to and including the latest AdaJEPA from June)? Please enlighten us as to exactly which line(s) of research you think were "scaled" at Meta, and then tell us which of these constituted anything even remotely resembling a complete testable intelligence?
FYI, it's been a long time since FaceBook/Meta even had a single head of AI. Since 2018 it has been split into two groups, FAIR and Generative AI, with LeCun being in the FAIR group, not as head, but as Chief AI scientist. LeCun only invented JEPA in 2022 (shortly after FaceBook became Meta), first writing about it in his "A Path Towards Autonomous Machine Intelligence" paper, perhaps unhappy with the work of the GenAI group, which he had no control over, that presumably was getting all the compute.
https://openreview.net/pdf?id=BZ5a1r-kVsf
> What are you talking about?
I was referring to Uszkoreit's personal telling (on YouTube) of the origin story of the Transformer, his motivations with the design, his initial personal failure to implement his idea in a performant enough manner to beat the current LSTM SOTA, and Noam Shazeer then throwing the kitchen sink at it and eventually coming up with the Transformer design.
Comment by famouswaffles 1 hour ago
Yeah. Presumably, lots of synthetic data is being generated, experiments being run, but post-training is still a largely automated process.
>Have you looked at all the published JEPA research both while LeCun was at Meta, and since (up to and including the latest AdaJEPA from June)? Please enlighten us as to exactly which line(s) of research you think were "scaled" at Meta, and then tell us which of these constituted anything even remotely resembling a complete testable intelligence?
I have. My point isn't that his ideas are trash or that he should stop working on them. My point is it's not "gone very far" because he's taking it as far as he can, which isn't very far. He's not had a lack of influence, resources or will, either from his time at Meta or now with his billion dollar startup. He's had far more of it than most, if anything. That there's not much to show for it so far is not for a lack of trying.
>I was referring to Uszkoreit's personal telling (on YouTube) of the origin story of the Transformer, his motivations with the design, his initial personal failure to implement his idea in a performant enough manner to beat the current LSTM SOTA, and Noam Shazeer then throwing the kitchen sink at it and eventually coming up with the Transformer design.
So it's what I thought. This is just regular research unless an inordinate amount of time was spent on it and that's not the case.
Comment by HarHarVeryFunny 17 minutes ago
I wouldn't really agree - I'm no fan of LeCun, but the problem with JEPA isn't that it's a bad idea, or can't go very far, but just that it's not much of an idea in the first place!
It's no secret that our brain basically works by prediction, and what we're predicting is necessarily the external world as we perceive it though our own senses, aka latent representations, aka JEPA.
So, you COULD take this unoriginal smidgen of an idea and built it out to a full model of a human/animal brain, whether or not it's LeCun's intention to do so (he seems more interested in just the representational / world model aspect to it), but he certainly hasn't done so yet, nor created any research manifesto indicating that as his intent.
The fact that JEPA implementations to date are using pre-trained Transformers doesn't seem inherent to the approach - one could, with more effort, still predict latent representations (i.e. sensory feedback) but do so using a new real-time learning algorithm based on prediction failure.
LeCun seems more of an academic / research director than a builder, and I would never have put much stock in him being the one to build an animal brain.
Comment by drdeca 4 days ago
They are still different concepts of course, but I imagine that once one is achieved, the others aren’t far off.
Comment by p1esk 4 days ago
I’m just trying to understand the implications of the current frontier model capabilities.
Comment by drdeca 3 days ago
But there are also many cognitive tasks that the typical educated person would do better at than these models.
Like, e.g. long-term managing what a vending machine gets stocked with and what prices the items should be sold for. Or, various things like that.
When there are essentially no more tasks like that, then we’ve reached AGI.
Comment by p1esk 3 days ago
Comment by fxd 4 days ago
Comment by meander_water 4 days ago
"highly autonomous systems that outperform humans at most economically valuable work"
https://time.com/article/2026/08/26/openai-sam-altman-interv...
Comment by fxd 4 days ago
We cannot in a declarative sense define what is economically valuable work even now let alone into the future.
People take what they can get for pay. Very few individuals can demand a wage. The value of employment is obviously designed around that, not some arbitrary definition of “valuable”.
Of course an AI will accept $0/hr, it doesn’t mean it does the job.
Anyone who could accurately define the value of work would be wildly successful without having to try.
That is not a useful definition for me unfortunately.
Comment by jryle70 4 days ago
Do you want to take a stab at defining/quantifying it? I'm s afraid anything specific you can come up with will also be useless.
Comment by fxd 4 days ago
I said it’s impossible to define that value, not intelligence. I’ve seen a dozen or so useful definitions of intelligence.
Comment by benlivengood 4 days ago
The consensus now seems to be that once you've got human-level intelligence and planning and executive function then you get recursive self-improvement that can eventually autonomously solve the robotics and world-modeling and other portions of human-equivalence.
Comment by eagerpace 5 days ago
Comment by fxd 4 days ago
I wonder at what point consciousness is necessary… that is, if you can have anything like that without it.
To the point that solving consciousness (and combining it with intelligence) is what gives you the autonomous, recursive, self-improving thing otherwise it can only drive in the dark and make big mistakes.
To your point I think - it’s why we don’t see too many non-conscious advanced biology (it rarely survives against those with it).
Comment by mdp2021 4 days ago
How can you conflate the two.
> solving consciousness
We are very much not interested in that. We just need a proper problem solver.
Comment by submain 4 days ago
Comment by mdp2021 4 days ago
Comment by fxd 4 days ago
For example, how do you know that “feeling pain” is not a functional prerequisite for a task. And that consciousness is a prerequisite for feeling pain
Comment by mdp2021 4 days ago
Comment by grantcas 4 days ago
Comment by StopTheMods 4 days ago
Comment by ajjahs 4 days ago