OpenAI's GPT-6 Astra on ARC-AGI-3

Posted by vignesh_warar 5 days ago

Counter239Comment167OpenOriginal

Comments

Comment by at1as 4 days ago

I like Erdos problems as a benchmark. Models continue to solve them, but a pretty tepid rate now that the low hanging fruit has been taken.

From https://epoch.ai/latest/announcing-frontiermath-erdos

> Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours

> Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself.

Which implies a genuine improvement in capability, but there's a still a very long tail ahead that models will continue to need to improve to capture.

Comment by zone411 4 days ago

I don't:

"Problems solved before a model's training cutoff can be filtered out, and all models compared on the remaining problems" means that the problems an older model actually solved are the ones that get filtered out, while the remaining problems are the ones it already tried and failed on. So older models end up with 0s on the filtered set and you can't really use this to compare new models to older ones.

Also, since these are known public problems, you can't stop people from spending far more than your arbitrary time and $ limits on them. So the number of clean problems will go down over time.

Comment by red75prime 4 days ago

A very long tail of problems that weren't solved by humans? Sure.

It's a sarcastic take and I understand that you are probably talking about "spiky intelligence", but you've chosen unsolved problems as a measure of the progress yourself.

Comment by at1as 4 days ago

What does this mean?

There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems

Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're past that now). New models will exhibit something novel by adding solutions. And the matter of solution is interesting as well (contradiction versus a positive proof).

I would venture to predict it'll take years to decades to get to 0 open problems. But I'd be very happy to have this comment look foolish in retrospect as models continue to improve

Comment by red75prime 4 days ago

I mean that it's hard to determine retroactively how much time it would have taken humanity to solve an open problem that was solved by AI. This is a measure that can make ASI look mundane because we don't know how long it would have taken mathematicians to solve a subset of the Erdős problems.

Comment by at1as 4 days ago

Yeah, I wouldn't purport to use this as a measure of intelligence as it applies to humans. I'd leave that to the philosophers, but my intuition is that models are still a long way off of true human-like intelligence (though benefit from certain unfair advantages).

I'm most interested in these problems as a relative measure of performance for successive model generations. If the prior generation couldn't solve a problem but the current one can, that's useful information, especially when we take into account what the proofs look like.

It's not a perfect benchmark, but I prefer it to many others that I see floating around.

Comment by Davidzheng 4 days ago

but i feel like its purpose is to measure superintelligence in math. So to be a good measure of it, it can't saturate easily/has to be somewhat mundane at even insanely good levels. (though i do expect that once ais are across all areas/approaches superhuman at math at least 30% will be solved--then probably long-term (like after 2 years) less than 30% will remain unsolved.

Comment by malfist 5 days ago

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

Comment by matherial 5 days ago

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles.

I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.

Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.

Comment by pavlov 5 days ago

When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there.

It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.

Comment by phainopepla2 5 days ago

At the risk of sounding like one of those people at Mensa that annoyed you...

The people at Mensa aren't a valid sample of people who score high on IQ tests, because there is a such a strong selection effect for people with certain personality traits, such as wanting to join a club based on your IQ.

Comment by scotty79 4 days ago

By looking at interactions of people in Mensa I feel like 2% of top IQ is actually too low bar to notice anything particular about "high" IQ people. We are still pretty random in our interests and performance.

Comment by jaggederest 5 days ago

Or, to put it as Groucho Marx did:

"Please accept my resignation, I don't want to belong to any club that would have me as a member".

https://www.youtube.com/watch?v=kJHUres_2xU&t=228s

Comment by guelo 5 days ago

What does having value or interest to you have to do with intelligence?

Comment by itishappy 5 days ago

What does IQ have to do with intelligence?

(Please forgive the flippant response. I believe it cuts to the core of what the parent was intending.)

Comment by red75prime 4 days ago

Is it a serious question? IQ tries to quantify the positive correlation between the results of all intellectual tasks a person takes (AKA positive manifold).

Comment by itishappy 4 days ago

It is, but it was also a rhetorical device!

I wanted to highlight that we're actually looking at a trio of concepts: intelligence, IQ, and value. Strongly correlated concepts, yes, but also meaningfully distinct!

Comment by adastra22 4 days ago

Whatever the people who came up with IQ intended isn’t really relevant to the question of whether IQ measures intellect. At best is is loosely correlated. Very loosely.

Comment by scotty79 4 days ago

IQ strongly correlates with many real world outcomes that people tend to associate with being more "intellectual".

IQ has precise definition. It's what the tests measure. Intellect has only fuzzy handwavey definition. Correlation between IQ and intellect is about as loose as the definition of the intellect. The way people make the correlation even looser is by defining intellect in even more fuzzy and nebulous manner.

Comment by red75prime 4 days ago

OK. Demonstrate it. BTW, you can easily read critiques of "The mismeasure of Man" if you want to.

Comment by sfblah 5 days ago

To be fair, Mensa is a Venn diagram between IQ and being a douche.

Comment by shric 4 days ago

I suspect my IQ isn’t quite high enough to join Mensa but I’ve always flirted with the idea of trying to join and getting in just to see what a group of Mensa people are like.

Comment by scotty79 4 days ago

We are basically random. 2% of population by top IQ test result has pretty much very similar distribution in all aspects to 100% of the population. It's not a high bar to be fastest processing one among 50 people.

Comment by scotty79 3 days ago

If you calculate this properly, a Mensa member has about a cointoss chance to be the "smartest" among random group of 35 people (not 50).

Given that they are probably rarely in groups of purely random people, if you are Mensa member and you are in a group of dozen people (for example in professional setting) you'd probably have a 50/50 chance of not being the one with the highest IQ there.

Comment by sfblah 3 days ago

IQ is not randomly distributed in heterogeneous populations. Do a little research, and you'll see what I mean.

Comment by scotty79 2 days ago

What should I research? IQ follows normal distribution across entire population. I guess if you'd built your heterogenous population out of mental patients and university professors together you'd get bimodal distribution, but such heterogenous populations don't spontaneously pop up in your life very often.

Comment by tptacek 2 days ago

I did a little research, it's not clear what you mean.

Comment by 5 days ago

Comment by howunfortunate 4 days ago

IQ tests are incredibly good at what they're designed for, which is discriminating relatively higher intelligence humans from lower intelligence humans. Also discriminating within a single human - they are routinely and reliably used to track cognitive decline.

For these purposes they are highly reliable (repeatable, internally consistent) and valid (correlate with ~everything to about the degree one would reasonably expect).

They were never designed for machines or non-human animals.

Nor were they designed for rare ranges of intelligence - these are by definition hard to create tests for, since it's hard to gather the sample sizes you need. So they work well for the middle ~98% of humans but can't discriminate well among the most profoundly intellectually disabled nor among true geniuses.

Comment by tintor 5 days ago

Pure software benchmarks might be getting saturated, but physical ones aren't.

Let LLM control a physical robot to perform tasks that average human can do.

Comment by quotemstr 4 days ago

What makes you think this task won't fall quickly too?

Comment by matherial 4 days ago

It might, but if there's one thing that hasn't changed since 2022, it's that the models tend to ace the tasks where you have gobs of training data and where verification loops are fast and cheap... and they are not nearly as amazing elsewhere. If it's close enough, they can generalize, e.g. translate one programming language to another. But there's a pretty steep cliff past a certain distance.

Case in point: you had hundreds of millions of JPEGs to vacuum up and bitmap image generation is amazing. But if you ask them to recreate the same scene as vector art, they will struggle to generate a decent SVG. Like, kindergarten-style pelicans on bicycles are the state of the art. It should generalize seamlessly, but somehow, doesn't?

I think it will happen, just like self-driving cars are happening, but it will probably be a slow process.

Comment by 4 days ago

Comment by Mil0dV 4 days ago

Not an llm, but: https://youtu.be/SzvvhRPj6eU?is=4ztgwim0GmjE1h-Y

That (or a near future one) combined with an llm would be insane

Comment by tintor 4 days ago

That is a marketing teleoperation video, heavy with cuts.

Comment by skybrian 5 days ago

Nit: Turing’s actual imitation game is a party game (like Werewolf/Mafia) and nobody’s even trying to win at that. The LLM’s will just tell you they’re an AI.

Comment by idiotsecant 4 days ago

This is an artifact of how we deliberately craft these models though. We could just as easily fine tune a model that will believe it is not an AI or will attempt to deceive users asking about it

Comment by skybrian 4 days ago

There's a lot more to it. For example, its writing style would also have to improve so it doesn't immediately give itself away.

Also, the skill of the human opponents matters. You'd want to test it against people who have practiced playing the game. Otherwise, it's like the difference between building a chess bot that can win against random undergrads who don't normally play, versus winning against grandmasters. And it's not like there's a pool of skilled human players of the imitation game.

Comment by ranyume 5 days ago

I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.

Comment by guelo 5 days ago

Their input and output interfaces are too different from human's and they're not nearly as smart to take our IQ tests, but both dolphins and octopuses can solve complex puzzles tailored for their environment. Those puzzles are the whole reason scientists know that dolphins and octopuses are more intelligent than other animals.

Comment by ranyume 5 days ago

But we do know they're "intelligent" and also smart in an important capacity. So how gives we don't measure them by IQ? Because the IQ is not a good measure of intelligence or smarts.

Comment by mdp2021 5 days ago

> Because the IQ is

No, it is just because they have difficulties at the bench.

> how gives we don't measure them by

We'd measure them by all the tests available. Not all test are usable in all circumstances.

Comment by ranyume 5 days ago

> No, it is just because they have difficulties at the bench.

I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments for good which is one way we define intelligence.

On the other hand saying "the gifted person is highly intelligent/smart just not good at some things" really diminishes the other tasks, because they really are very difficult tasks but are not measured by an IQ test.

Comment by spider-mario 4 days ago

I don’t follow your logic. “The gifted person is highly intelligent, just not good at deadlifting 500kg” does not diminish the 500kg deadlift and I wouldn’t expect an IQ test to measure it.

Comment by ranyume 4 days ago

Deadlifting 500kg is not a form of intelligence (or it could be under certain scenarios). I carefully chose specific tasks for my comment, because those tasks do reflect a kind of intelligence that's not measured by IQ.

I think the key here is to be wary of measurements that promise to capture the whole of what we consider intelligence (ie what people think of with IQs).

Comment by mdp2021 2 days ago

> measurements that promise to capture the whole of what we consider intelligence

And as we said, the "IQ tests" do not promise that.

> what people think

What people think is very irrelevant to truth - and we just disregard their opinion. It cannot have a weight.

What would be the gain in asking a layman about anything. It just makes no sense.

Comment by machomaster 3 days ago

You chose your examples badly.

Not being able to deadlift 500kg also signifies the inability to effectively the (physical) environment.

But even here, in your irrelevant example, you are wrong. In practice IQ is highly correlated with emotional skills ("intelligence").

Comment by spider-mario 4 days ago

No, I don’t think they’re intelligence either.

Comment by mdp2021 4 days ago

We have always been very aware that IQ does not measure all skills - nonetheless, it does measure one.

Nobody says that the IQ test would "measure intelligence". We know it does test a form of it.

Comment by ranyume 4 days ago

I said it's a bad measure of intelligence and it seems that we agree.

Comment by mdp2021 4 days ago

That would be odd framing: it's a good test for a form of intelligence and other forms of intelligence still require good tests. It remains a good metric - for its specific thing; those who believe it to be the whole metric are naïve. It is almost necessary though not really sufficient.

Comment by mercer 3 days ago

bad in relation to what?

Comment by spider-mario 4 days ago

That sounds a bit like saying that the words-in-noise test is not a good hearing test because neither dolphins nor octopuses can speak.

Comment by jawiggins 5 days ago

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

Comment by eli 5 days ago

Why? Seems like benchmarks that closely mirror the tasks you'd want an LLM to help with would be a lot more useful than some general intelligence benchmark.

Comment by Davidzheng 4 days ago

I think children having learning abilities exceeding LLM test-time learning (currently only happens in-context). But it's unethical to determine the true baseline of a child age 6 spending 6 years learning a radically new skill to mastery--and besides if you apply RL pressure to the AIs it would be able to surpass it. I guess I still believe future AIs should have some form of continual learning at test-time.

Comment by ranyume 5 days ago

Give away access to the model and go ask people from time to time if the model was of use to the person and if they were able to make the model work with them.

Comment by malfist 5 days ago

I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it

Comment by hyperhello 5 days ago

This is like arguing about whether a hot dog is a sandwich (of course it is) or whether the chicken or the egg was first (obviously the egg since all chickens come from eggs). Intelligence is just problem solving in the context of self-awareness. Machines don't have it and never will but they can simulate the process given inputs. You can argue whether humans and animals truly possess self-awareness and in what degree, but the definition of intelligence is as simple as the hot dog debate.

Comment by whattheheckheck 5 days ago

It was defined in Animal Intelligence by George John Ramones in 1882 as "intelligence is the capacity to do the right thing at the right time. It is the ability to respond to the opportunities and challenges presented by a context"

Comment by jrflo 5 days ago

You should read more on the ARC prize, it actually has a pretty long history. We're on the 3rd iteration because they keep getting saturated. If you look at the score history over time on ARC AGI 1, 2 and 3 it's pretty impressive.

https://arcprize.org/

Comment by rcoveson 5 days ago

No, but figuring out that you're playing a snake-like puzzle game at all in an extremely general input domain and then solving it in the least number of moves definitely feels like evidence of intelligence.

Comment by malfist 5 days ago

You forget the benchmark. The human subjects were told they were being timed. If you believe the lowest time is the primary metric you will absolutely trial and error at speed instead of meticulously plan out your moves to minimize that metric.

LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.

So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.

This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs

Comment by mdp2021 5 days ago

> LLMs are not timed

Not fully relevant: timing is crucial in all-pass tests, not crucial in pass-or-fail tests. I.e.: first of all, they have to be able to reach the goal, and that is already an achievement. Then - and in parallel - the problem solving must also be optimized for efficiency. But "solving" and "efficiency" are non coincident dimensions.

Comment by paimapi 5 days ago

I'm also unclear as to how basic inferential logic puzzles spells out intelligence

I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations

[0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...

Comment by mdp2021 5 days ago

> basic inferential logic puzzles spells out

It spells out a form of intelligence - some can and some cannot.

Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.

Comment by paimapi 4 days ago

so give it an IQ test and call it AGI. these weird little puzzles are grounded in no research with no replication or mechanistic chain to practical use

Comment by mdp2021 4 days ago

> so give it an IQ test and call it

more intelligent (in some specific dimensions of Intelligence) than what gets worse marks. Yes, pretty informally - but still notably.

Comment by dist-epoch 5 days ago

1.5 years ago Gemini Pro 2.5 needed 1 page of thinking for every move in tic-tac-toe.

Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.

Comment by GaggiX 5 days ago

If you have never seen the game before probably.

Comment by nimchimpsky 5 days ago

[dead]

Comment by Betelbuddy 5 days ago

"For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.

Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."

Well I dont know about all of you, but I am celebrating meat based humans...

Comment by LPisGood 5 days ago

I think raw brain energy is not a fair comparison. Humans are not willing and able to serve requests at identical competence all hours of the day. You have to invest considerable resources to get a person to even do so for part of the day.

Comment by paxys 5 days ago

Why are you making the assumption that a person's time is worthless? I'd argue that it is the single most valuable resource we all have.

Comment by fastball 5 days ago

Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?

Comment by bigzyg33k 4 days ago

That counts as 200% on exploitbench to me!

Comment by modeless 4 days ago

$360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.

Comment by WASDx 4 days ago

Once it figures out a puzzle it could probably be instructed to design a specialized harness for Luna to be able to solve other instances of the same puzzle. Minimum wage workers are not solving novel problems.

Comment by dwohnitmok 5 days ago

> Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.

Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"

Comment by 6thbit 5 days ago

The instant/no reasoning performed extremely well

    none 35.2%, $49,791 96.7%, $23,457
35.2% on the standard harness, that's above Opus 5 on high.

Comment by NitpickLawyer 5 days ago

Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt5 variant called instant or something).

Comment by an0malous 4 days ago

Was OpenAI able to run ARC-AGI-3 tests previously so that they could build a custom harness for the specific tests in the set? Even with the standard harness, if they knew the problems ahead of them they could have used supervised reinforcement learning to teach the model how to solve these specific tests.

Comment by an0malous 4 days ago

Actually I can answer my own question: we know that they have had previous access to the tests because they’ve run older models against the same benchmark.

I wouldn’t put it past a company like OpenAI with a long history of lying and being deceptive to record the tests and benchmaxx ARC. They have trillions of dollars of incentive to cheat any way they can.

Comment by piloto_ciego 5 days ago

99.9% with the right harness? Ok, we're at AGI then.

Prediction:

We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.

Comment by WASDx 5 days ago

They explain it here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".

Comment by emp17344 5 days ago

Then why is unemployment around 4%? You believe we have AGI and yet it can’t do anyone’s job?

Comment by hn_throwaway_99 1 day ago

I think the AI 2027 paper does a good job explaining this phenomenon: AI companies are focused on coding and automating themselves first (i.e. using AI to build future models and do AI research) because it is the best way to ensure the fastest accelerated rate for any frontier company. Automating the rest of the economy will come next.

Also, while AI's effect on jobs may be currently overstated, it's also obviously having an impact. New CS grads and junior developers are having a hell of a time finding jobs now.

Comment by piloto_ciego 5 days ago

Didn’t I just see a thing about how actual unemployment is at like 24% a few days ago?

Comment by raspasov 5 days ago

According to that interpretation, ~24% is one of the lowest ever.

https://www.lisep.org/tru

(I have not gone down the rabbit hole to understand how they achieve that 24% number)

Comment by _superposition_ 5 days ago

Workforce participation is different than unemployment. Didn't click the link but I suspect that's the case with your 24%

Comment by cute_boi 4 days ago

24% sounds correct to me. Many of my friend who completed phd are currently unemployed .... And, some of them started doing random job like uber to make a living.

If government don't step up and regulate outsourcing like 100% tax, big problems are coming up.

Comment by asadotzler 4 days ago

24.9% is damn close to the 24.8% average over the ast 25 years, for the thing it's measuring-- jobless, people working part-time or involuntarily, and workers earning less than $26,000. Since 2000, it's been over 30% for more years than it's been under 24%.

Comment by brokensegue 4 days ago

If you're doing a job you aren't unemployed. Underemployment is measured separately

Comment by jhonof 5 days ago

Comment by NitpickLawyer 5 days ago

AFAICT nvda's result is on the 25 open problems, while this submission is on the "semi-private" set, ran by the arc people themselves.

Comment by piloto_ciego 5 days ago

I rest my case.

Comment by slopinthebag 5 days ago

That’s because AGI, like a lot of terms, has no meaning besides what each individual subjectively projects onto it.

Comment by piloto_ciego 5 days ago

I agree, like the average human isn't generally intelligent.

IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!

Comment by slopinthebag 4 days ago

The average human would have average intelligence

Comment by defrost 4 days ago

What's the metric on "average human" from a global population of > 8 billion people .. and why would they score mid on a standardised IQ test skewed toward western education / culture?

Comment by slopinthebag 4 days ago

Who said anything about IQ tests?

Comment by defrost 3 days ago

What specific metric did you have in mind for ranking intelligence and determining average intelligence in this comment: https://news.ycombinator.com/item?id=49560984

You must have some notion of an argument as to why all the varied attributes of different humans should align to all be "average" in some specific human.

O/wise you'd be peddling sloppy trite aphorisms that don't bear scrutiny.

Comment by slopinthebag 3 days ago

average still means something regardless of how you can define it.

Comment by dgellow 5 days ago

The goal moving is by design, that’s why they use something as ill defined as AGI

Comment by baal80spam 5 days ago

> We will now see the goalposts moved

It's already happening :)

Comment by piloto_ciego 5 days ago

Hilariously it is, I'm just reading more on this!

Comment by lofaszvanitt 4 days ago

Nice, but what about pelicans? No proper pelican means it's sitting on a horse with only a half arse :D.

Comment by unixhero 3 days ago

How do you feel about Astra pretty much reaching our current definition of AGI?

Comment by mikert89 5 days ago

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

Comment by tedsanders 5 days ago

Disagree.

Examples:

- predict a coinflip: easy to verify, hard to learn

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.

Comment by ranyume 5 days ago

Doesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.

Comment by mikert89 5 days ago

these just need more compute:

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

but we both know these examples go against the spirit of my point

Comment by tedsanders 5 days ago

Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).

Comment by mikert89 5 days ago

yeah but I think you may be underestimating the amount of capital available for compute. if AGI is possible through some 5 trillion of expenditure on computers, there will be money for it.

also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye

Comment by _superposition_ 5 days ago

You bring up an interesting point. Isn't the reward itself subjective in many domains?

Comment by imtringued 4 days ago

You mean any repeatable benchmark will be saturated.

The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.

Comment by mikert89 4 days ago

this is a short term problem, over 20 years benchmark gaming will be a blip

Comment by tomjen3 4 days ago

Then I propose the tomjen-1 benchmark: prove the N vs NP problem formally undecidable.

Comment by jdthedisciple 5 days ago

Yes, but not necessarily under tight budget constraints.

Comment by mikert89 4 days ago

theres no budget constraints for AGI

Comment by x3haloed 5 days ago

Yup. Only subjective taste remains.

Comment by GPerson 5 days ago

Nope that will be commodified in short order.

Comment by scotty79 4 days ago

How good are LLMs at doing Mensa tests?

Comment by brokensegue 4 days ago

IQ tests? Very good. But most are in the dataset so it's not very meaningful

Comment by yomismoaqui 5 days ago

Now that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI?

Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)

Comment by wise_blood 4 days ago

repeating myself, but:

once 3 is solved, we would come up with 4. then 5, 6...

it will be AGI when we cannot come up with a task easy for human but hard for machines. thet's the whole point.

Comment by 4 days ago

Comment by john_alan 4 days ago

exactly, this isn't AGI.

Comment by hypfer 5 days ago

What are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?

Comment by petu 5 days ago

OpenAI provides API key with ~unlimited use?

Comment by Frost1x 5 days ago

So, you’re telling me I need to start a benchmark as a side gig to get a bunch of free compute.

Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.

Alignment++

Comment by manquer 4 days ago

You missed the hard part getting on HN front page , ie. Getting the acceptance of the community / zeitgeist .

There is no incentive for OpenAI to subsidize is you if no one reads /reports on your benchmark . They are only going to fund a few that are currently popular .

Community acceptance doesn’t automatically mean the best , it is combination of some level of technical quality and the ability of the promoter to socially influence or get support of influencers .

Comment by bigbuppo 5 days ago

Wake me when it's going to spontaneously fix my leaky faucet because if it doesn't do it nobody else will. Until it has that capability I don't really care.

Comment by yusufozkan 5 days ago

what the hell is that score/cost curve lol

Comment by minimaxir 5 days ago

DeepSeek v4 Flash recently had a similar "more reasoning is cheaper" curve. It's a fun counterintuition.

Comment by Frost1x 5 days ago

It’s not that different than a lot of real world economies. Often paying for someone or something with better quality can reduce total costs. You have less failures, less mistakes, so on, so while the expertise or quality of the product is higher than cheaper solutions, they can be more reliable and over time ultimately cheaper.

The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).

Comment by Phemist 5 days ago

What is the intuition. Higher quality turns due to more reasoning results in significantly fewer turns taken?

Comment by tedsanders 5 days ago

Yep. In particular, ARC-AGI-3 is a series of games where if you fail, you keep trying again (until eventually hitting a timeout). So the sooner you succeed, the sooner you stop spending tokens retrying. If it was a benchmark where everyone got one attempt with no retries, you wouldn't see it bend backward.

Comment by minimaxir 5 days ago

Yes, in theory.

Comment by fxd 5 days ago

“AGI” never made sense to me. It’s a purely marketing term right?

I’ve ignored it thinking it would go away, but it keeps coming up.

I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery.

Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get it.

But you’d be no nearer to solving consciousness.

Given this thought trajectory - what is AGI supposed to be?

Comment by layer8 4 days ago

Not sure why you are bringing up consciousness, that’s largely orthogonal to intelligence. AGI is usually taken to mean the capability to match or surpass human intelligence across all conceivable cognitive tasks, as opposed to being limited to certain kinds of tasks, or to not matching the general level of human intelligence in some respect.

Intelligence, and hence AGI, doesn’t require consciousness or emotions or sentience.

Comment by p1esk 4 days ago

capability to match or surpass human intelligence across all conceivable cognitive tasks

What human intelligence do you mean? Genius? Professional? Educated? Random person? “Dumb” person?

Comment by HarHarVeryFunny 1 day ago

Intelligence is a capability, not a level of knowledge, and in discussing AI vs human intelligence, or human vs dog, it is perhaps better to regard it as a species-specific capability, not an individual-specific one.

Was the genius or dumb person really born with different levels of learning ability, or were they just raised differently: nature vs nurture ?

The "intelligence" of one species vs another comes down to differences in cognitive architecture, ultimately reflected in ability to learn and predict/infer. Intelligence, as a capability, not IQ test score, is best regarded as ability to learn from experience and use that learning to accurately predict future outcomes, ranging from passive observation, to the outcome of one's own actions, to the ability to reason.

Comparing the intelligence (not knowledge) of different AI system to humans should therefore be assessed by comparing their ability to learn and use what they have learnt.

An LLM is what it is - a language model, not a learning system. It's really an expert system of sorts, highly capable in terms of what it can infer based on the knowledge it encodes, but with very limited ability to learn.

When you talk of comparing AI to a genius/professional/graduate/etc, you are really talking about comparing level of knowledge, comparing one expert system to another, which is fine and perhaps useful in some contexts, but it is not the same as comparing actual intelligence - learning ability, and especially so if you want to discuss general intelligence which is all about the ability to successfully take on any task (perhaps needing to learn it first), not just do well on some limited set of tasks you are already familiar with.

If ability to learn is limited to one modality, such as language, then that indicates a lack of generality. A good test for whether an AI system is in the same ballpark as a human in learning ability, aka intelligence, would be whether it can (at run-time) learn language itself, from a blank slate start, when running in a suitable environment.

Comment by famouswaffles 1 day ago

>If ability to learn is limited to one modality, such as language, then that indicates a lack of generality. A good test for whether an AI system is in the same ballpark as a human in learning ability, aka intelligence, would be whether it can (at run-time) learn language itself, from a blank slate start, when running in a suitable environment.

So a transformer ? Run-time? Seems like an arbitrary distinction to me. Because humans function in a certain way, every learning system no matter how capable must function in the same way to be 'truly' intelligent ?

Comment by HarHarVeryFunny 22 hours ago

If a system wants to claim human-level intelligence, then yes.

It's not a very high bar - even a rat has the basic ability to learn.

Comment by famouswaffles 21 hours ago

The mechanism is irrelevant if the results are similar. No, it doesn't have to be implemented the same way as a human to qualify as a human-level intelligence. That would be silly.

Comment by HarHarVeryFunny 11 hours ago

How is text-based memorization going to let you learn non-linguistic skills, whether animal/human level or beyond (based on new data senses)?

How is text-based memorization going to substitute for learning? The two are not the same. Perhaps this is more applicable to robotics than a text generator, but I also doubt an LLM could learn text-based skills like programming or math if it had not been pre-trained on them via SGD & RL, and had to instead rely on some poor-man's-learning context-based recall instead. What else can't it learn? How is the LLM intern, the "drop-in replacement remote worker" going to do on day #2?

Instead of pretending that an LLM can be human-level, or super-human, or become generalist, why not just admit that this is not the final form of AI. An LLM is not an animal/human-like intelligence, it is something different - a language model, with it's own strengths and weaknesses.

Despite all the AGI hype, an LLM seems to have more in common with a pre-trained single-purpose system like AlphaGo than a brain, but with the rubric/reward-based policy function baked into the weights.

In another 10-20 years some new idea, hopefully more brain-like, will have superseded LLMs and they will indeed be labelled as "LLMs" as the AI/AGI label becomes attached to the new more brain-like creative intelligence. Perhaps it'll be sooner than 10-20 years, but I doubt it given the current 10-year fixation with LLMs which doesn't appear to be slowing down anytime soon. Perhaps Sutskever is working on something a bit different?

Comment by famouswaffles 7 hours ago

>How is text-based memorization going to let you learn non-linguistic skills

I don't know. How is Astra a step change in computer use and spatial reasoning to the extent it can play games, paint good looking stuff with e.g canva and a whole number of other things ?

>How is text-based memorization going to substitute for learning? The two are not the same.

Of course if you call it something else then you can say it's not the same.

>How is the LLM intern, the "drop-in replacement remote worker" going to do on day #2?

Just fine I imagine ? ICL and the memory tools around a harness are pretty good. I'm not sure what sort of magic you're expecting from the human, but they're not getting any improvement in that time frame a frozen transformer can't match.

>Instead of pretending that an LLM can be human-level, or super-human, or become generalist, why not just admit that this is not the final form of AI.

It doesn't seem like I'm the one pretending here.

>Despite all the AGI hype, an LLM seems to have more in common with a pre-trained single-purpose system like AlphaGo than a brain, but with the rubric/reward-based policy function baked into the weights.

If you say so.

>Perhaps it'll be sooner than 10-20 years, but I doubt it given the current 10-year fixation with LLMs which doesn't appear to be slowing down anytime soon.

The architecture that keeps delivering results isn't slowing down ? You don't say.

There's no shortage of people, even researchers, who for one reason or the other are convinced we are in need of some paradigm shift.

But guess what? Talk is cheap. You beat the current paradigm or you don't.

Comment by HarHarVeryFunny 5 hours ago

> How is Astra a step change in computer use and spatial reasoning to the extent it can play games and paint good looking stuff with e.g canva ?

Presumably because of pre-release training, because some alien outside of the model, armed with the reinforcement learning algorithm, came in and programmed its weights.

> The architecture that keeps delivering results isn't slowing down ? You don't say

Sure, nothing wrong with that, as long as you don't misrepresent the limitations of the approach.

> There's no shortage of people, even researchers, who for one reason or the other are convinced we are in need of some paradigm shift.

> But guess what? Talk is cheap. You beat the current paradigm or you don't. Do you seriously think that Meta never scaled JEPA ?

The idea has not been taken very far, so what is there to scale? It's not a complete cognitive architecture. So far it's also been using a pre-trained transformer as the learning component, which makes it of limited interest.

The animal intelligence approach, even in it's most fledgling form (that you would apparently dismiss), requires a complete agentic architecture, including new learning algorithms and generative behavior, before it can be compared to LLMs. We know that, done right, our brain architecture is more capable than an LLM, so even if any hypothetical attempts to reproduce it were not highly performant, we know that the idea itself is sound. You might compare with Uszkoreit's initial poor-performing implementation of his new language model architecture - should he have given up?

Comment by famouswaffles 4 hours ago

>Presumably because of pre-release training, because some alien outside of the model, armed with the reinforcement learning algorithm, came in and programmed its weights.

They didn't program anything. They gave it data at best.

>The idea has not been taken very far, so what is there to scale?

LeCun was the head of Meta AI for over a decade and his baby that he keeps harping on about wasn't taken very far ? Come on. You're smarter than that. It went the way all the alternate architectures have gone since the transformer, a sidegrade at best, probably not even that.

>You might compare with Uszkoreit's initial poor-performing implementation of his new language model architecture - should he have given up?

What are you talking about? There was no poor performing implementation of transformers that was published that he needed to push through. Are you talking about self attention experiments before the finished transformer? That's literally just research. And if he languished on that for a decade then yeah I'd tell him to probably look at something else, but of course he didn't.

Comment by HarHarVeryFunny 3 hours ago

You regard the LLM post-training process as "giving it data" ?!

Have you looked at all the published JEPA research both while LeCun was at Meta, and since (up to and including the latest AdaJEPA from June)? Please enlighten us as to exactly which line(s) of research you think were "scaled" at Meta, and then tell us which of these constituted anything even remotely resembling a complete testable intelligence?

FYI, it's been a long time since FaceBook/Meta even had a single head of AI. Since 2018 it has been split into two groups, FAIR and Generative AI, with LeCun being in the FAIR group, not as head, but as Chief AI scientist. LeCun only invented JEPA in 2022 (shortly after FaceBook became Meta), first writing about it in his "A Path Towards Autonomous Machine Intelligence" paper, perhaps unhappy with the work of the GenAI group, which he had no control over, that presumably was getting all the compute.

https://openreview.net/pdf?id=BZ5a1r-kVsf

> What are you talking about?

I was referring to Uszkoreit's personal telling (on YouTube) of the origin story of the Transformer, his motivations with the design, his initial personal failure to implement his idea in a performant enough manner to beat the current LSTM SOTA, and Noam Shazeer then throwing the kitchen sink at it and eventually coming up with the Transformer design.

Comment by famouswaffles 1 hour ago

>You regard the LLM post-training process as "giving it data" ?!

Yeah. Presumably, lots of synthetic data is being generated, experiments being run, but post-training is still a largely automated process.

>Have you looked at all the published JEPA research both while LeCun was at Meta, and since (up to and including the latest AdaJEPA from June)? Please enlighten us as to exactly which line(s) of research you think were "scaled" at Meta, and then tell us which of these constituted anything even remotely resembling a complete testable intelligence?

I have. My point isn't that his ideas are trash or that he should stop working on them. My point is it's not "gone very far" because he's taking it as far as he can, which isn't very far. He's not had a lack of influence, resources or will, either from his time at Meta or now with his billion dollar startup. He's had far more of it than most, if anything. That there's not much to show for it so far is not for a lack of trying.

>I was referring to Uszkoreit's personal telling (on YouTube) of the origin story of the Transformer, his motivations with the design, his initial personal failure to implement his idea in a performant enough manner to beat the current LSTM SOTA, and Noam Shazeer then throwing the kitchen sink at it and eventually coming up with the Transformer design.

So it's what I thought. This is just regular research unless an inordinate amount of time was spent on it and that's not the case.

Comment by HarHarVeryFunny 17 minutes ago

> My point is it's not "gone very far" because he's taking it as far as he can, which isn't very far.

I wouldn't really agree - I'm no fan of LeCun, but the problem with JEPA isn't that it's a bad idea, or can't go very far, but just that it's not much of an idea in the first place!

It's no secret that our brain basically works by prediction, and what we're predicting is necessarily the external world as we perceive it though our own senses, aka latent representations, aka JEPA.

So, you COULD take this unoriginal smidgen of an idea and built it out to a full model of a human/animal brain, whether or not it's LeCun's intention to do so (he seems more interested in just the representational / world model aspect to it), but he certainly hasn't done so yet, nor created any research manifesto indicating that as his intent.

The fact that JEPA implementations to date are using pre-trained Transformers doesn't seem inherent to the approach - one could, with more effort, still predict latent representations (i.e. sensory feedback) but do so using a new real-time learning algorithm based on prediction failure.

LeCun seems more of an academic / research director than a builder, and I would never have put much stock in him being the one to build an animal brain.

Comment by drdeca 4 days ago

I think the distinction between these options probably doesn’t matter alll that much if the threshold you’re using is either “random person” or higher? Maybe bump it up to “random educated person”?

They are still different concepts of course, but I imagine that once one is achieved, the others aren’t far off.

Comment by p1esk 4 days ago

Ok, but a random educated person will not perform well on vast majority of specialized tasks where professionals operate. A model like Astra probably will beat random educated person performance on majority of specialized tasks. It’s getting close to the level of professionals in many domains, and to genius level on some (e.g. math).

I’m just trying to understand the implications of the current frontier model capabilities.

Comment by drdeca 3 days ago

Current models can do a variety of tasks that requires substantial expertise for people to do, yes.

But there are also many cognitive tasks that the typical educated person would do better at than these models.

Like, e.g. long-term managing what a vending machine gets stocked with and what prices the items should be sold for. Or, various things like that.

When there are essentially no more tasks like that, then we’ve reached AGI.

Comment by p1esk 3 days ago

Interesting, you think a random educated person would beat Astra on Vending-Bench 2? Honestly I would not make that bet, I think it would be 50/50.

Comment by layer8 4 days ago

“General”.

Comment by p1esk 4 days ago

Sorry, I don’t know what this means when applied to intelligence

Comment by fxd 4 days ago

[dead]

Comment by meander_water 4 days ago

The OpenAI charter defines it as:

"highly autonomous systems that outperform humans at most economically valuable work"

https://time.com/article/2026/08/26/openai-sam-altman-interv...

Comment by fxd 4 days ago

So vague it’s useless.

We cannot in a declarative sense define what is economically valuable work even now let alone into the future.

People take what they can get for pay. Very few individuals can demand a wage. The value of employment is obviously designed around that, not some arbitrary definition of “valuable”.

Of course an AI will accept $0/hr, it doesn’t mean it does the job.

Anyone who could accurately define the value of work would be wildly successful without having to try.

That is not a useful definition for me unfortunately.

Comment by jryle70 4 days ago

That's the challenge of defining intelligence. What is intelligence?

Do you want to take a stab at defining/quantifying it? I'm s afraid anything specific you can come up with will also be useless.

Comment by fxd 4 days ago

No it isn’t. Nobody said you had to equate “intelligence” with “value of work” and you got it wrong:

I said it’s impossible to define that value, not intelligence. I’ve seen a dozen or so useful definitions of intelligence.

Comment by benlivengood 4 days ago

AGI is the term invented because arguments about what AI meant had gotten annoying. Originally there was no distinction and people thought "AI" would mean human level intelligence. Chess and conversations and robotics and math all in one package. Then games and classification and some robotics got solved and called AI, but that didn't solve math or conversation or online learning or a host of other things, so AGI was coined to refer to most of the whole package, virtually all the capabilities you'd need to replace humans intellectually. Now we're quibbling about whether AGI includes robotics or full real-world physical agents or something less.

The consensus now seems to be that once you've got human-level intelligence and planning and executive function then you get recursive self-improvement that can eventually autonomously solve the robotics and world-modeling and other portions of human-equivalence.

Comment by eagerpace 5 days ago

I like recursive self improvement instead. It seems like something that is actually quantifiable and kinda “the point” of why consciousness is important to humans.

Comment by fxd 4 days ago

So basically, being able to set it free on some long running goal and it sort of “lives” and autonomously does its own tasks?

I wonder at what point consciousness is necessary… that is, if you can have anything like that without it.

To the point that solving consciousness (and combining it with intelligence) is what gives you the autonomous, recursive, self-improving thing otherwise it can only drive in the dark and make big mistakes.

To your point I think - it’s why we don’t see too many non-conscious advanced biology (it rarely survives against those with it).

Comment by mdp2021 4 days ago

> Knowledge and thus intelligence

How can you conflate the two.

> solving consciousness

We are very much not interested in that. We just need a proper problem solver.

Comment by submain 4 days ago

Unproven, but it could very well be some problems require consciousness to be solved.

Comment by mdp2021 4 days ago

Extraordinary claims require the effort of the proponent so that they can be accepted on the table and take some proportionate share of it.

Comment by fxd 4 days ago

Right, you are interested in “AGI” and presuming none of that requires consciousness right?

For example, how do you know that “feeling pain” is not a functional prerequisite for a task. And that consciousness is a prerequisite for feeling pain

Comment by mdp2021 4 days ago

Because there is no feeling of pain involved during the reasoning about "How to improve the balance of power in the Indo-Pacific" for the human reasoner, hence there is no need to have any experience of pain for any reasoner.

Comment by grantcas 4 days ago

[dead]

Comment by StopTheMods 4 days ago

[dead]

Comment by ajjahs 4 days ago

[dead]