Artificial Analysis Intelligence Index v4.2
Posted by nojs 4 days ago
Comments
Comment by jascha_eng 4 days ago
https://artificialanalysis.ai/evaluations/omniscience
> measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. But this index measures how often it is correct while penalizing wrong responses so that a high score means you can trust this model more and when it doesn't know it is more likely to tell you that it really doesn't know rather than making shit up.
Fable also performs a lot better than opus 5 here which correlates very strongly with perceived strength despite the models performing similarly on e.g. DeepSWE
Astra is a big jump from sol and performs the same or slightly better than fable here.
Comment by anon373839 3 days ago
IMO, what really makes a model useful is its ability to process information within its context reliably and faithfully. I don’t care if it hallucinates George Washington’s favorite color, but I do care about it hallucinating the results of tool calls.
Comment by jascha_eng 3 days ago
Comment by gizmodo59 4 days ago
Comment by testycool 3 days ago
If you customize the filter to add Gemini 3.1 Pro you'll see it ranks 5th in AA-Omniscince Index. Yet by default you wont see Gemini 3.1 Pro.
I find this to be very unhelpful and confusing.
Comment by jascha_eng 3 days ago
Comment by jesuslop 4 days ago
Comment by yorwba 3 days ago
Comment by jascha_eng 3 days ago
Comment by re-thc 3 days ago
It's not useful because the benchmarks often measure the wrong thing. They're here yapping about AGI and yet the benchmarks treat it like a trained dog. Fetch this. 100 points.
Each "problem" in these benchmarks likely has more than 1 solution that can be considered correct and even should be graded in many ways. Yet we see in many benchmarks higher effort (or thinking levels) don't help because the benchmark penalizes for doing "more" than what the answers asks for. So what did you ask for?
In human school you often get marks on the process and not just the end result. Thinking tokens have been cut. All we group on is things like cost, turns and time but not the what else.
Comment by Buoylog 4 days ago
Comment by redox99 4 days ago
The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.
Comment by paimapi 4 days ago
like they have words that are dressed in scientific language on their site like
"We estimate a 95% confidence interval for Artificial Analysis Intelligence Index of less than ±1% - based on experiments with >10 repeats on certain models for all evaluation datasets included in Artificial Analysis Intelligence Index v4.2."
but where's the outcome dataset justifying this? how did they get that probability? what was the specific methodology of the tests? what variables did they account for?
there's a major difference between scientific sounding and being truly empirically rigorous. the 'research' in AI intelligence feels somehow even less trustworthy than supplement-funded studies because those are at least subjected to scrutiny by peers without profit motives
Comment by Wheen 3 days ago
It's just high school statistics: https://en.wikipedia.org/wiki/Confidence_interval
It's not a probability. It's essentially an assertion that if the test were run 100 times, the result would be within the interval 95 times.
Comment by elvin_d 3 days ago
Comment by tancop 3 days ago
Comment by paimapi 3 days ago
2) there are definitely problems with journals and publishing but, like many similar absurdly reductive pronouncements, the argument that the current mode of scientific inquiry is bunk is both wrong and lacks nuance
the problem with modern publishing is that private equity is buying up publishers [0]. these publishers are then giving peer-reviewers no time and zero pay to do the necessary work of review [1] while also charging exorbitant rates for access. this is leading to worsening quality of the published research along with highly overburdened researchers who are stuck between shrinking funding [2] and their myriad other professional obligations
to just say that 'journals' are bunk is ignorant in a harmfully anti-empirical way. the process of empirical research and review is the entire reason why we see realworld results. foundations comprised of bullshit crumble fast but for some reason or another our economic system is highly driven to enshittifying everything it touches
to have defenders who claim AA is 'pushing the boundaries of the scientific method' sounds like the screeching refrain of anti-intellectual cargo cults, apeishly mimicking the features of rigor and methodology while avoiding any real accountability
[0] https://issues.org/how-academic-science-gave-its-soul-to-the...
[1] https://www.insidehighered.com/news/faculty/books-publishing...
Comment by elvin_d 3 days ago
Arxiv is dominated by CS [1] and nothing really changed over the years [2]. Saying about pushing the boundaries of scientific methods it was another rock into journals and the status quo on how it's done like p-hacking and other data manipulations. Being in open with community notes would would have a difference.
Empirical research can exist without journals, science can be for masses and journals are antithetical for sharing making the knowledge exclusive.
[1] https://info.arxiv.org/about/reports/submission_category_by_... [2] https://en.wikipedia.org/wiki/ArXiv?useskin=vector#/media/Fi...
Comment by x312 3 days ago
Completely discredits the index if it just gets modified to match social media vibes.
Comment by kingstnap 4 days ago
Theoretically it's unscientific to do tweaks like this but in reality this is what actual science is because you need to see the results understand the problems in your experimental designs.
Now the interesting thing is that you could repair the bias problem. The main issue is that you will do these tweaks after a bad result not after they look fine.
Maybe we could commit ahead of time that the experiment will be reanalyzed after results regardless of what they are. Instead of only when the results disprove the hypothesis.
So like artificial analysis committing to a fixed cadence of index updates instead of when the results start looking jank.
Comment by baq 3 days ago
This needs to be evaluated per task, jagged frontier yadda yadda. I would be not at all surprised if sol was better at some things than Astra just like people still use opus 4.6 and for good reasons.
Comment by cma 3 days ago
For knowledge type questions is that necessarily true? In the past we've seen things like Google's models degrading on general knowledge after the preview releases while improving on code/tool use, presumably due to catastrophic forgetting from the additional training. Their preview would be free, get lots of agentic use from users, then additional training on that and probably additional automated RL.
However Astra is on an entirely new base model so I also wouldn't expect it to be worse.
Comment by andriy_koval 3 days ago
I asked Astra to make some changes, it built crazy overengineered code, switched back to Sol, said code is too complicated, rewrite it, and it made lot cleaner code.
I feel some models could chaise complicated benchmarks too much in expense of simpler tasks quality.
Comment by cbg0 3 days ago
Comment by big-chungus4 3 days ago
Comment by weird-eye-issue 3 days ago
Comment by villish 3 days ago
Comment by weird-eye-issue 3 days ago
They are literally frequently discussed and it's why there are different benchmarks for different domains.
Comment by villish 3 days ago
Labs know the only thing the public even discusses on model releases are benchmarks, so they devote a majority of training on just benchmaxxing. It’s marketing.
Muse Spark looks great in benchmarks. Everyone I know who has tried it (Rust & C++ projects) has determined it’s a resounding “meh”. That doesn’t mean it’s not a great tool for frontend devs, I wouldn’t know.
Comment by throw10920 3 days ago
Comment by throwaway13337 3 days ago
A glance at their new index shows that whatever they're measuring, it isn't useful.
Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad.
I haven't tried muse spark 1.3. But it must have been a miracle since 1.2 to hit that rank.
Video game journalism vibes all over this.
Comment by nojs 3 days ago
Comment by WASDx 3 days ago
Comment by dist-epoch 3 days ago
Comment by Catloafdev 3 days ago
It's really easy to shit on AI benchmarks, but that noise is useless unless you're offering a solution or a better benchmark.
Comment by jjcm 3 days ago
Comment by __jl__ 4 days ago
Many labs used increased thinking to boost benchmark scores and performance. Most of the Chinese models were doing that for a while. Google and Anthropic as well.
Not OpenAI. 5.6 already was much more token efficient than other models and Astra beats Sol in token efficiency by a wide margin.
Edit: Just to make the point: Astra (max) has the 2nd highest score and the third lowest output tokens (among the models shown by AA).
Comment by jsnell 3 days ago
I don't really understand why they keep doing this. Either run and report everyone at multiple effort levels, or run everyone at only one.
But alsi, token efficiency seems pretty artificial? For example tokenizers are different from model to model. The cost/perf Pareto frontier seems a lot more meaningful (and Astra does very well at that too, just to be clear. It seems to be a great model.)
Comment by Scaevolus 4 days ago
This also makes it much harder to monitor its reasoning.
Comment by ssivark 3 days ago
Comment by water-drummer 3 days ago
Comment by ssivark 3 days ago
But if we're actually talking benchmarking intelligence -- as evidenced by "Artificial Analysis Intelligence Index" and the rest of the reporting -- then comparing systems with different amounts of recurrence on output tokens is flawed.
Comment by __natty__ 3 days ago
Comment by dist-epoch 3 days ago
The only thing they are selling and why people look at them is trust that they do honest evaluations.
Comment by marmarama 2 days ago
These kinds of "research" companies are there to validate people's preconceived ideas and purchasing decisions rather than being genuinely unbiased. And they are very useful for that, both for consumers and for marketers.
Comment by AnodicElegy 4 days ago
Comment by CuriouslyC 3 days ago
Comment by pixl97 3 days ago
Comment by stared 3 days ago
Compare and contrast with ARC-AGI, BabaIsBench (https://quesma.com/benchmarks/babaisbench/), or MazeBench (https://mazebench.com/blog?post=introducing-mazebench).
In particular, in one Baba Is Bench post (https://quesma.com/blog/baba-is-aug-2026/), while quoting a Pareto frontier chart from AA, I noted:
> Intelligence Index vs. Cost per Intelligence Index Task from Artificial Analysis. Note that it is based on score of benchmarks like Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond - not necessarily fluid intelligence like in abstract puzzle games of ARC-AGI-3 or Baba is You.
Comment by sanxiyn 3 days ago
Comment by dgacmu 3 days ago
(2.1 is useful for knowing they can all one-shot straightforward scripts, of course.)
Comment by swingboy 3 days ago
Comment by dgacmu 3 days ago
It's still providing strong discrimination between weaker models.
Comment by theagenticleade 3 days ago
Comment by sheepscreek 2 days ago
Comment by aurareturn 3 days ago
Comment by mmmmbbbhb 2 days ago
Comment by nthypes 4 days ago
Comment by AnodicElegy 4 days ago
Comment by lousken 4 days ago
Comment by redox99 3 days ago
Comment by 6thbit 4 days ago
Clearly that would move things around.
Comment by thereitgoes456 3 days ago
Comment by dannyw 3 days ago
ARC-AGI takes it very seriously: they've never tested Fable, because they won't run on the eval set without ZDR.
Another possible way to look at this is that any benchmark reporting Fable scores is potentially contaminated.
Comment by theycallmeritik 3 days ago
Comment by MoreThanMe 3 days ago
Comment by snezhadianpm 3 days ago
Comment by ahmedelsama 3 days ago