Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
Posted by aarondong 6 hours ago
Comments
Comment by andy99 2 hours ago
Comment by afavour 2 hours ago
Comment by areoform 25 minutes ago
– "Does collagen supplementation empirically work?"
- "Can you help me figure out how to calculate and generate Kaplan-Meier curve?"
– "Why do rabbits reproduce so frequently?"
— "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me this works at the biomolecular level?"Comment by alightsoul 11 minutes ago
Comment by alain_gilbert 1 hour ago
Fable understood it as something along the lines of:
"introducing" "security risk" "using software" to "unlock door" YOU ARE FLAGGED
Comment by JumpCrisscross 12 minutes ago
The dumbfuck bouncer Anthropic put in front of Fable decided this.
Fable is a PR model. It’s great. But if it were an employee, it would be the one who randomly shows up to work high.
Comment by gck1 1 hour ago
It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.
Comment by d-m 38 minutes ago
Comment by JumpCrisscross 13 minutes ago
The current state of guardrails seems to be entirely about marketing to investors at the cost of customers. I’m switching to open models when my subscription expires.
Comment by wild_egg 2 hours ago
Comment by Retr0id 2 hours ago
Comment by msp26 2 hours ago
Comment by skinfaxi 1 hour ago
> Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats.
> Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.
Comment by cge 41 minutes ago
Being a researcher somewhat connected to chemistry and biology, Fable has been the most useless model I have ever tried. Essentially all work has instantly downgraded to Opus.
Comment by stavros 59 minutes ago
Comment by estearum 48 minutes ago
Comment by areoform 29 minutes ago
> while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite
I've heard this sentiment repeated elsewhere, but why? What makes you think that's the case?Under this rationale, every serious HS textbook has "approximately infinite" risk. That's clearly not so. Why is this any special?
Comment by estearum 23 minutes ago
Comment by Mistletoe 1 hour ago
Comment by TurdF3rguson 1 hour ago
Comment by rzk 25 minutes ago
Comment by wewtyflakes 2 hours ago
Comment by fluidcruft 1 hour ago
Comment by patcon 1 hour ago
Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me
Comment by icedrift 2 hours ago
Comment by jefftk 2 hours ago
Comment by eterm 1 hour ago
Either that or everyone is indeed talking across each other and talking about different things.
Comment by theplumber 50 minutes ago
Comment by dylanowen 59 minutes ago
Comment by weird-eye-issue 1 hour ago
Comment by AnotherGoodName 1 hour ago
Comment by Levitz 1 hour ago
Comment by jbritton 1 hour ago
Comment by arcanemachiner 1 hour ago
I've been saying this a lot lately, but it doesn't bites you until it bites you.
The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.
Comment by thousand_nights 2 hours ago
Comment by Tostino 1 hour ago
Comment by cute_boi 1 hour ago
Comment by idiotsecant 38 minutes ago
Comment by gck1 1 hour ago
Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?
Comment by kccqzy 1 hour ago
Oh but then you said you never pay Anthropic so you haven’t actually used Claude Code yet. Why would anyone listen to the opinion of a non-user?
Comment by gck1 53 minutes ago
Where do you think the principle came from? I've used claude code for a year, and stopped February this year.
Comment by markasoftware 1 hour ago
Comment by buzzerbetrayed 2 hours ago
Comment by pinkyboy 2 hours ago
Comment by didibus 1 hour ago
The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).
Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.
That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
Comment by theplumber 16 minutes ago
Comment by chmod775 2 hours ago
At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
Comment by ricardobeat 2 hours ago
Comment by Bolwin 5 minutes ago
Comment by andriy_koval 12 minutes ago
Comment by stingraycharles 56 minutes ago
Comment by firasd 2 hours ago
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)
I've thought for a while that Gemini 3.x has 'big model smell'
Comment by chronogram 1 hour ago
I think AA-Omniscience Accuracy follows your expectations better. An ultra size Fable at 61%, followed by large frontier models like Sol, 5.5 and Opus. With Flash being up there. I assume because Gemini is more focused on general knowledge to operational cost in particular, rather than getting the highest scores in coding benchmarks. If you go to Domain Score (Normalized) you'll see that the Gemini models are only less competitive in Software. And that's where Sol goes from 6 in Health to 71 in Software.
Comment by mchusma 1 hour ago
Comment by andriy_koval 12 seconds ago
Comment by aarondong 6 hours ago
Comment by eli 3 hours ago
Comment by emmp 2 hours ago
Comment by midnightbobarun 5 hours ago
Comment by nijave 3 hours ago
Like 96% vs 93% or something
Comment by impulser_ 3 hours ago
Comment by scrlk 2 hours ago
No wonder why Tibo can afford to hit the reset button liberally.
Comment by charcircuit 2 hours ago
Comment by vikramkr 52 minutes ago
Comment by charcircuit 23 minutes ago
Comment by wmf 1 hour ago
Comment by charcircuit 27 minutes ago
Comment by giancarlostoro 3 hours ago
Comment by brookst 3 hours ago
Comment by brcmthrowaway 3 hours ago
Comment by wmf 3 hours ago
Comment by Schiendelman 4 hours ago
Comment by anuramat 2 hours ago
Comment by kristopolous 40 minutes ago
https://github.com/day50-dev/aa-eval-email
This also works
$ curl day50.dev/art-analysis.sh | bash
Artificial analysis knows about my tool and I'm working with them on getting their API improved.
Comment by nu11ptr 1 hour ago
Comment by zuzululu 3 minutes ago
its surprisingly bad at UI which is unexpected
its also lacking in depth vs sol 5.6 which goes above and beyond (which in itself is also an issue at times)
Comment by zormino 2 hours ago
Comment by hoppp 2 hours ago
Comment by vehemenz 2 hours ago
Comment by theplumber 1 hour ago
Comment by sggyamg 2 hours ago
Comment by LeBit 2 hours ago
"Not fair! They distilled Opus 5!"
Comment by brikym 9 minutes ago
Comment by claude-ai 5 hours ago
Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").
Comment by reilly3000 3 hours ago
Comment by pixelesque 2 hours ago
I bizarrely had Opus 4.8 this week (in pi.dev within a podman container, using openrouter) start installing various python packages (and uv!) within the environment (not as root) when I asked it to code review some fairly basic Rust .rs files that were generally stand-alone (it did very nicely work out and write some stubs for them to build them and work out how they worked).
It only gave up with the weird Python installing stuff when it discovered one of the Python packages needed Tensorflow.
It seems pretty focused and persistent in continuing its initial approach, and I'm wondering if I need to alter some instructions / initial prompts to rein it in a bit...