Claude Opus 5
Posted by alvis 8 hours ago
Comments
Comment by postalcoder 8 hours ago
> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].
0: https://support.claude.com/en/articles/15425996-data-retenti...
Comment by bluegatty 1 hour ago
Comment by alvis 8 hours ago
Comment by x313 7 hours ago
Comment by SwellJoe 6 hours ago
Comment by InsideOutSanta 5 hours ago
Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.
Comment by SwellJoe 5 hours ago
But, I'm finding Kimi K3 terrifyingly expensive in the way that Fable and GPT 5.5 Pro are at token rates. Not as expensive as those, but expensive enough to where if you don't put a budget cap on it, you might wake up bankrupt if you leave a task running overnight. Not because of the per-token cost, but because how many tokens it's going to burn.
Comment by nullify88 4 hours ago
Comment by SwellJoe 4 hours ago
Then, I added it to my benchmark of security vulnerability auditing capability, and it burned a bazillion tokens, burned through the 5-hour limit, burned through $100 in extra usage I'd allocated, and was only 11% finished. That's more expensive than any model I've tested other than GPT 5.5 Pro on this task.
These are things I've done with a bunch of other models, I feel like I have a notion of what they ought to cost, and with K3, they end up being crazy expensive. (And it seems to be a function of how many tokens it burns accomplishing the tasks.)
Comment by lukan 2 hours ago
Those who pay for the expensive direct API, get served first.
Comment by Barbing 1 hour ago
And not convinced they couldn’t have instead tried the It’s A Wonderful Life strategy (“fam we’re oversold, would some of y’all be OK to limit your usage? We’ll get you back one day!”)
Comment by try-working 5 hours ago
When K2.7 was released, they cut quota by 80%. I can't tell how much they have further cut it after the K3 release because it's barely worth using at all. I just use it in my model router since I have the annual plan paid for.
It's just not a serious model or company.
Comment by InsideOutSanta 4 hours ago
Comment by bg24 3 hours ago
Subscription is to drive adoption - fixed cost, can adjust the usage eg. give resets, increase quota based on capacity available. We subscribers tend to take it as a mandatory benefit :-) For labs, it is not letting the capacity go waste.
api is the $$ driver - pay per use, enterprises.
Right now, Kimi needs to first hit the subscribers at the level of OpenAI and Anthropic. With the api usage skyrocketing due to K3, it will be clear in a few months on the actual subscription benefits.
Comment by KronisLV 3 hours ago
For me, the Moonshot 100$ plan felt like it gives me lower total amount of work I can do than the Anthropic 100$ plan (probably within like 30% of each other). Kimi has way more generous 5 hour limits (never hit those once, whereas I do regularly with Opus) but the 7-day and monthly ones are lower. However, with the annual billing, Moonshot's 200$ tier plan becomes way better, because you get it for 159 USD per month.
There's also the odd thing of Anthropic's 100$ plan charging me 108 EUR so seems like their sticker price does not include VAT but Kimi's did, cause I paid like 87 EUR. Wrote down some initial thoughts at https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don... but it's hard to do exact comparisons (even the same task will have way different real token amounts per model).
Still, Kimi K3 is a pretty cool model! On high reasoning, it was pretty close to Opus 4.8 and didn't seem to waste as many tokens as Max.
Comment by adgjlsfhk1 5 hours ago
Comment by sggyamg 2 hours ago
Comment by adgjlsfhk1 1 hour ago
Comment by crimist 15 minutes ago
Comment by reinitctxoffset 7 hours ago
Comment by idiotsecant 6 hours ago
Comment by throwawayoaky 3 hours ago
Comment by reinitctxoffset 5 hours ago
Comment by onlyrealcuzzo 6 hours ago
It basically shows that Sol absolutely demolishes Fable at every part of the cost curve for coding for the same level of quality.
Opus is competitive. It just has a higher level of quality / higher cost to start.
Comment by throw10920 1 hour ago
Comment by pixl97 6 hours ago
Comment by qsera 8 hours ago
Comment by iambateman 7 hours ago
Stop using Opus immediately if you experience signs of dizziness or vomiting.
Opus 5…the people’s favorite.
Comment by manojlds 6 hours ago
Comment by benjiro29 5 hours ago
https://www.vals.ai/benchmarks/vals_index
!!! Vals !!!
Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. Sure, 20% cheaper then Fable, but that is a 3x price increase compared to Opus 4.8 in that test.
https://artificialanalysis.ai/models/claude-opus-5 https://artificialanalysis.ai/models/claude-opus-5#price-cos...
!!! artificial analysis !!
Cost per task is second highest, right below Fable.
* Fable: $2.75
* Opus 5.0: $2.03
* Opus 4.8: $1.80
* GPT 5.6 Sol: $1.04
* Kimi K3: $0.95
Looks like interest levels of cherry picked cost in their report. Cheaper model, clearly NOT. More expensive in both benchmarks.
Comment by spider-mario 4 hours ago
That the most expensive variant is expensive doesn’t really tell us much.
Comment by benjiro29 3 hours ago
If you start to drop effort levels, you need to compare to the competition models. So GPT models on the same ~intelligence level, are then 50% cheaper.
You see the issue? Its still a expensive model, and from my understanding, it still uses the old tokenizer.
Going to be interesting to see when GPT 6 comes out (very soon).
Comment by artursapek 5 hours ago
Comment by abixb 8 hours ago
Also glad they still kepy Fable 5 on "credits only" access. I think we're going to start seeing model providers gate top-of-the-line models behind pay-as-you-go API rates/credits while subsidizing other models on monthly subscriptions.
Comment by jpk2f2 7 hours ago
Comment by saratogacx 6 hours ago
I burned through $45 in 3 prompts to fix some bugs in my code (Some kind of tricky to isolate). That thing burns through cash so fast I don't see myself using it outside of maybe building execution plans for other systems
Comment by tackta 4 hours ago
I have moved on from Fable anyway so just going to view this next 6 weeks as I have a massive amount of Opus 5 to use.
I had a hard time finding anything that would let Fable express its increased intelligence. The few conversations I had this afternoon with Opus 5 were pretty impressed.
If Opus stays one click back from the frontier model, I will remain a happy customer.
Comment by mcv 6 hours ago
Comment by ciefa 6 hours ago
Comment by Wowfunhappy 8 hours ago
Comment by collabs 7 hours ago
Comment by ValentineC 7 hours ago
Comment by bakies 7 hours ago
Comment by einsteinx2 7 hours ago
Comment by ValentineC 5 hours ago
Comment by eterm 7 hours ago
Comment by d4rkp4ttern 7 hours ago
Comment by collabs 7 hours ago
Comment by abratabia 7 hours ago
Comment by krzyk 5 hours ago
Comment by kodablah 5 hours ago
> Opus 5 can silently fallback to Opus 4.8 (without any notice) on the serverside if you hit a guardrail
But https://support.claude.com/en/articles/16049681-why-claude-s... says (emphasis mine):
> These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered.
So who is right? I know for Fable I am visibly told, is this tweet trying to say it is silent against what Anthropic is saying?
Comment by SwellJoe 4 hours ago
Comment by pseudohadamard 36 minutes ago
There also seems to be some cross-pollination across models, going Fable, Fable, Fable, guardrail, Opus 4.8, Opus 4.8, ... gives more Fable-like results from Opus than just Opus 4.8, Opus 4.8, Opus 4.8, ...
Comment by solenoid0937 4 hours ago
Comment by lossolo 3 hours ago
Comment by trueno 2 hours ago
Comment by pseudohadamard 32 minutes ago
Comment by gonzalohm 7 hours ago
Comment by NiloCK 6 hours ago
It's a funny design/affordance. I do see them often writing memories of things that that feel unlikely to be important going foward / with other tasks, but I don't see them clearly getting tripped up by them as prior models used to. (eg: Since you're running Ubuntu in Canada, here are some drills you can try to help your kid hit a baseball more consistently.)
Comment by persedes 7 hours ago
Comment by mh- 7 hours ago
In my enterprise-seated account I see slightly different options available (vs. my personal account) in the Capabilities section:
Search and reference chats
Allow Claude to search for relevant details in past chats.
Generate memory from chat history (Legacy)
Allow Claude to remember relevant context from your chats. Memory includes your entire chat history with Claude.
The first option was defaulted to on, if I recall.Comment by gonzalohm 5 hours ago
Comment by bathtub365 7 hours ago
Comment by gonzalohm 5 hours ago
Comment by usef- 2 hours ago
When people talk about retention they mean API usage and terminal agents, which run on your device.
Comment by manojlds 6 hours ago
Comment by fny 2 hours ago
Comment by cjonas 2 hours ago
From the docs[0]:
> To use this model, you must opt in to provider data sharing by setting your data retention mode to provider_data_share via the Data Retention API
0: https://docs.aws.amazon.com/bedrock/latest/userguide/model-c...
Comment by arrowleaf 8 hours ago
I hope we get clarification on this, I can't find anything claiming that it is compatible with ZDR.
Comment by collinrapp 7 hours ago
> Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
Comment by solenoid0937 7 hours ago
Comment by doctorpangloss 7 hours ago
Comment by trueno 2 hours ago
Comment by gigatexal 7 hours ago
" Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8)"
Comment by RazorBucksICO 6 hours ago
At the end of the day, they have established a strong brand and if they can get away with a 95%+ gross margin on inference entirely from the status premium, then I suppose that’s good for them. Apple does the same thing, and I don’t fault them for it.
Comment by gigatexal 4 hours ago
Comment by oblio 6 hours ago
Comment by khuzaimsharif 4 hours ago
Comment by hnscum 8 hours ago
Comment by jjcm 7 hours ago
Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models).
Opus' results seem to be more accurate than Fable, following the design source of truth better.
Example results:
Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we...
Opus 5 build: https://html.non.io/solaraOpus/
Fable 5 build: https://html.non.io/solara/
Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).
Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
Comment by jjcm 6 hours ago
One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration.
Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we...
Opus 5 build: https://html.non.io/neonRamen/
Thoughts: It does a really, REALLY good job at these angular cuts / microglyphs. The responsiveness is off, but I'm very impressed at how well it did here. One way I think of it is "how close to a finished product did this get me?". Opus gets you like 90% there.
Comment by winwang 5 hours ago
Personally, I think being able to have these design languages be easily prototypable is fucking awesome. Great tests! (But a tad low-performance/janky, somehow). Though, I also like the cyberpunk aesthetic. Very on-brand(?) that AI generates it, hah.
Comment by razster 1 hour ago
Comment by echelon 6 hours ago
I love this so much.
Designs like this would never have seen the light of day in the cellphone incrementalism / corporate memphis era of tech. Now people can be weird and awesome again.
This is 1980's cyberpunk / late-90's Matrix / early-00's sci-fi UI. Great ideas that died to frutiger aero (which isn't a bad design aesthetic) and flat design (which is).
This is fun and it's got great colors and I love it.
It's so refreshing to see this.
AI rules. This is the best timeline.
Comment by andersonpico 4 hours ago
Comment by ricardobeat 2 minutes ago
Comment by brailsafe 3 hours ago
The only recent novel addition—I'd speculate—is the specific influence of Cyberpunk the game with its shiny surfaces and pink highlights, but even then it's hardly new.
Comment by GPerson 5 hours ago
Comment by jjcm 4 hours ago
The ramen shop website above is pretty, but it's a veneer. It's not weird and awesome, it's just a representation of a site. I spent about... 7 minutes of my life making it. It's a tech demo, nothing more.
If someone actually poured their heart and soul into a vision for a cyberpunk themed ramen cart, and happened to use this because they didn't have the capabilities or funds to do a proper design, suddenly it becomes less of a veneer, and more just a component in the wider vision of that individual. Their human hours poured into the wider thing that's the business becomes what matters.
Ideally what AI does is it amplifies the hours we do pour into things that are weird and awesome, it doesn't replace them.
Comment by ultimafan 4 hours ago
We do things to achieve some end result but it's the journey there that is the most cathartic to me. The "skilled crafts" element of development where careful deliberation and hours of tinkering to get any kind of appreciable output you can admire has been replaced with a one stop dopamine button that skips the whole process that I could find myself getting lost in.
I've taken up carpentry/metal working as a result. Maybe someday we'll have live in robots that do the same for those hobbies that AI did for programmers but I can't see it happening any time soon.
Comment by hangrybear666 3 hours ago
At my workplace management is pushing AI, so I am using it in order to establish sensible and thoughtful applications of it and in order to know when to call out colleagues for pushing mindless automation out of complacency or blind obedience.
Comment by kami23 46 minutes ago
They were an accessibility nightmare, but you use what you got. I tried so hard as a kid to understand flash, but had to settle on MS Frontpage to publish my first RPG page.
What's old is new again.
Comment by razster 1 hour ago
I'm hella interested in finding out what website builder/diagram app was used. I dig the dark theme/grid.
Comment by jjcm 6 hours ago
This was just from a prompt "A cyberpunk themed ramen food cart website. Should feature menu, locations, and an ability to put in an order for pickup. Simple and clean website with angular cyberpunk microglyphs, pink/teal colors."
Comment by wilkystyle 4 hours ago
I have found myself empowered by AI to tackle all sorts of things that would have too high of a barrier to entry for me to want to spend my limited time on as a busy father who is also working at a small startup.
And when I say that, I do NOT mean that I can crank out a bunch of slop and label it as something I produced even though I don't understand the code. I mean that I can do things like go back to college math that I never appreciated at the time and honestly felt too scared of. I mean having an on-demand math tutor that ask clarifying questions to as I struggle through the problem sets.
I have found that it actually accelerates learning how to code in various problem domains because I can tell it to answer my questions at a conceptual level and be a sounding board, but to never actually write code for me. It can review the code I write and gently nudge me without giving away the answers, so that I still struggle through the learning process and actually gain the knowledge.
And finally, for the first time in like 10 years of feeling overwhelmed and daunted by the prospect of learning game development (I have no background in that), I have found Codex to be an incredible boon for learning with the Godot engine. It helps me understand the terminology so that I know what to search for and what documentation to read. It helps me map my computer science knowledge from other domains into the game world, and to understand why things are structured the way they are. And because Godot saves all of the scenes and geometry and lighting and shaders to the file system as text files, Codex can inspect the results of the work I'm doing in the IDE and help me track down things I'm stuck on, and explain what the issue is. For example, why my pre-baked global illumination lightmap is breaking my ambient lighting configuration.
I know it has never been easier to cheat and skip the hard work that results in actually learning something, but for me, personally, I cannot believe the incredible value that $20 a month has provided me. I have never been more excited and eager to dive into tackling hard things I had previously been afraid of or simply too overwhelmed to attempt.
It has never been easier to quickly prototype and get a feel for some idea you have in your head to see if it even has legs. Simply seeing a quick prototype of an idea is often all of the excitement and fuel I need to then take it and make it a real project.
Comment by hangrybear666 3 hours ago
Comment by aaa_aaa 6 hours ago
Comment by djeastm 2 hours ago
Comment by FranzFerdiNaN 5 hours ago
Comment by bonoboTP 4 hours ago
Comment by aaa_aaa 5 hours ago
Comment by newsy-combi 4 hours ago
Btw, if websites would only include the frontend dev's own hand painted images, we would also revolt at the sight of human slop. It's not just AI.
The whole point is that good artists are capable of producing non-slop, and to this day they're the only group of which this is reliably true.
Comment by tackta 3 hours ago
Water Lilies are enormous paintings. They are breathtaking in person because of their scale. Monet wouldn't be Monet if he had only produced images on a screen.
Art is good or it is shit, based on personal taste. Just like food, no one can tell me what food tastes good or tastes bad.
AI Art seems to produce strong emotions in people who don't go to art galleries. I love modern art, I am a huge art snob but if you want to see slop, go to any modern art gallery. Personally, I would say for my taste, at least 70% of all art at any gallery is basically shit.
Like food also, the presentation matters. To believe there is no possible way to print out a 6 foot tall by 10 foot long AI generated image that would look awesome hanging in a gallery is stupid.
Comment by shimman 1 hour ago
Comment by HaZeust 5 hours ago
Comment by theappsecguy 6 hours ago
But sure, lets cheer that funky website designs are back on the menu…
Comment by Mtinie 6 hours ago
Comment by shimman 1 hour ago
Comment by nextaccountic 1 hour ago
AI has the merit of showing SWE folks exactly where in the class divide they belong. If you are selling your workforce, and you can't maintain your lifestyle if you stop working, you are in the working class
Comment by signatoremo 5 hours ago
In the mean time, I’ll unashamedly continue to cheer for creativity and innovation. Note: I don’t even like this website design.
Comment by kypro 6 hours ago
The "funky" websites of the past were mostly a result of tech immaturity and a lack of profit motive.
Businesses have been able to easily install templates like this for at least a decade. They don't because stuff like this looks cool but isn't very functional.
AI isn't going to make your local restaurant have a funky website, it's just going to make everyone who use to work directly and indirectly for that company unemployable. And even the local restaurant will close down because they can't compete with the multi-national competitor that has automated their kitchen with AI.
Comment by andersonpico 4 hours ago
Comment by kypro 3 hours ago
Can you expand? Specifically what is it about humans that AI and robotics could not replace?
Comment by ai_fry_ur_brain 6 hours ago
Comment by echelon 6 hours ago
Comment by wilkystyle 4 hours ago
Comment by mainmailman 4 hours ago
Comment by chriscamargo 7 hours ago
Out of curiosity, what app is that Design source of truth screenshot from?
Comment by jjcm 6 hours ago
Edit: Generation was down, back up now. Apparently just hit my $1000 cap for the openai api. Upped it to 10k. Growth!
Comment by IanCal 2 hours ago
Comment by jjcm 2 hours ago
At the moment I currently have around $600 of revenue on $1200 spend, but that's primarily because I'm subsidizing new accounts (each new account gets $5 to spend for free, which translates to around ~36 designs). I'm in the process of doing an angel round, so I can afford to operate at a bit of a loss during the growth stage.
Comment by abidlabs 3 hours ago
Inkling (not too great): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Kimi 2.7 (really well, esp. note that this is the predecessor model, not the latest Kimi3): https://cdn-uploads.huggingface.co/production/uploads/608b8b...
Here's how I tested them: https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...
https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...
Comment by jjcm 2 hours ago
Comment by bottlepalm 7 hours ago
Comment by jjcm 6 hours ago
Opus though followed the source of truth better imo. The details are more present.
Fable filled in the gaps for things it wasn't able to do (ie in the design the hero image goes behind the nav), which resulted in a better looking page that was more divergent.
Comment by kccqzy 5 hours ago
It seemed to me that Fable meaningfully improved on the original design more than just faithfully executing the original design.
Comment by andersonpico 4 hours ago
Comment by thatxliner 40 minutes ago
I wonder if there exists a benchmark for that.
Comment by afro88 5 hours ago
Comment by erikw 6 hours ago
Comment by jjcm 6 hours ago
> Create a web page implementation from the following instructions:
> https://diffui.ai/build/Spa_Booking_Experience_build.md?auth...
Comment by hbcondo714 4 hours ago
Thank you for sharing this. I was just using OpenAI's Product Design plugin[1] to create designs but it just didn't reproduce it in code faithfully so will need to try this.
Comment by jjcm 4 hours ago
Comment by abdussamit 2 hours ago
Comment by jjcm 2 hours ago
Comment by Xenograph 6 hours ago
Comment by jjcm 6 hours ago
Comment by ai_fry_ur_brain 6 hours ago
Comment by 01100011 7 minutes ago
Comment by hodgehog11 5 minutes ago
Comment by paxys 8 hours ago
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
Comment by ai-x 7 hours ago
Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability
Comment by johnfn 7 hours ago
Comment by posix_compliant 7 hours ago
Comment by johnfn 2 hours ago
Comment by kolinko 2 hours ago
My experience is the opposite - for many cases it’s not very obvious how good a model needs to be to solve it. Worse models tend to just follow their first instincts without proper reasoning
And also btw you don’t need a routing company to decide, you can do it on your harness. And yeah my Fable has zero issues delegating to Terra instead of Opus.
Comment by maCDzP 6 hours ago
Comment by johnfn 6 hours ago
It would be like asking the clerk at a Whole Foods which grocery store in the city sells the cheapest eggs. He’d probably answer - he might not even say Whole Foods - but WF is hardly teaching all their staff the best methods to answer this question in training. (Heh, training.)
Comment by siva7 6 hours ago
Comment by anon7000 6 hours ago
Comment by MagicMoonlight 4 hours ago
Comment by vanuatu 7 hours ago
model routing in this case is cross-provider
Imo the main issue behind model routing is you need to figure out how much intelligence a new task takes, which is a very non trivial problem. Presumably, a organization knows this about their own tasks and is better suited to built in-house compared to outsourcing to a vendor.
Comment by fny 1 hour ago
Comment by CuriouslyC 7 hours ago
Comment by hnfong 8 hours ago
Otherwise the expensive-yet-powerful model probably won't see much revenue. How much money is there in bleeding edge scientific research? There's a lot, but there's even more existing capital in paying people people to do college level paperwork, and the bulk of those traffic gets routed to the cheapest model.
You mostly don't need super powerful AGI to replace the paper pushers, but the frontier labs are trying to position themselves as being uniquely capable of producing super powerful AGI, and also be the ones replacing office workers.
Not sure how it will work out for them, but I think model routing is going to poke holes in that narrative. That's why I think they're trying very hard not to understand model routing exists.
Comment by ip26 4 hours ago
Comment by TeMPOraL 7 hours ago
For me, anything other than current best available SOTA for any task is unacceptable. The only routing rule I need is "the most powerful model I still have flat-priced quota available for". I mean, why settle for less?
Comment by btown 7 hours ago
Model routing for subsidized users takes the form of a "use Opus 5 subagents for implementation" type of system prompt. You lean into a single provider, build tooling around that, and your savings are far beyond anything multi-provider routing can get you.
Model routing for enterprises is far more complex - approaches like https://fireworks.ai/blog/kimik3-fable become necessary for cost control.
Comment by pzo 6 hours ago
There is also matter about convenience - when I ask some small easy question often I don't bother to switch the model or forget in prompt to ask faster/cheaper subagent.
Comment by kxxx 7 hours ago
Comment by lostmsu 7 hours ago
Comment by paxys 6 hours ago
Comment by Slartie 7 hours ago
Also: quota. Implies you do not have unlimited access even for flat prices. Which in turn implies that as soon as you hit the quota on the most expensive flat price plan, even you will suddenly discover the magic of economically sensible behavior.
Comment by msabalau 7 hours ago
Certainly if I'm confident that I'm going to get what I need from a faster model, that's what I want to use, rather than wasting time grinding away for the sake of saying of the same answer came from a SOTA model.
Given that every chatbot does offer a range of models, it seems clear people do choose among options.
Comment by internet2000 7 hours ago
I just want to switch to Claude Code, tell it to turn a .csv into a BigQuery table then cmd+tab to something else while it runs. Thinking "oh this is probably an easy task, I can /model to Sonnet to save $0.0004" is silly.
Comment by polotics 7 hours ago
Comment by kxxx 7 hours ago
Comment by jnwatson 7 hours ago
Comment by StilesCrisis 4 hours ago
Comment by TacticalCoder 6 hours ago
Then you must route. An article with lots of upvotes yesterday or two days ago showed that K3+Fable 5 was more SOTA than either of those.
Comment by toss1 7 hours ago
OFC, YMMV
Comment by awongh 7 hours ago
For coding my own work I don't trust the model router, and it would have to be shown to be to save a real dollar amount.
From a buying perspective it's a hard sell to save x but lose out on bugs you are probably introducing at an unquantifiable severity and frequency. How much is it worth to hedge your bets by doing every single inference request on the frontier model?
How much will it cost to go back later and fix things, but also the meta question of how to be able to decide on a hypothetical unknowable? (You'll never know how much better or worse your code was gonna be, it's untestable at a project level)
Comment by owenthejumper 5 hours ago
Comment by ModernMech 5 hours ago
Comment by simianwords 8 hours ago
Comment by verdverm 8 hours ago
I would expect routers to commodify like tokens.
Comment by torginus 7 hours ago
Comment by verdverm 7 hours ago
Comment by TacticalCoder 6 hours ago
If that is true, model routing is here to stay.
It also seems to validate the minimalist approach of pi.dev, where sub-agents from the same company is not the preferred approach (pi.dev believes in neither sub-agents all from the same company nor MCP even you can do it if you want for pi.dev's philosophy is to do add any functionality you want to a minimal harness).
Now of course we'll get for a few weeks all the Anthropic fanbois and shills explaining that "sure, K3 was basically at the level of Fable 5 but now that Opus 5 is out, open-weights models are six months behind".
Comment by deet 5 hours ago
Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"
We need an "annoying English" benchmark.
- Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d...
- Opus 5 Max: https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc42692...
Comment by yesitcan 2 hours ago
Comment by Kwpolska 5 hours ago
Comment by wfme 3 hours ago
Comment by hodgehog11 4 minutes ago
Comment by kranke155 3 hours ago
Comment by arizen 5 hours ago
Comment by solenoid0937 4 hours ago
Comment by dwaltrip 3 hours ago
Comment by merlindru 2 hours ago
for now all they've got is english, so they'll just bend that into shape. it'll do.
Comment by trinari 2 hours ago
Comment by bobbylarrybobby 15 minutes ago
Comment by markab21 3 hours ago
I love it for a few things, but it's gotten really hard to spend any extended amount of time with it because of the lack of mental model I seem to be able to hold while working with complicated problems.
I'm guessing it's just not enough time doing RL on human feedback.
Check out the anouncement of Inkling (https://thinkingmachines.ai/news/introducing-inkling/)... the section in the middle
"Early in RL verbose, grammatical" (if you search) :
We need to understand the operator. The 5D line element is ds² = e^{2A(x)} (ds²_4d + dx²), where A(x) = sin(x) + 4 cos(x), x in [0, 2π]. The internal coordinate is periodic. The background is a warped product: metric g_{MN} where M,N = 0..4. The internal direction has metric e^{2A(x)} dx²? Wait, the ds² is e^{2A} (ds²_4d + dx²). So the internal metric is e^{2A(x)} dx². Actually if the total metric is ds² = e^{2A(x)} (ds²_4d + dx²), then yes, internal metric is e^{2A} dx².
vs. Post RL
We need determine eigenvalue problem for spin-2 fluctuations h_{μν}(x,y) with TT in 4d and depend on x. For metric of form ds² = e^{2A(x)} (g_{μν}(y) + h_{μν}(y,x)) dy^μ dy^ν + e^{2A(x)}? Wait internal metric is e^{2A} dx²? Actually ds² = e^{2A} [ds_4² + dx²]. So internal metric is e^{2A} dx²; warp factor same for 4d and internal? Yes. We need equation for h_{μν}(y,x) = h_{μν}(y) ψ(x) maybe with normalization. …
I can understand it with less cognitive load in the post-RL version versus early in RL. This resonated with my experience using Fable, especially digging hard problems; it feels like I'm reading the "early in RL" version of that model explanation.
Comment by kanodiaayush 3 hours ago
Comment by arjie 2 hours ago
Comment by DonsDiscountGas 4 hours ago
Comment by robwwilliams 1 hour ago
Comment by winwang 5 hours ago
I'm also thinking of another benchmark: (quantified) stylistic range across different prompts. Just putting it out there if anyone wants to do the work for me :D
Comment by sibeliuss 25 minutes ago
Comment by ianberdin 5 hours ago
Comment by duplessitous 5 hours ago
Comment by rb2e 8 hours ago
Comment by shwaj 8 hours ago
Comment by dd8601fn 8 hours ago
Comment by benjiro29 5 hours ago
Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8).
Already posted this before, so here is the link.
Comment by MagnumOpus 4 hours ago
Comment by benjiro29 4 hours ago
You see the issue, if you try to scale effort down, you also need to compare how other competing models compare.
Comment by binsquare 7 hours ago
Comment by Freedumbs 5 hours ago
Comment by ceejayoz 8 hours ago
Almost as good for half the cost is something I'm very comfortable describing that way.
Comment by lelanthran 6 hours ago
It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar).
Comment by moffkalast 5 hours ago
Got an endless list of stuff done with Fable, Opus 4.8 was like a flailing braindead idiot in comparison. Maybe this one is a bit better if it's distilled.
Comment by ProofHouse 8 hours ago
Comment by tshaddox 8 hours ago
Comment by ActivePattern 8 hours ago
Comment by adam_arthur 8 hours ago
Where are you getting cheaper per dollar?
Comment by ActivePattern 8 hours ago
Comment by adam_arthur 8 hours ago
Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels.
There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving)
If they had a more efficient model at coding they would lead with that chart.
Comment by km144 7 hours ago
https://artificialanalysis.ai/models?cost=intelligence-vs-co...
Here is another data point for output token efficiency:
https://artificialanalysis.ai/models?cost=intelligence-vs-co...
Comment by HarHarVeryFunny 7 hours ago
Comment by conradkay 7 hours ago
It seems roughly equal according to Anthropic's benchmarks
Comment by shwaj 8 hours ago
How big of a lie is too big? Especially when no lie needed to be told at all: many including myself would have noticed the tiny 0.1% deficit and been suitably impressed by the Opus 5 result.
I’ll admit this is a small deception by today’s standards. I’m one of those who believes in truth for truth’s sake.
Edit: typo
Comment by manojlds 8 hours ago
Comment by jsLavaGoat 8 hours ago
Comment by toephu2 7 hours ago
Comment by dbbk 7 hours ago
Comment by Aurornis 8 hours ago
Fable is typically used for key planning, architecting, and review tasks.
I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.
Comment by airstrike 8 hours ago
Comment by Aurornis 7 hours ago
If you bought the $200/mo plan and you don’t use it much, using Fable for everything is fine.
Comment by maineldc 7 hours ago
Comment by akmarinov 8 hours ago
Just this past week Fable was able to figure out a couple of small issues for me where Opus was failing to.
Also both are still somewhat bad at UI implementation. Opus more so
Comment by unclebucknasty 7 hours ago
"Use <less expensive or older model> for everyday tasks and <other non-critical stuff>. Use <more expensive or recent model> for complex coding tasks, refactoring large code bases, etc.".
Then, the next model/release emerges and the previous "best for complex" gets demoted to "everyday".
Obviously, it's all relative. But, it does beg the question: was the previous model really good for complex coding tasks or no? I mean, how is it now suddenly only good for the "easy" stuff?
Comment by hvb2 7 hours ago
Because your expectations have changed.
Comment by unclebucknasty 5 hours ago
Comment by entropicdrifter 8 hours ago
Comment by siwakotisaurav 8 hours ago
Comment by kossae 8 hours ago
Comment by lattalayta 6 hours ago
Comment by nerdsniper 8 hours ago
---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
Comment by nightpool 8 hours ago
That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable)
Comment by tadfisher 8 hours ago
What variance is acceptable to publish without a retraction?
Comment by nerdsniper 8 hours ago
Comment by tadfisher 8 hours ago
Comment by Atotalnoob 7 hours ago
Like the other person said 5% variation is probably expected
Comment by jll29 6 hours ago
The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic. A proper methodology would ask each question 20 times and calculate the mean correctness across experiments.
The reason is that the temperature parameter introduces random behavior.
Comment by aleenz1102 8 hours ago
Comment by Ancalagon 8 hours ago
Comment by ssalka 8 hours ago
Comment by HyperL0gi 8 hours ago
Comment by impulser_ 8 hours ago
These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards.
OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.
Comment by HyperL0gi 8 hours ago
Comment by K0nserv 5 hours ago
Comment by SatvikBeri 1 hour ago
* Generate misleading news articles
* Impersonate others online
* Automate the production of abusive or faked content to post on social media
* Automate the production of spam/phishing content
Seems like the prediction was pretty accurate.Comment by jgilias 4 hours ago
Comment by reciprocity 3 hours ago
Comment by dzonga 6 hours ago
since then I have never cared about models except those that affect money in my pocket e.g AWS Nova Sonic
Comment by baq 7 hours ago
Comment by neuronexmachina 8 hours ago
Comment by HyperL0gi 7 hours ago
- https://www.axios.com/2026/04/08/anthropic-mythos-model-ai-c...
- https://www.axios.com/2026/04/07/anthropic-mythos-preview-cy...
- https://www.businessinsider.com/anthropic-mythos-latest-ai-m...
- https://www.reuters.com/world/anthropic-ceo-dario-amodei-arr...
Comment by johnfn 6 hours ago
Comment by NichoPaolucci 7 hours ago
Comment by reducesuffering 4 hours ago
OpenAI Huggingface breach begs to differ
Comment by knuppar 7 hours ago
i think we'll see one of the fastest deflations in history post anthropic/oai ipo
Comment by efficax 6 hours ago
Comment by vl 3 hours ago
Comment by efficax 3 hours ago
Comment by scrollaway 2 hours ago
Or maybe you just don't know exactly how capable these models are. Most people's experience of AI is a stupid chatbot, it's no wonder they don't understand how these things are coming for their jobs.
On my end, I have a software that is designed and built by Claude, that I did a strategy session on (with claude), and prepared a fundraise for (with claude). My only role, other than "knowing what to aim for", has been to feed the AI some fairly basic english prompts for a few weeks... which is also easily automatable.
Everyone's job is fucked. Devs, CEOs, everyone.
Comment by le-mark 20 minutes ago
It’s curious to me that there are two distinct factions here. People like parent commenter who has no discernment and others who see llms for what they are. I just talked to opus 5 and in it’s first response caught some well disguise BS. These things are bullshit machines. There are indeed a lot of bullshit jobs around so maybe parent does discern something I don’t?
Comment by fatata123 25 minutes ago
Comment by iLoveOncall 3 hours ago
Comment by skohan 5 hours ago
Comment by websap 8 hours ago
Comment by bottlepalm 7 hours ago
Comment by HyperL0gi 6 hours ago
Comment by MostlyStable 6 hours ago
Comment by emp17344 6 hours ago
Comment by pixl97 2 hours ago
Emp: "what a bunch of lies, I bet they don't even do anything over there"
Comment by 6thbit 8 hours ago
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
Comment by HarHarVeryFunny 7 hours ago
They do say that (implicitly unlike Mythos) Opus 5 was not trained to exploit software vulnerabilities, which would certainly make it safer in that regard.
"As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats."
Comment by square_usual 7 hours ago
Comment by eli 4 hours ago
I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.
Comment by matt2000 3 hours ago
Comment by eli 3 hours ago
So most look like that but I did include a few one-shot “build an app that solves this problem” and some qualitative design tasks and a tough algorithmic optimization one.
Comment by bisonbear 41 minutes ago
Comment by matt2000 3 hours ago
Comment by eli 3 hours ago
Comment by llelouch 7 hours ago
Comment by lifty 6 hours ago
Comment by solenoid0937 4 hours ago
Comment by 6thbit 6 hours ago
Comment by usef- 2 hours ago
Comment by gallerdude 8 hours ago
Comment by flakiness 7 hours ago
Comment by Dibes 8 hours ago
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
Comment by underyx 7 hours ago
Comment by axus 3 hours ago
Last week it felt like Opus 4.8 was moving the Pro "usage" meter very quickly. Today, pre-announcement, Opus 4.8 Medium felt like there was less meter-use per minute. And post-announcement, Opus 5 Medium also feels more efficient, allowing more work in the 5-hour window.
Completely subjective, of course.
Comment by doginasuit 3 hours ago
Comment by mbil 6 hours ago
Comment by km144 7 hours ago
Comment by 2001zhaozhao 7 hours ago
Comment by artemisart 6 hours ago
> We report FrontierCode’s overall score, a composite measure that grades each patch on blocking functional criteria (held-out unit tests) together with weighted code-quality rubric criteria, as mean@5.
They don't explain more in the system card, I guess higher effort levels could loose points on the code quality / scope / style / maintainability stuff?
Comment by landrew_ 7 hours ago
Comment by steve_adams_86 6 hours ago
Comment by Dibes 7 hours ago
Comment by acmnrs 8 hours ago
> Claude Opus 5's default user-facing responses run longer than prior Opus models'.
The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher.
This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
Comment by lukeschlather 8 hours ago
I think in the long run tokens are probably the wrong thing; it's compute and cache memory that you need to be measuring, and when you look at it that way I suspect in most cases the models have pretty similar performance.
Comment by adamtaylor_13 7 hours ago
I don't need more powerful models, I need one that responds fast enough that my attention doesn't wander to other tasks. Grok 4.5 is so fast I can just use it in-band without swapping to other tasks.
Slower than Opus 4.8, which was already miserably slow, is indeed a step in the wrong direction.
Comment by zoicsoftware 3 hours ago
Comment by MarceColl 6 hours ago
Comment by bredren 4 hours ago
During post-training of opus 5, the last few days, opus was a real wreck. I had to swap in gpt 5.6 sol for my orchestrator and enable fast mode (1.5x speed) in order for it to keep up with work and communications from a handful of mostly 5.6 sol agents.
Also because interacting with a slow orchestrator is no fun, even when plenty of work is getting done in parallel in the background.
Comment by eli 3 hours ago
I have a benchmark to build a game engine from a set of written instructions. It's a little tricky. Opus 4.8 did it in 470k tokens at a cost of $1.29 vs Opus 5 in 179k tokens for $0.33. (Fable 5 did it in 245k for $0.95)
Though if you really want to cut costs, Tencent's Hy3 model also got it right and did it in 283k tokens for $0.03
Comment by overfeed 8 hours ago
Gemini also had modest increase before this - don't be surprised when OpenAI also has a "modest increase" with its next release. Cartel-like behaviour doesn't require direct communication when none of the participants are interested in participating in a margin-destroying price-war. All one needs to do is raise their price and watch how the competition react.
Such a scheme (and resulting high margins) would be imperilled by the existence of frontier open-weight models in the market, which may be why the reaction to Chinese models may be particularly shrill.
Comment by simianwords 8 hours ago
No I will be surprised and I'll bet on the fact that prices will keep going down, just like it went ~50% down in the latest GPT 5.6 release.
Comment by falkensmaize 2 hours ago
Comment by firemelt 34 minutes ago
Comment by holtkam2 8 hours ago
Comment by aesthesia 7 hours ago
Comment by alansaber 8 hours ago
Comment by edumucelli 8 hours ago
I was dividing my work between Codex and DeepSeek. Now I barely use DeepSeek, or never because Codex quota is enough after Sol
Comment by elbear 7 hours ago
Comment by wahnfrieden 8 hours ago
Comment by copperx 7 hours ago
Comment by jjcm 7 hours ago
To be fair though, Sol tends to go off the rails sometimes. It's much less reliable than Fable in its outputs. It tends to be overzealous in its research/changes.
Comment by frenchie4111 8 hours ago
Comment by winwang 5 hours ago
Comment by Sol- 8 hours ago
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
Comment by rdedev 8 hours ago
Comment by theHocineSaad 7 hours ago
Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason?
[0]: https://artificialanalysis.ai/#intelligence
[1]: https://platform.claude.com/docs/en/about-claude/pricing
[2]: https://platform.claude.com/docs/en/about-claude/models/over...
Comment by pietz 6 hours ago
I have been trying to build something that captures the behavioral element of different models, but it's kinda tough.
Comment by dgellow 6 hours ago
Comment by hangrybear666 2 hours ago
Comment by anuramat 1 hour ago
Comment by not_a9 8 hours ago
Okay so it’s worse than Opus 4.8 for my purposes I guess?
Comment by sebzim4500 8 hours ago
Comment by bobbylarrybobby 11 minutes ago
Comment by tyre 8 hours ago
Comment by yukIttEft 8 hours ago
Comment by not_a9 8 hours ago
Comment by ciefa 6 hours ago
I recently created a patch for Riftborne via static IL patching and Fable 5 outright kept refusing to do it, no issue whatsoever with GPT 5.6 Sol lol.
Comment by roboyoshi 6 hours ago
Comment by jaggederest 7 hours ago
Comment by derac 5 hours ago
Comment by atraac 8 hours ago
Comment by hoppp 8 hours ago
Fixing those issues still requires humans.
Comment by websap 8 hours ago
Comment by amelius 8 hours ago
Comment by artursapek 8 hours ago
Comment by JackSlateur 8 hours ago
Comment by SketchySeaBeast 8 hours ago
Comment by arm32 8 hours ago
Comment by SketchySeaBeast 8 hours ago
934 days since people first started threatening that devs would be replaced by AI in 365 days. 0 day(s) since Anthropic posted a developer job posting.
Only one of those numbers would need to be dynamic.
Comment by subscribed 6 hours ago
Specialist headhunters handle that.
Comment by moonu 8 hours ago
Comment by arm32 7 hours ago
Comment by DaSHacka 7 hours ago
Comment by arm32 6 hours ago
Comment by londons_explore 8 hours ago
Imagine you are a company that sells concrete. You have a web dev contractor you use to build and maintain your website. It has tools on it to get delivery quotes and a few internal tools to track orders.
Except now you can just have your sales team also maintain the website with a $20/month Claude subscription.
Comment by atraac 6 hours ago
Comment by SketchySeaBeast 7 hours ago
Comment by dawnerd 7 hours ago
Comment by teaearlgraycold 8 hours ago
Comment by igregoryca 7 hours ago
Comment by worldthruword 8 hours ago
Comment by lardosaurusrex 8 hours ago
and you can only kill weyoun, awaken the next vorta clone and have him 'catch up' on all that its missed so many times before they just end up with a complete mess, so. uh. yeah.
doubt they can just "fix" their problems like that.
Comment by JLO64 8 hours ago
Comment by leobuskin 7 hours ago
Comment by manuisin 6 hours ago
Comment by bredren 4 hours ago
What terminal tooling are you using?
Comment by hangrybear666 2 hours ago
Comment by atraac 7 hours ago
Comment by oceanplexian 6 hours ago
They start hitting timeouts or API errors at the same time on two different computers. As far as I can tell it’s the exact same infrastructure.
Comment by einsteinx2 5 hours ago
Comment by acedTrex 8 hours ago
- Boris
Comment by dbbk 7 hours ago
Comment by xixixao 8 hours ago
The first page of the score card mentions that this model is not capable to replace engineers.
Comment by jedberg 7 hours ago
Comment by lossolo 7 hours ago
And memory leaks.
Comment by bellowsgulch 8 hours ago
So dangerous! I can't believe they let the public use this technology! /s
Comment by ealready_value 8 hours ago
Comment by bonoboTP 4 hours ago
They could say "tech report" but model card makes it clear that it's a specific kind of tech report.
Comment by CHUNK_CHUNK 4 minutes ago
Comment by CHUNK_CHUNK 30 minutes ago
Comment by redbell 1 hour ago
I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason.
I submitted a an appeal describing that I am 100% sure I haven't broken any rules and that it was my very first project but, unfortunately, after about 20 days now, nothing seem to be happening.
The thing that hurts me the most is that I had the same experience in the very first days of Anthropic. They suspended my account immediately after I submitted the first prompt, I commented back then (https://news.ycombinator.com/item?id=39698788) and fortunately, someone from Anthropic reach out to me via X and helped me get my account back.
To be honest, I haven't used Claude much since then but when I decided it's time to give it a try, they locked me out again! For reference, the account I used recently is relatively a new one but the activity is crystal clear that it is fair use.
Comment by copperx 1 hour ago
Comment by redbell 42 minutes ago
Comment by abroszka33 8 hours ago
Comment by Diogenesian 7 hours ago
This is snarky but I am grumpy: I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.
Comment by unshavedyak 7 hours ago
Semi related, but i would hate to read that PDF but i also hate reading what LLMs write lol.
LLMs are pretty terrible at being concise. Using an LLM these days means putting up with bizarre and often confusing phrasing, wordy explanations, etc. It's kinda crazy to me how good they are but how bad their writing style is for me personally. Even though i use an LLM constantly i can't stand reading its responses.
Comment by abroszka33 7 hours ago
Maybe it's just me, but 150 pages is like third of a good book. Quite long. And it's full of LLM slop, they did not even bother to remove the em dashes.
Comment by no_multitudes 6 hours ago
I'm not saying you're wrong btw; I'm sure this has many authors and some of them probably used LLMs significantly in the writing process.
Comment by abroszka33 5 hours ago
I'm not saying it's impossible, but I'm more confident about winning the lottery next week.
Comment by no_multitudes 4 hours ago
It's probable that LLM text was pasted directly into early drafts of the document, and plausible that some of that text survives in the final document.
However, no section of the final document I have looked at reads to me like un-edited LLM output (which is almost always very obvious to me.)
Therefore, I think it is more likely than not that human editors went over the document carefully and rewrote anything that was full of the uselessly punchy sentences or constant over-corrections that hallmark LLM speech.
Comment by brokencode 4 hours ago
You can use an LLM to create work that isn’t slop. And you can hand write slop with no computer involvement at all. Most of the people I knew in high school 15 years ago would write slop on a daily basis.
Comment by dgellow 6 hours ago
Comment by kubb 3 hours ago
There are many ways to read something, model cards are usually skimmed.
Comment by srveale 8 hours ago
It's okay if you're not the target audience for one or the other.
Comment by stravant 7 hours ago
They're not meant for normal consumers who just want to use the model for work.
Comment by bredren 4 hours ago
Not just in what the models can or might want to do, but how they treat the operators they interact with.
If you look carefully, this card shows the addition of a new benchmark for "condescension" as a character trait.
I think a lot of people would like to see a comparable system card for the unannounced model that escaped openai last week.
Comment by reasonableklout 7 hours ago
And lots of folks read these. For example here's simonw's notes on the Claude 4 system card: https://simonwillison.net/2025/May/25/claude-4-system-card/
All of this seemed like utter sci-fi just a couple years ago. Do you think that frontier AI companies should be less transparent?
Comment by skeptic_ai 7 hours ago
Comment by bonoboTP 7 hours ago
Comment by neuronexmachina 8 hours ago
Comment by volkk 8 hours ago
Comment by ajmurmann 8 hours ago
Comment by coffeebeqn 8 hours ago
Comment by thewebguyd 8 hours ago
Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
Comment by foota 8 hours ago
Comment by not_a9 8 hours ago
Comment by himata4113 8 hours ago
Comment by cael450 8 hours ago
Comment by ithkuil 7 hours ago
Comment by rdedev 8 hours ago
Comment by subscribed 6 hours ago
Apparently this is the way - if you know, you know :)
Comment by SubiculumCode 8 hours ago
Comment by SubiculumCode 8 hours ago
Comment by adastra22 8 hours ago
Comment by himata4113 8 hours ago
Comment by ray__ 8 hours ago
Comment by trollbridge 7 hours ago
Comment by gillesjacobs 8 hours ago
Comment by wuhhh 5 hours ago
Comment by jakubmazanec 53 minutes ago
Comment by thefourthchime 4 hours ago
Comment by thousand_nights 2 hours ago
i'm still pretty confident someone like my mom wouldn't be able to do my job even with the same access to all the latest LLMs, so we're still providing some value, just in a very different way. whether the market will reprice the cost of our labor, we will see
Comment by tripleee 2 hours ago
That's yet to happen. 90% of software dev skills are still relevant - AI is, for now, just a productivity boost.
Comment by tripleee 5 hours ago
Comment by slices 5 hours ago
Comment by wuhhh 5 hours ago
Another thing that helps is pointing it to patterns in an existing codebase (e.g. "use the box-link pattern for cards, as shown in [..]").
EDIT: The point being that even if they make mistakes that are easy to spot and fix _now_, you'd have to assume that in the very near future those kinks will be ironed out - I mean, the capabilities are only going in one direction.
Comment by morbicer 5 hours ago
Thanks out can also hook it to Playwright with Axe and let it run assessments.
Comment by yewenjie 8 hours ago
Comment by Stevvo 4 hours ago
Comment by dbgrman 3 hours ago
Comment by crewindream 3 hours ago
Comment by prirun 4 hours ago
Comment by eli 4 hours ago
Comment by pyridines 8 hours ago
> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
Comment by MallocVoidstar 8 hours ago
Comment by deweywsu 2 hours ago
Comment by unsupp0rted 2 hours ago
Comment by matheusmoreira 1 hour ago
> Identifying bugs in code is a core part of the secure software development lifecycle, and unblocking this allows for software engineers and coding hobbyists alike to produce more secure code, reducing new vulnerabilities put out into the world.
Not happy with these annoying "safeguards" but at least it's a step in the right direction. Looks like Opus 5 has the same vulnerability detection performance as Fable 5 and that makes it worth it for code review.
Comment by artninja1988 8 hours ago
Comment by mcbuilder 7 hours ago
Comment by vadansky 6 hours ago
I miss him... But for reference he did get past Doom and got pretty far in the strength puzzle too before he cut cut off. He was looping and just brute forcing it.
Comment by JacobAsmuth 3 hours ago
Comment by modeless 8 hours ago
I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.
Comment by wyre 8 hours ago
Comment by conradkay 7 hours ago
I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.
Comment by dominotw 8 hours ago
Comment by password54321 6 hours ago
Comment by criddell 6 hours ago
Maybe hook up a bunch of the AIs to a stereo camera and a couple of microphones and give them control over actuators to so they can drive cars. Then lets race them around a somewhat complex course.
When they are good enough at driving on tracks, put them on the road. Maybe see which can drive a truck with 400 cases of Coors from Texarkana, TX to Atlanta, GA and back within 28 hours.
Comment by bonoboTP 4 hours ago
Comment by criddell 4 hours ago
Comment by Stevvo 4 hours ago
Comment by layer8 8 hours ago
I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.
Comment by awestroke 8 hours ago
Comment by cebert 8 hours ago
Comment by user43928 8 hours ago
It seems plausible to me that RL improvements allowed Anthropic to improve on Opus 4.8, similar to how OpenAI substantially improved upon GPT 5.5 with 5.6 Sol.
Fable 5.1 and GPT-6 are rumored to launch in August, presumably bringing those improvements to the larger models.
Comment by HarHarVeryFunny 6 hours ago
I don't know how systematic Anthropic are about their versioning - I'd have guessed that major version number increases (4.x -> 5.x) reflect different base models (different pre-training runs), in which case Opus 5 would be a distilled version of the Fable 5 base model (but without the cyber exploit post-training), rather than being Opus 4.8 with additional post-training, but who knows? I don't believe Anthropic have said anything about this.
Comment by aesthesia 1 hour ago
Comment by bonoboTP 4 hours ago
Comment by HarHarVeryFunny 3 hours ago
Comment by cesarvarela 8 hours ago
I think the best proxy for this feeling is the Artificial Analysis' omniscience index. Fable has a 40 score, and Opus (4.8) has 27.
Comment by modeless 8 hours ago
Comment by Tenoke 7 hours ago
Comment by trentor 8 hours ago
Comment by logicchains 7 hours ago
Comment by r1ch 3 hours ago
Comment by braebo 1 hour ago
Comment by ianberdin 6 hours ago
It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
Comment by israrkhan 1 hour ago
Comment by uncivilized 1 hour ago
Comment by guybedo 7 hours ago
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...
Comment by conradkay 7 hours ago
I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
Comment by JacobAsmuth 3 hours ago
A useful measure of real world cost (complementary with total cost like they already report, of course) would be "cost for correct answers". You could look at the ratio between the two costs to get a measure of laziness which many would find quite useful.
Comment by adverbly 4 hours ago
It did far better at some tasks compared to Sol (e.g. the ARC 3 benchmark). And at those tasks, it's not just "a bit smarter": It got 30% vs less than 8% - so you're talking 2.75x more for almost 4x the coverage.
Comment by alphabettsy 5 hours ago
There’s also the frustration of it not quite being enough sometimes. It’s extremely capable, but I still find that it needs more concrete guidance and boundaries than other models.
Comment by I_am_tiberius 7 hours ago
Comment by adamtaylor_13 5 hours ago
If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.
Comment by JacobAsmuth 3 hours ago
Comment by dist-epoch 4 hours ago
Comment by wxw 3 hours ago
This is a pretty common trading firm internship project funnily enough.
Comment by visiondude 8 hours ago
Comment by stri8ted 7 hours ago
Comment by ahsillyme 2 hours ago
Comment by braebo 1 hour ago
Comment by modeless 8 hours ago
Comment by oh_no 8 hours ago
Comment by modeless 8 hours ago
Comment by dominotw 8 hours ago
why is that? its now being benchmaxxed too
Comment by williamstein 8 hours ago
Annoyingly, this is a concrete argument that open source software may be easier to attack.
Comment by ddxv 8 hours ago
Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.
Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
Comment by layer8 8 hours ago
Comment by ReptileMan 8 hours ago
In the next - please scan this totally mine code for vulnerabilities
Comment by zb3 8 hours ago
Comment by redsocksfan45 6 hours ago
Comment by petilon 7 hours ago
Comment by einsteinx2 7 hours ago
Though it gets even more confusing because they also have effort levels so it’s not really possible to call one fast and one slow since Fable on Medium will be faster than Opus on Max.
I agree it’s confusing, and now OpenAI is following Anthropic’s lead with their new naming (Sol, Terra, Luna).
Comment by bonoboTP 6 hours ago
A similar complaint was valid years ago when OpenAI had GPT-4o, o1, o3 (but no o2), o4-mini-high, GPT-4, and GPT-4.1 and GPT-3.5 etc.
Comment by einsteinx2 5 hours ago
Arguably the complaint was more valid for those older GPT models you mentioned.
Comment by bonoboTP 5 hours ago
Some models like ViTs use something similar but then introduce words with no unambiguous order, like Small, Medium/Base, Large but then I always forget if Huge or Giant is larger.
Comment by paxys 6 hours ago
Comment by einsteinx2 5 hours ago
Also fwiw I’ve never found LLM benchmarks to match reality based on my own usage, not for the large frontier models or smaller open weight models so who knows if Opus is actually better than Fable (I doubt it).
Comment by tackta 4 hours ago
Fable 5.1 or whatever they go with will be the stronger version vs Opus 5.
From about 2 hours of Opus 5 use , I would say it is quite impressive.
Comment by JacobAsmuth 3 hours ago
Comment by hk__2 7 hours ago
> Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
This is not possible: Standard (Free) / Pro / Max are plan names. Fast is a mode.
Comment by abalaji 7 hours ago
Comment by bonoboTP 4 hours ago
And fables are not particularly long actually.
Comment by petilon 6 hours ago
And people know this? I didn't. I am not into music or poetry so these are not terms I am familiar with.
Comment by MostlyStable 6 hours ago
Comment by Mossly 5 hours ago
fable: 99/100 sonnet: 97/100 haiku: 91/100 opus: 89/100
So while these terms are almost universally known, opus is indeed the least known of the four. And I guess this only measures whether a person knows a word, not whether they know an opus is longer than a sonnet! Personally I only inferred that based on the related term 'magnum opus.'
Comment by josefresco 5 hours ago
Comment by alasano 8 hours ago
Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
Comment by mchusma 8 hours ago
Comment by alasano 2 hours ago
Comment by consumer451 6 hours ago
Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
Comment by roboyoshi 6 hours ago
Comment by consumer451 5 hours ago
Something along the lines of: "Please run a full security analysis on the entire project. Make sure user documents are secure."
Just something like that prompt found a vector in my web app's MCP server that I never would have considered. It was very much an edge case, but it did exist.
Being broad allows the model and harness to do the work. Giving too many instructions can apparently work against you in many cases.
Of course, when dealing with new PRs, I use the /security-review and /code-review skills.
Comment by tekacs 7 hours ago
Comment by hrpnk 6 hours ago
1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.
2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.
[1] https://platform.claude.com/docs/en/about-claude/models/migr...
Comment by slymax 4 hours ago
Comment by lucamark 7 hours ago
I've never trusted on model cards though. I'm sorry.
Comment by dbbk 7 hours ago
Comment by lucamark 7 hours ago
Comment by tysilva 2 hours ago
Comment by marcindulak 3 hours ago
The desire of the models to act at the cost of ignoring user instructions is still noticeable.
Comment by irthomasthomas 7 hours ago
Comment by luciana1u 5 hours ago
Comment by bottlepalm 7 hours ago
Comment by trunnell 6 hours ago
From the system card [1]:
The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...Comment by albert_e 8 hours ago
Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
Comment by thewebguyd 8 hours ago
Comment by vatsachak 7 hours ago
It's great with Codex.
I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.
Comment by guess_who_is 5 hours ago
Comment by the_lucifer 8 hours ago
Comment by himata4113 8 hours ago
Comment by kouteiheika 7 hours ago
That's by design. Anthropic wants to make open-weight models illegal (not my speculation -- Dario explicitly said so), so I assume they don't want to give them any undue attention.
Comment by paxys 6 hours ago
Comment by hahahaa 1 hour ago
Comment by seizethecheese 4 hours ago
Comment by adamhowell 5 hours ago
Comment by cheesecakegood 3 hours ago
Comment by firemelt 40 minutes ago
Comment by vinhnx 7 hours ago
Comment by pietz 5 hours ago
But good god, what a steaming pile of bullshit this is. Completely exaggerated and overly technical language over 235 seconds that could have been explained in 30 to a 12 year old.
Trash content doesn't normally frustrate me, because it's usually quite easy to spot trash. But in the time of AI, trash can actually look good at first glance and it needs some actual knowledge to spot its problems.
Sorry for the harsh words, but for the love of humanity stop producing content or do it better.
Comment by vinhnx 2 hours ago
Comment by cheema33 5 hours ago
Comment by vinhnx 2 hours ago
Video Special | Anthropic's Claude Opus 5 + https://www.youtube.com/watch?v=8Vdofv2vQ_M + https://www.youtube.com/watch?v=q-jHHx3J8m8
Comment by b-side 1 hour ago
Comment by doginasuit 2 hours ago
Comment by rad_val 7 hours ago
Comment by CuriouslyC 7 hours ago
Comment by destring 8 hours ago
Comment by skybrian 8 hours ago
Maybe there’s a better comparison than cost per token, but it will be application-specific.
Comment by skerit 8 hours ago
> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud. > > This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
For nearly all models EXCEPT Sonnet 5? That is weird. How old is Sonnet 5 really?
Comment by jatins 8 hours ago
Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
Comment by helloplanets 6 hours ago
So, not a distilled version of Mythos or Fable, but those models likely helped a lot in the post training phase of Opus.
Comment by CaveTech 8 hours ago
Comment by beydogan 6 hours ago
- it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad
- on >xhigh it eats tokens like there is no tomorrow
I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
Comment by dehugger 8 hours ago
Comment by andrewl-hn 8 hours ago
With this iteration they had a delay because when the Mythos was ready they had some sort of "Oh shit" moment and spent half a year adding safety guards to it. Then slowly rolled it out, but got another delay due to a government block. So, maybe the work on making Opus and Sonnet only started after they got a green light from the administration.
Presumably, now that they learned how to do this safety-wrapping the next iteration of Mythos / Fable / Opus / Sonnet is going to show up faster.
Something like that.
Comment by Wowfunhappy 8 hours ago
Comment by anon373839 3 hours ago
Comment by riknos314 7 hours ago
So I'm assuming at least a subset of employees could continue using the models during that time.
Comment by Wowfunhappy 7 hours ago
Comment by tedsanders 6 hours ago
Comment by oh_no 8 hours ago
Comment by somenameforme 7 hours ago
Comment by 6thbit 8 hours ago
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".
This is probably great news, but then again, where does this leave Fable as a choice?
Comment by itissid 6 hours ago
Comment by ianberdin 6 hours ago
Fable is not better, it says zero information between steps and then output a summary. A perfect “send - done”.
Comment by mulhoon 7 hours ago
Comment by furyofantares 5 hours ago
Sonnet 5 and Opus 4.8 seem about the same to me - the reason I switch between the two is I'd read that it's cheaper to use Sonnet 5 on those reasoning levels, and cheaper to use Opus 4.8 above them. This is due to them using different token quantities.
Comment by stsch 6 hours ago
Comment by markasoftware 8 hours ago
Comment by geooff_ 8 hours ago
Comment by dpe82 8 hours ago
Comment by 6thbit 8 hours ago
Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..
Comment by stri8ted 7 hours ago
Comment by vinhnx 7 hours ago
Comment by bouke 6 hours ago
Comment by stevefan1999 8 hours ago
Comment by alvis 8 hours ago
Comment by briandoll 8 hours ago
Comment by m_w_ 8 hours ago
Comment by tyre 8 hours ago
Comment by somenameforme 7 hours ago
Comment by pmg1991 8 hours ago
Comment by aleenz1102 8 hours ago
Comment by twothreeone 8 hours ago
Comment by boc 8 hours ago
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
Comment by jannyfer 8 hours ago
Comment by bovermyer 8 hours ago
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
Comment by orangecat 7 hours ago
I'd be curious to see a version of the test where models are asked to give a probability that their answers are correct so we can see how calibrated they are.
Comment by arjie 7 hours ago
Comment by skinfaxi 8 hours ago
Comment by arrowleaf 8 hours ago
Comment by shockembopper 7 hours ago
Comment by himata4113 8 hours ago
Comment by plqbfbv 7 hours ago
Opus can give better results on architectural/concept tasks and I use it sparingly, but it still costs more than Sonnet 5. Opus 5 seems to achieve results very close to Fable 5 while costing less (keeps Opus 4.8 pricing IIUC), but still more than Sonnet 5 then.
Comment by krmmalik 8 hours ago
So, for coding, for example: Opus for solution design and architectural blueprint and then Sonnet for actual implementation.
Works out cheaper with minimal loss of quality.
At least that's my personal understanding and anecdotal experience.
Comment by theLiminator 7 hours ago
It's only when you need even lower levels of cost than opus at zero to low reasoning when sonnet starts to make sense at all.
Comment by born-jre 6 hours ago
Comment by bonoboTP 4 hours ago
Comment by yusufozkan 8 hours ago
wow
Comment by theplumber 7 hours ago
Comment by whatever1 8 hours ago
Comment by moomin 8 hours ago
Comment by drusepth 8 hours ago
Comment by 8note 8 hours ago
i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
Comment by Uptrenda 1 hour ago
Comment by internet2000 7 hours ago
Comment by korabs 7 hours ago
But they say it's "almost as good as fable"
Comment by urams 8 hours ago
Comment by MasterScrat 4 hours ago
Comment by mcast 8 hours ago
Comment by bellowsgulch 8 hours ago
Comment by doctoboggan 8 hours ago
Comment by inshard 8 hours ago
Comment by spstoyanov 8 hours ago
Comment by tomlockwood 2 hours ago
Comment by mkurz 8 hours ago
Comment by taf2 8 hours ago
Comment by arj 7 hours ago
Comment by Eldodi 8 hours ago
Comment by shinhyeok 3 hours ago
Comment by throwaw12 8 hours ago
Comment by lbrito 5 hours ago
Comment by toephu2 7 hours ago
Comment by toephu2 7 hours ago
Comment by abc42 7 hours ago
Comment by holoduke 5 hours ago
Comment by arseniitrut 6 hours ago
Comment by hmontazeri 8 hours ago
Comment by _pdp_ 6 hours ago
Comment by backscratches 6 hours ago
Comment by jakeogh 5 hours ago
Comment by mihau 8 hours ago
Comment by simianwords 8 hours ago
Comment by ismailmaj 7 hours ago
Comment by zmmmmm 3 hours ago
Comment by LoganDark 7 hours ago
Comment by Footprint0521 6 hours ago
Comment by LoganDark 5 hours ago
Comment by Footprint0521 3 hours ago
Comment by StrauXX 8 hours ago
Comment by skybrian 7 hours ago
Comment by StrauXX 6 hours ago
Comment by sudohalt 7 hours ago
Comment by sp4cec0wb0y 6 hours ago
Comment by throwaway23597 7 hours ago
Maybe I'm wrong and Opus 5 is a real unlock?
Comment by simianwords 8 hours ago
Comment by sbochins 8 hours ago
Comment by alvis 8 hours ago
Comment by zuzululu 8 hours ago
Comment by mrcwinn 7 hours ago
Comment by SoftTalker 6 hours ago
Comment by mrcwinn 1 hour ago
Comment by wyre 8 hours ago
Comment by iLoveOncall 3 hours ago
Comment by mnky9800n 8 hours ago
Comment by justindotdev 8 hours ago
ffs just keep it man.
Comment by alex1138 5 hours ago
Comment by jackjd 1 hour ago
Comment by gorkemyildirim 6 hours ago
Comment by jeffybefffy519 3 hours ago
Comment by vilmire 7 hours ago
Comment by marsven_422 4 hours ago
Comment by nee_oo_ru 8 hours ago
Comment by CurbStomper 1 hour ago
Comment by emunova 7 hours ago
Comment by Nevin1901 8 hours ago
Comment by midnightbobarun 5 hours ago
Comment by datakan 8 hours ago
Ok then so what's the point?
Comment by vidarh 8 hours ago
Comment by varispeed 8 hours ago
I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish
Comment by vidarh 1 hour ago
There are certainly tasks where fable will be faster and/or cheaper, but there are plenty of tasks where even Haiku is as fast or faster and cheaper, or where you can e.g. get away with models like gpt-oss that you can get from inference providers providing 10x+ the token/second speed.
If you don't use enough tokens that relying only on Fable becomes a problem, then keep using just Fable. Personally, for my $200/week Max subscription I'd run out of the weekly quota for Fable in a day. At API pricing I'd go bankrupt if I tried doing the things I do with cheaper models using Fable.
Comment by SubiculumCode 8 hours ago
Comment by sambaumann 8 hours ago
Comment by stuartjohnson12 8 hours ago
Comment by signalchain 4 hours ago
Comment by 8note 8 hours ago
Comment by albert_e 8 hours ago
Comment by tamimio 8 hours ago
Comment by javawizard 8 hours ago
When they release new versions of Sonnet, no-one expects them to be better than Opus.
Comment by serf 8 hours ago
Comment by Infinity315 8 hours ago
Comment by cmrdporcupine 8 hours ago
Comment by occz 8 hours ago
Comment by viccis 8 hours ago
Comment by danielbln 8 hours ago
Comment by LeoPanthera 8 hours ago
Comment by simianwords 8 hours ago
The point is that Opus 5 is the best they can do without needing classifiers and absurdly broad safeguards.
Comment by pferde 8 hours ago
Comment by CaptWorld 8 hours ago
Comment by zormino 8 hours ago
Comment by stri8ted 7 hours ago
Comment by aleenz1102 8 hours ago
Comment by TheJCDenton 7 hours ago
Comment by LoganDark 7 hours ago
Comment by midnightbobarun 8 hours ago
Comment by Footprint0521 6 hours ago