Elevated Errors for Opus 5
Posted by TimCTRL 1 day ago
Comments
Comment by bashtoni 1 day ago
Anthropic seem to be both letting their competitors outplay them and making unforced errors (like the anti Open Weights models stuff and the frequent Fable->Opus downgrades for 'safety').
Outages are inevitable in this early high growth, rapid development era, but tactical errors are not.
Comment by cbg0 1 day ago
OAI waits until overall cluster demand goes down before giving out a reset, they just happen to have much more spare capacity than Anthropic.
Comment by smoe 1 day ago
I don’t think Codex’s 99.98% uptime vs Claude Code’s 99.44%, is just due to OAI having more hardware to run it on.
Comment by CuriouslyC 1 day ago
Comment by odo1242 1 day ago
Having too much demand on limited servers means that less of those servers are going to be available for redundancy.
Comment by catlifeonmars 1 day ago
> it's not like they haven't thought about doing the obvious.
Comment by genidoi 1 day ago
Comment by chaos_emergent 1 day ago
You can also get a drift of it in their announcements - OpenAI has been far more aggressive in datacenter build-out partnerships as a more top-down allocation strategy since 2024, and Anthropic has rented out existing compute from infrastructure-heavy companies that are losing the model race in what seems like a more desperate ad-hoc compute acquisition strategy.
In all, Dario’s conservatism was in service of not going bankrupt, but any inspection of his argument that an overallocation of compute a year too early results in bankruptcy falls flat: there’s so much demand for mechanized intelligence in the world today that it’s easy to rent out compute if you have too great a supply.
[1]: recall that he was trying to court the saudis into investing $1T to build his own chips
Comment by lilytweed 1 day ago
Comment by usef- 1 day ago
Though Dario's reasoning still applies this year, as the scale keeps increasing. Very interesting to know how it will go.
Comment by bob1029 1 day ago
https://qz.com/openai-investor-memo-compute-advantage-anthro...
Comment by tjoff 1 day ago
Comment by beering 20 hours ago
Comment by dchftcs 1 day ago
Comment by ACCount37 1 day ago
Comment by cbg0 1 day ago
There's less clear info about this from Anthropic's side, but given their frequent issues and lack of resets and overall less usage given to users on their subscriptions plans, it's a pretty simple conclusion to draw.
Comment by solenoid0937 1 day ago
we are in the "millennial lifestyle subsidy" era for AI where companies ruthlessly undercut each other in an attempt to win marketshare, before then ratcheting up prices
Comment by hangrybear666 1 day ago
Comment by odo1242 1 day ago
Comment by solenoid0937 1 day ago
Comment by londons_explore 1 day ago
Best would be a free week/month for anyone who sent a request during the downtime.
Comment by solenoid0937 1 day ago
This is bizarrely out of touch. It would be a courtesy but it's not a "reasonable expectation" at all.
At best you could maybe reasonably expect a prorated refund for the time of unavailability, but then Fable was still available.
Comment by ShinTakuya 1 day ago
Comment by londons_explore 1 day ago
Instead you should be compensated for your losses - the time and money it took you to repair that tyre.
Same with a web service. If I pay for 24 * 7 service and it's not working when I go to use it, I want compensation for my time and effort resolving the matter.
For some that'll be small - eg. Searching Google instead, taking 30 seconds. For others it'll be a big headache.
But downtime of 1 hour is almost certainly more costly for the consumer than 1/30/24 * subscription price.
Comment by solenoid0937 1 day ago
Your contract doesn't have an SLA. If you are important enough, you can certainly negotiate one - and it will be far more expensive than your prosumer subscription.
And even then you will not get a 1:168 or 1:720 SLA like you are expecting - that's simply ludicrous and totally out of touch with every reality. Even if you were a lucrative enterprise customer (which you are not) they'd tell you to get lost.
> If there is a hole in the street and you fall down it, you wouldn't expect a refund just for the 1 second
Actually, you wouldn't get a refund at all unless you suffer damage or injury, which is an high bar to prove in court and a totally separate ball game.
Comment by 1123581321 1 day ago
Comment by gozucito 1 day ago
Enough to offset any KV cache miss expenses is the absolute minimum.
The smart play is to refund customers something like $20-50 worth of unsubsidized credits.
Those are high margin and only the equivalent of 45 minutes of Fable usage once the KV cache reload costs are factored in.
Yet customers often need a few credits to finish a job without waiting 5 hours or days for their reset.
Keep in mind openAI already gives me free resets I can use when I want which are very useful to me. Seems a no brainer to tack those to subscriptions, actually.
It gives the user a little more flexibility, a little more control over their tools.
Comment by furyofantares 1 day ago
Comment by Azantys 1 day ago
Comment by einsteinx2 1 day ago
Comment by jm4 1 day ago
I wish it didn’t go down so frequently, but it does. Still, I get an enormous amount of value. I realize they probably prioritize API users over subscribers and I’m ok with that. I use the API and openrouter for the things that need to be resilient to outages.
I would be embarrassed to try to ask for a refund or free credits. They already give credits far in excess of what you pay for and you have practically the entire month to use them outside of a couple hours of downtime.
Comment by bigDinosaur 1 day ago
Comment by anonzzzies 1 day ago
Comment by close04 1 day ago
When you miss 1h of your job, do you give back a month for free?
Comment by croemer 11 hours ago
Yesterday it was down for 1.5hr, today it's 1hr for incident 1, and the second one is open for 10min now.
Comment by bashtoni 23 hours ago
Claude Code built a lot of good will amongst developers. OpenAI are playing a great tactical game to try and catch up by offering things like weekly usage limit resets after outages.
Of course, once the finance types get control after the rapid growth phase is complete the inevitable enshitification will begin with comments exactly like yours.
Comment by throwaway314155 1 day ago
Comment by solenoid0937 1 day ago
I think Anthropic is past this point. They are trying to get to profitability. They literally don't have enough compute to serve their demand.
OpenAI on the other hand I can totally see doing this. They need to gain marketshare and gain it fast or they're screwed.
Comment by matt-p 1 day ago
Comment by vikramkr 16 hours ago
Comment by hedgehog 1 day ago
Comment by stevefan1999 1 day ago
Comment by anon_anon12 1 day ago
Comment by theptip 1 day ago
Comment by hmokiguess 1 day ago
Comment by Wowfunhappy 1 day ago
Comment by vitally3643 1 day ago
I could accomplish a lot more with those tokens if it just asked for the package to be installed.
Comment by Wowfunhappy 1 day ago
Comment by vikramkr 16 hours ago
Comment by vikramkr 16 hours ago
Comment by hmokiguess 10 hours ago
Look at features they release, such as Dynamic Workflows, that spin up 100+ sub agents then reconcile the result.
Does that make sense now?
Comment by vikramkr 6 hours ago
Now with that said, opensi's ultra mode is absolute trash - they should just have stolen Claude code's implementation - and there are clearly modes like max reasoning where they'll burn double the tokens to get another 10th of a percent performance to win benchmarks, which you should basically never be using. They're not making those versions at the expense of more efficient reasoning levels though and I don't mind that they exist (as long as they don't get made the default mode - openai - fix your ultra mode already).
Comment by prologic 1 day ago
Comment by CuriouslyC 1 day ago
Comment by tetha 1 day ago
It's not clear if or when these tools become as essential as power or internet to a developers or admins work, but if they do? Giving everyone a quarter day off at least generates some good-will, unlike forcing them to sit around unproductively because of working hours.
Comment by sajithdilshan 1 day ago
Comment by vidarh 1 day ago
While there is some lock-in for the hardest tasks, for most of what people do LLMs are rapidly becoming a commodity.
In fact, right this minute I have Opus benchmarking Kimi alongside itself to determine which workflows we'll switch to using Kimi by default and Opus as the fallback instead of vice versa. It has already conceded Kimi does better on several tasks.
Comment by ulimn 1 day ago
Comment by vidarh 1 day ago
If I needed to fall back on OpenRouter, it'd be paid per token, of course, but that'd take 4 major providers being down, in which case I suspect I'll have more critical things to worry about.
As it is, I use more than I could with a single $200 Claude Max subscription anyway, so I distribute my tasks across multiple providers.
Even without that it's cheap insurance given the cost of lost time.
Comment by rokkamokka 1 day ago
Comment by danielbln 1 day ago
Comment by vidarh 1 day ago
Comment by ShinyLeftPad 1 day ago
Comment by vidarh 1 day ago
Comment by ShinyLeftPad 1 day ago
Comment by vidarh 1 day ago
Comment by ShinyLeftPad 1 day ago
Comment by vidarh 1 day ago
Comment by ShinyLeftPad 15 hours ago
Comment by Eddy_Viscosity2 1 day ago
Comment by vikramkr 16 hours ago
Comment by senko 1 day ago
Comment by jdthedisciple 1 day ago
Comment by mikeydiamonds 1 day ago
Comment by luciana1u 1 day ago
Comment by claaams 1 day ago
Comment by Schiendelman 1 day ago
Comment by 7734128 1 day ago
Comment by rvz 1 day ago
So Claude is now on vacation and is unavailable.
Comment by conorcleary 1 day ago
Comment by threatofrain 1 day ago
Comment by mrcwinn 1 day ago
Comment by fannning 1 day ago
Comment by dotdev_prem 1 day ago
Comment by afdsaifdoi 1 day ago
Comment by throwaw12 1 day ago
Comment by blurbleblurble 1 day ago