Corporate America is getting hooked on open-source AI
Posted by aaraujo002 4 days ago
Comments
Comment by cmiles8 4 days ago
Unless they both dramatically slash prices then they’re in big trouble. Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any hope at a successful IPO.
However the cold reality for both is that there is zero moat to a model anymore. It’s a pure commodity. Those selling compute and access to open models are gearing up to wipe the floor with Open AI and Anthropic.
Comment by Gareth321 4 days ago
The moat right now is a) the hardware, b) the electricity, c) the intelligence, and d) scalability.
On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision.
On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places. In California or Europe this could be $30-60k per year.
As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less.
The last major moat is the ability for subscriptions to scale with need. This means easily adding and removing licenses. This is far easier than purchasing extremely expensive hardware (and managing it), and selling it if/when internal demand changes. It's the same reason companies use contractors. The ramp up/down costs are very high.
The only real moat that local LLMs have right now is privacy.
Comment by jurgenburgen 3 days ago
OpenAI and Anthropic rent their compute from AWS & friends. When we say large enterprises are moving to open weight models it means they are cutting out the middleman and renting the compute directly from AWS instead of giving OpenAI and Anthropic a margin.
> they have at least a 3-6 month head start
This is a moat of nothing. Our company still hasn’t gotten access to Fable so switching to open weight models would mean getting access to similar quality models. In some orgs they are still on 2025 models.
Comment by sedansesame 4 days ago
All it takes is one incident, and all your company's internal data will start showing up in public users' chats. You can rely on a contract to prevent this, or you can guarantee it by using a locally hosted model you fully control.
When combined with the cost savings and good enough performance mentioned in the article, this can become a huge selling point.
Comment by autoexec 4 days ago
Corporate doesn't care at all about privacy. It's why the run outlook and windows and let Microsoft scoop up all of their company secrets. It's why they hand every scrap of data they have on their customers to salesforce and surrender their data to Atlassian and use Confluence and JIRA over countless alternatives.
All companies care about is that when data breaches happen publicly they can point the finger at someone else.
Comment by fluidcruft 2 days ago
And with things like Fable mandating data retention at Anthropic it makes things business records which US Federal Administrations have a history of obtaining without warrants and under mandatory gag orders like national security letters. They generally won't issue those directly to law firms (because they can and will fight), but IT companies generally comply quietly and have government contracts to think about.
Comment by a34729t 3 days ago
Comment by calgoo 3 days ago
Comment by robrorcroptrer 4 days ago
Comment by anon373839 3 days ago
Comment by moate 4 days ago
If however, our internal IT group didn't read the manual correctly and set up our cloud service access parameters incorrectly so that anyone who was trying to probe our defenses could access the formula, Jim from IT is looking for a new job.
See the difference?
Comment by autoexec 4 days ago
You'd never hear about if it the cloud service was taking your corporate data and using it for countless other things though. Insider trading, selling leads to your competitors, pushing ads to your clients/customers with offers of their own products, using trends in what your company and others are doing to guide development of their own products and services. Done carefully, you'd never be able to trace any of it back to them and it'd take a whistleblower for you to find out it happened in the first place.
I don't imagine any cloud service that has massive amounts of data which could easily make them a ton of money is just ignoring it.
Comment by robrorcroptrer 3 days ago
I doubt they want to risk their reputation, trust and agreements to make money in the ways you describe.
Comment by autoexec 3 days ago
Comment by amazingamazing 4 days ago
Comment by dingaling 3 days ago
Depreciation is designed to _encourage_ purchasing of useful local tools, by incrementally matching fractions of the cost of the tool to the revenue it generates over its useful life. The fact that a graphics card might have a book value of $0 after five years of depreciation is a feature, not a bug.
Since the invention of corporation tax it has also had the benefit of offsetting tax over the same period, instead of just one big offset in the first year.
Comment by Gareth321 3 days ago
I'm not challenging the concept of depreciation. It's a necessary tax function. I'm explaining the business case for local LLMs is poor.
Comment by archagon 4 days ago
Comment by cmiles8 4 days ago
OpenAI and Anthropic are fighting to win a race (build the biggest baddest model) that has no prize. The prize is mass adoption at scale at the best price, which is why companies are rapidly shifting to open model. They don’t need to pay 10x for a model that’s provides no practical additional benefit.
Comment by justinhj 4 days ago
Comment by TomGarden 4 days ago
Comment by moate 4 days ago
Comment by tancop 3 days ago
The important part is if they break the contract or raise their prices you can always move to a different provider. You get lower cost and lower risk at the same time which is extremely rare in business. That's just not possible for closed models where your only options are the official branded API or Azure/Bedrock.
Comment by wolvoleo 3 days ago
It's the same with the subscriptions. Local models can't compete because they're simply giving too much value for money. They're effectively subsidised by Big AI. Again something that won't last.
Comment by subarctic 4 days ago
> As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less.
I would argue that the reason they cost so little is because anyone can run open models and offer them as a service, so there's actual competition and the price is closer to cost. i.e. if the open models were just as intelligent as frontier models but cost the same to run as they do right now, the price wouldn't be higher (unless demand went up so high that marginal cost to provide more of the service went up, due to scarcity of hardware and or electricicy).
On the other hand, if what you're saying is the frontier labs have some pricing power due to their models being better, and that is the reason they are able to charge more than the companies providing open models as a service, then I would agree.
Comment by jjav 3 days ago
I'll grant they are superior at least right now. But also, they are too expensive.
We ($work) are finding that it is best to build engineering discipline around AI usage (who would've thought!) and use the cheaper models like Cursor Composer.
Using Opus we can blow through an entire month budget in an afternoon, so while more powerful, it is no longer practical except for rare very complex tasks.
Comment by alxfoster 3 days ago
Comment by xoa 3 days ago
One of the basic questions/concerns here though is that it's not like the AI places are getting the GPUs for 10x less. It's true they have some economies of scale, but they also have some waste, and frankly in this particular case it's not clear they get that much gain over what a lot of businesses could achieve. The biggest traditional gain for central providers is that a lot of typical computing usage is burst-y, and in turn local kit might be underutilized. But with LLMs heavy users tend to use them all the time assuming their tokens allow it (and in the case of local hardware there's nothing stopping you, quite the contrary), they can use it directly interactively or leave them to go overnight on something too.
So it's reasonable to suspect that the reason subscriptions are only a fraction of the cost is that we're in a bubble seeing these companies losing money in an attempt to gain some sort of durable advantage. Just as every previous time, there is the chance that the music stops at some point, and they need to crank up pricing or pull other schemes to actually make money. Of course, it can be a good deal in the mean time, you basically get to suck down investor money for nothing, but it's also not unreasonable to at least be consider fallbacks. Even beyond questions of control and risk etc. I know at least a few places that are now genuinely considering questions like "what happens if a datacenter we depend on gets droned" that would have never had an iota of thought devoted to them even 5 years ago.
>On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places.
I don't think that's "surprising" at all, everyone knows about power use. And this seems like it gets heavily into what you're defining as "usable" and is also more useful to define in terms of cost-per-employee vs total. Obviously a bigger business will have a higher line number total even if the cost per employee is identical, but simultaneously can be expected to be making more revenue to pay for it.
If we're defining an average of a dedicated 5090 pulling 1 kW for every single employee (presumably some people wouldn't use it all the time, but others would then pull the compute for other work), running 24/7 (to cover people running stuff when they're away), then that'd be 8760 kWh per year. At my not particularly cheap New England location that'd be about $1900 per employee per year at the generalized residential rate (~$0.22/kWh), or $156 per month. That doesn't seem radical if it really does boost productivity. However, there is a lot of room to go lower. I'd expect a business to run backup anyway, and these days there are a lot of incentives to do that at least partially with batteries. That also opens up rate shifting as another way to pay back the cost. If we change to time of day pricing, that's 8 hours of peak pricing with the rest off-peak. 8 kWh of battery can now be had for a few thousand. And the off-peak rate is only ~$0.14/kWh, cutting the cost per year by about $700 to $1200 per employee per year. Solar power is also usually far more valuable to use yourself then sell back to the grid, and also continues to plummet in price.
None of this is to say that it makes sense for every place at all, but it's close enough to the the line that the math is at least worth exploring, or could at least lower the cost enough to be worth it given other things. It really comes down to how much extra value the company (or individual) expects to come out of it per month.
>In California or Europe this could be $30-60k per year.
Dunno about Europe, but at the kinda prices I see for California I'm really surprised more places aren't trying to move a lot of usage to battery+renewable.
>The only real moat that local LLMs have right now is privacy.
I don't think resiliency and control are things that can be taken for granted anymore, particularly on the global scale. War and terrorism is getting worse again. International relations are getting nastier, and governments have the power to just order places cut off. If LLMs aren't particularly valuable to a business, then why an expensive subscription? But if they are particularly valuable, then insurance is something leadership should be contemplating.
Comment by epistasis 4 days ago
But Claude also makes it really hard to do that, so what am I even really paying for? Time to extract all my data, put it into a sqlite with FTS5 and make sure I never rely on the overly-opinionated, low-thinking PMs from these giant orgs again.
Of course, that "easy" step has lots of partial solutions like CTK (Conversation Toolkit) or MyChatArchive and I haven't found the perfect one yet, ideally it'd be something that dumped everything into Obsidian or an Obsidian-alike, but surely somebody is working on that? I'd pay $5/month for somebody to solve that problem for me, as long as I still owned the data...
Comment by jjav 3 days ago
Don't rely on chat history. Have it write and maintain summary files that you can import into different sessions, at least for anything important.
Comment by unrented7977 3 days ago
Comment by chrisweekly 4 days ago
Comment by epistasis 4 days ago
What I'm doing instead: syncing coding agent sessions to a central backup location, and for cloud LLM chat providers I'm occasionally exporting data.
Still need something to automate that syncing, and provide search and viewing.
Comment by rpastuszak 4 days ago
I’d try:
1 exporting my data (I imagine it’s common outside of GDPR?)
2 asking Claude to convert it to an easily digestible format :)
Comment by ownsearch 8 hours ago
Comment by piva00 4 days ago
Over time I can imagine us becoming mostly open models on our deployments when hardware is more accessible and the need for expensive frontier models is constrained to very few use-cases that might demand their capabilities.
Comment by Karrot_Kream 4 days ago
Comment by piva00 4 days ago
Just too many products to exhaustively list, it's a big tech company, we're trialing a lot under the sun to find workflows, tools, integrations, including a lot of bespoke internal research, there's a large ML department since almost the inception of the company.
Comment by Karrot_Kream 4 days ago
Comment by r_lee 4 days ago
Comment by calebkaiser 4 days ago
Maybe there is some threshold where the price/quality math for your standard business tips in favor of smaller models and self-hosting the entire stack. I'd certainly love that.
Comment by nutjob2 4 days ago
It's not about money, it's about control. Companies have lots of money and want control over their key technology.
Comment by calebkaiser 4 days ago
The same dynamics that define the public cloud ecosystem are at play here. What AWS sells you is access to appropriate hardware and turnkey infra for your needs. Looking at the cloud industry over the last 20 years, I find it hard to believe that it is impossible to build a moat or a huge business around this.
Comment by ijidak 4 days ago
The problem is serving is a skill readily mastered by the hyperscalers. That's their MO.
All they need is weights to serve. And the open models provide that.
OpenAI is relatively well placed in that they have inference chips they've designed and they own compute.
Comment by jopsen 4 days ago
Comment by epistasis 4 days ago
Comment by r_lee 4 days ago
Comment by epistasis 3 days ago
But the "open" models are right there too, and I'll be using them more in the next weeks as I get better chat interfaces.
Comment by bluefirebrand 4 days ago
Comment by 0xbadcafebee 4 days ago
Comment by lenerdenator 4 days ago
They were set up as a public benefit company and their returns were capped at 100x. They (well, Sam) went out of their way to put themselves in this death march to the IPO. If they had just done what Mark Zuckerberg did with Muse, they're not in this position.
Do you know how bad you have to be at the tech business to make Mark Zuckerberg look like a prudent-yet-visionary leader?
Comment by alexashka 4 days ago
You're giving him too much credit if you think he figured anything on his own - this is the guy who thought 'metaverse' was a good idea.
Comment by TitaRusell 4 days ago
Comment by lenerdenator 8 hours ago
The job of a CEO is to have a vision that is compelling for the market and an ability to execute upon it, thus the "executive" in "Chief Executive Officer". Zuckerberg doesn't have this. What he does have is an absolute majority of voting control over Meta's shares, so his actual ability to do his job doesn't matter.
It's almost as if creating incentivizing systems that don't hold people to account is a bad idea.
Comment by lenerdenator 4 days ago
Even that guy could read the situation.
Comment by pianopatrick 4 days ago
Comment by nutjob2 4 days ago
I think there will be separate and huge markets for models, hardware and compute. That will maximize competition and innovation.
Why? Because even Blind Freddy can see the huge usefulness and power of these (and future non-LLM) models and no-one in their right mind is interested in becoming OpenAI's or Anthropic's bitch. Those companies have tickets on themselves.
Given the recent behavior of tech companies and the US administration, no one trusts either anymore.
Comment by apefulsin 4 days ago
Comment by Mkengin 3 days ago
https://cloud.google.com/blog/topics/hybrid-cloud/gemini-is-...
Comment by phoghed 4 days ago
Comment by SSLy 4 days ago
Comment by motbus3 4 days ago
Comment by nxobject 4 days ago
Comment by fittingopposite 3 days ago
Comment by siruncledrew 2 days ago
Comment by dzonga 3 days ago
Comment by genxy 3 days ago
Having trained on your own chips, that is the impressive part.
Comment by g8oz 3 days ago
Comment by __rito__ 3 days ago
Comment by alfiedotwtf 4 days ago
Think about how crazy it is for non-US companies to use American AI providers - their marketing boasts that you can treat their models as co-workers, assign tasks, invite them to slack annd video calls, etc. Taken at face value, would you hire someone remote who lived in a country that commonly does random shit like deciding whether or not remote workers aren’t allowed to go to work?
Comment by disdegeneration 4 days ago
Comment by throwitaway222 4 days ago
Comment by vohk 4 days ago
Not everybody is coding or doing work that lends itself to burning tokens for warmth. Reuters for example seems to be more interested in using it for research, editing and formatting citations and the like. There's only so much of that work that needs doing, it doesn't always need to be real-time, and they probably don't see it scaling exponentially. They also need to be very aware and in control of their model's biases, or they risk it compromising their work output.
It's widely expected that all of the major providers will need to - and surely want to - drastically raise prices to justify the ludicrous amount of capital they're burning. Multiple companies have already talked about how their AI costs have exploded, and from what I understand that scale of enterprise is paying API rates. I would be disappointed if big business wasn't having a think about what that liability could look like. It's one thing to be reliant on a relatively "stable" vendor like Microsoft for Windows and Office, another to get AWS sticker shock, and then this is promising to be an order of magnitude worse.
Then just plain trust. What if ChatGPT starts recommending your competitors products, or the USA bars export of Anthropic's latest model (again, but for real this time), or they stop serving a model your business now depends on, and so on... That's a lot of risk to leave outside of your control.
Comment by usrnm 4 days ago
Comment by WarmWash 4 days ago
Keep in mind that a vanishingly small number of workers are SWE's churning millions of tokens daily.
Comment by throwitaway222 4 days ago
Comment by kittikitti 4 days ago
We've all seen the office spaces where there's 200 empty computers on a floor. Combined, it's something like 500 cores at ~3 Ghz each and around 3 TB of RAM. The networking is already there and software like exo already exists.
Comment by gizajob 4 days ago
Comment by mancerayder 3 days ago
Comment by gizajob 3 days ago
Comment by sam345 4 days ago
Comment by intrasight 4 days ago
Comment by cmiles8 4 days ago
I couldn’t care less which company drilled for the oil… it’s all the same to me. Models are increasingly no different.
OpenAI and Anthropic are a gas station saying “buy our gas for 10x the price!” When the world is looking at them saying it’s just gas, we’ll take the cheaper brand. We’ve tested your gas and it’s really no better than the stuff that’s 1/10th the price.
Thats why their present business plan is screwed.
Comment by blmarket 4 days ago
I agree 90% of the world can work with 87 gas, but there's always niche/luxury market where 93 can make small difference.
(edit: typo)
Comment by cmiles8 4 days ago
The crazy setup here is that even with that fraction of the pie these companies might be worth say $100 billion optimistically, which would be amazing in normal times. Problem is it’s a train wreck for their investors and the associated debt bubble if they can’t sustain a valuation of 1-2 trillion and the present setup does not put them on a course to that trajectory.
Comment by thereitgoes456 4 days ago
Comment by kelvinjps10 4 days ago
Comment by euroderf 4 days ago
Comment by cmiles8 4 days ago
When your competition has a tiny cost base compared to yours and lacks the bonkers future capital commits you made then that’s a terrible position to be in… hence their conundrum.
Comment by hungryhobbit 4 days ago
Take Open AI for instance: it has "zero debt" ... and $665 billion to $1.4 trillion in "long-term commitments".
Comment by jimbokun 4 days ago
For LLMs the costs of training and inference are a very significant part of the overall costs.
Comment by techpression 4 days ago
Comment by gmadsen 4 days ago
Comment by theseamusjames 4 days ago
Comment by intrasight 4 days ago
Comment by cmiles8 4 days ago
1. People buy Apple because of the broader ecosystem of products and the “it just works” aspect of that ecosystem. Other companies make phones with features that are objectively better but folks don’t switch because the Apple ecosystem is sticky.
Despite trying, neither OpenAI nor Anthropic has managed to move up the stack beyond “hey guys new model release today!” announcements that everyone yawns at.
2. Switching costs are real. It’s a PITA to switch not just the phone but everything else. Switching model providers is a line of code and takes almost no effort.
Apple has a true moat which is why they can command a premium. OpenAI and Anthropic have no moat which is why they’re in trouble.
Comment by pornel 4 days ago
Lots of people use old iPhones and don't care about some "up to" benchmark bumped every year, but are stuck with iMessage contacts, their stuff in iCloud, Apple Watch or apps that are not allowed by Apple to even mention they have Android versions.
Comment by cmiles8 4 days ago
It’s literally the least stickiest thing in the history of tech. Which is a big problem for these companies.
Comment by aff-vasileva 4 days ago
Comment by cmiles8 4 days ago
Comment by overfeed 4 days ago
Comment by Finnucane 4 days ago
Comment by makapuf 4 days ago
Comment by intrasight 4 days ago
Comment by transdev12 4 days ago
Source? The proliferation of labs building competent models would seem to suggest the opposite.
Comment by intrasight 3 days ago
"To build or purchase the physical hardware required to store tens of petabytes of data and train a State-of-the-Art (SOTA) frontier AI model, you are looking at a capital expenditure (CapEx) ranging from $320 million to well over $1 billion."
Comment by simianwords 4 days ago
> Unless they both dramatically slash prices then they’re in big trouble
False, they have already done so many times.
> Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any hope at a successful IPO.
False, margins are higher and I can have a formal bet that prices will go lower.
> However the cold reality for both is that there is zero moat to a model anymore
False, LLMs are not fungible and there exists a natural moat. I like the behaviour of Fable, not the behaviour of Opus - the fact that many people speak about this is evidence.
Comment by fatcatsbestcats 4 days ago
Comment by phoghed 4 days ago
Not a tech company though, so maybe that differs and we’re just behind what’s en vogue. However we did get in on everything pretty early, building our ChatGPT RAG clone right about when azure got gpt-3.5-turbo on api
Comment by simianwords 4 days ago
Comment by techpression 4 days ago
Comment by senordevnyc 4 days ago
Comment by techpression 4 days ago
Comment by senordevnyc 4 days ago
Comment by cmiles8 4 days ago
We’re heading into corporate budget season for 2027 when all this is coming under a huge microscope in boardroom after boardroom across the country at a terrible time for companies trying to IPO.
Comment by khriss 4 days ago
I'm wondering how much of that is the harness vs the model. Overall, the 'feel' of a model seems to be largely due to the harness than the model itself.
Comment by GrinningFool 4 days ago
Comment by saghm 4 days ago
Is this a typo? Have you actually made a formal bet on a prediction market or something to put your money where your mouth is, or are you just saying that you could? There's a lot of things I could plausibly make bets on, but that doesn't mean that they're likely to happen.
Comment by simianwords 4 days ago
Comment by saghm 4 days ago
Comment by chrinic630 4 days ago
Comment by simianwords 4 days ago
Comment by throwitaway222 4 days ago
Classism is the new racism.
Comment by 23asgh1 4 days ago
Comment by harrymunro 4 days ago
Comment by novok 4 days ago
If the companies survive the next few years, which they probably will because they represent too much of US economic growth to allow them to fail, this gap will keep on expanding.
Starting from zero without distillation is a lot harder, a lot more expensive and a lot more work. OSS models is what a laggard does to get adoption. China's gov't might keep on sponsoring it as a counter GPU embargo thing, but when gov't get involved, usually the other side gets involved too.
As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors.
Comment by nvme0n1p1 4 days ago
One way or another they're getting the same results as proprietary labs, with a fraction of the hardware for a fraction of the cost. OpenAI can't keep raising funding rounds of $100billions to subsidize their compute costs and get results by brute forcing parameter count. And if they're having trouble keeping up with Chinese labs' efficiency, maybe they should stop worrying about distilling and instead hire some of the smart people responsible.
Comment by novok 4 days ago
To actually have something competitive and improved within the next 3 months and not be perpetually behind, you need your own independent model creation process. So to extend the metaphor, a complete movie studio with cameras, actors, staff, sets, budgets, etc. It's the right strategic move to do when you are GPU constrained, which the Chinese labs are, but it won't let you get past it.
A bunch of pedantic people will come out of the wood work citing a bunch of things saying that is not the case because of some detailed mechanics of how model training works and they will get fixated on some of the words I used, but zoom out to the level of what an AI lab is able to produce and this becomes evident.
Comment by nvme0n1p1 4 days ago
Per the article, companies are dropping OpenAI+Anthropic (partly) because of costs. If distilling is so simple and easy, why doesn't OpenAI take this "quick shortcut" and serve a self-distilled model externally, so they can charge reasonable prices and stop bleeding customers? Surely they can at least match the Chinese labs' efficiency, right? Wouldn't more customers and less opex look good for the IPO?
Comment by linkregister 4 days ago
Comment by nvme0n1p1 4 days ago
Comment by linkregister 3 days ago
Comment by nvme0n1p1 3 days ago
The open weight labs figured out some secret sauce that (so far) big name labs are unable to replicate, so instead of competing, they're going on the defensive with claims of distillation attacks.
Comment by linkregister 4 days ago
Comment by nvme0n1p1 4 days ago
Comment by cmiles8 4 days ago
They’re burning cash like there’s no tomorrow. They desperately need to show that they have a real business and not just a giant burning pile of cash doing academically interesting things. If they could simply sell models that are 95% as good at 1/10th the price they’d do that. They’re losing the enterprise sector because they’ve not done that.
Comment by linkregister 4 days ago
It certainly is possible that Anthropic and OpenAI are doomed to bankruptcy, but without a known drop in revenue it seems exceptionally confident to make such a certain claim.
Comment by cmiles8 4 days ago
Comment by linkregister 3 days ago
Comment by chvid 4 days ago
Chinese AI labs do well because they have the best AI people coming out of a huge talent pool.
Comment by mlboss 4 days ago
Comment by Gareth321 4 days ago
[There are many examples of distillation attacks.](https://cyberpress.org/anthropic-claude-ai-distillation-atta...) Of course, you could argue this is just a form of "learning" from other intelligent systems, and this is arguably what LLMs have been doing from the beginning. So I don't begrudge the Chinese labs doing it - they all do it. But let us not pretend that it's not happening.
Comment by janalsncm 4 days ago
1. Competitors to OpenAI and Anthropic are good because of distillation.
2. OpenAI and Anthropic will come up with some methods for preventing distillation in the future.
Both of these are dubious imo. For RLVR tasks like coding in particular, you definitely don’t need continuous distillation to improve, otherwise OpenAI and Anthropic themselves would not be able to improve because there is no better model to distill from.
Comment by achrono 4 days ago
Think about it: even if they distill the shit out of frontier models, the model still gotta learn, right?
If anything, as you can see from the K2 Horizon release, aggressive (self-proclaimed) reliance on distillation does not result in a model that has remotely any frontier capability. Try asking K2 Horizon to write iambic pentameter for instance, or even give it the car wash prompt. I tried both these on the Q8 quant for the 7B model and the results were depressing.
Comment by achrono 4 days ago
Comment by novok 4 days ago
> As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors.
To help you further understand, the Chinese labs will never admit they were distilling until they are better than US labs with their own non-distilling process or some sort of espionage-like public reveal shows it. So you need to look at secondary indicators. Much like how the chinese government lies about their economic stats so 3rd parties use secondary indicators to figure it out.
Comment by achrono 4 days ago
Nathan Lambert has made this distinction explicitly: he thinks Chinese labs likely innovate heavily on distillation, while saying he wouldn't call it a crucial factor in their post-training capabilities, in part because the RL work still has to be done by the lab itself -- just like what I said above.
https://www.interconnects.ai/p/how-much-does-distillation-re...
Comment by upcoming-sesame 4 days ago
Comment by krupan 4 days ago
Comment by PhilippGille 3 days ago
Isn't that a bit overgeneralized?
There's more than weights for the Olmo models for example: https://allenai.org/olmo
Similar for Nvidia's Nemotron models IIRC.
Artificial Analysis has an "openness" ranking: [1]
[1] https://artificialanalysis.ai/models?model-filters=open-sour...
Comment by hashstring 3 days ago
Comment by ms_menardi 3 days ago
Comment by krupan 2 hours ago
Comment by NullPrefix 3 days ago
Comment by ms_menardi 2 days ago
proprietary software is data and code, and while it's true that code is data, the difference is that some data is intended to be for control flow, while in a model it is all the same stuff.
You can point to any single part of an LLM's model and say "this here is a weight" but if you point to any single part of a binary file you will have no idea what you're looking at.
And yes, you could decompile the binary, but that still doesn't give you the entire picture, and all proprietary software does the shady shit on servers these days anyway so you're not even going to find anything interesting.
Comment by 8note 1 day ago
what if i want to change out the middle part of the training, and then still run the rest as before? or if i want to cull a bunch of the initial training set?
a binary is also just numbers, but we know there's other parts to it
Comment by syntaxing 4 days ago
Comment by majorchord 4 days ago
Comment by epistasis 4 days ago
The most insidious advertising in the world is about to be surfaced as people use LLMs to look for product recommendations.
Comment by WarmWash 4 days ago
Comment by epistasis 4 days ago
You are very confused about the security and what can happen here.
Comment by WarmWash 4 days ago
American local models don't hand over data either, which would nullify the point of the comment.
Comment by g8oz 3 days ago
Comment by 8note 1 day ago
Comment by majorchord 4 days ago
Comment by epistasis 4 days ago
Second, the "ALL" qualification is extremely wrong, as only small fractions of these products in the US could ever be sourced to the human rights violations cited here.
Comment by Kuyawa 4 days ago
Comment by cousinbryce 4 days ago
Comment by majorchord 4 days ago
Comment by SamInTheShell 4 days ago
Comment by Henchman21 4 days ago
Because they've been trained to think "cloud-first" for a decade?
Comment by phoghed 4 days ago
Comment by r_lee 4 days ago
but if there's roughly Sonnet 4.6 level capable open small models, then I'd be impressed
Comment by skybrian 3 days ago
Comment by spopejoy 2 days ago
Comment by petcat 4 days ago
This makes sense since corporations require legal certainty, and using an open model from an American company (probably) provides them some level of indemnity, and also someone to sue.
Comment by Waterluvian 4 days ago
Comment by petcat 4 days ago
Counterparty risk is a lot more straight-forward to evaluate when dealing entirely within the US, with US companies.
Comment by honr 4 days ago
But if you let LLMs talk to people (customers, for example) directly, then yes, you need an LLM provider that you can hold responsible.
Comment by rolisz 4 days ago
Comment by Phemist 4 days ago
The link is annoying enough to find that I can imagine "Mea Culpa" being an effective enough strategy for businesses moving into the ML/AI field, changing their tune after they get caught, but matured their own software to stand on its own feet.
Comment by unrented7977 3 days ago
Why play ball with a hostile government when you can host your own frontier models?
Comment by RajuChacha108 4 days ago
Comment by _the_inflator 4 days ago
The thing is that needs more attention is reverse engineered a LLM which is highly fascinating. I tried it, but it seems I am not there yet to put it mildly. It requires serious effort.
I am just speculating but can LLMs be sleepers? You write software and it seeds traces here and there under certain conditions that pose a serious security risk.
Or a kill switch?
I don’t know. I distrust Chinese LLMs but even more due to training data.
It is after all not a Western model. Different biases and the might be subtle but nevertheless substantial.
In short: no open source LLM may be usable without additional Finetuning for certain valid use cases.
The real value is versioning and autonomy as well as lot more stable answering despite model rot.
Also testing and the supporting systems are easier to maintain.
It is mainly an infrastructure challenge.
Comment by kakacik 4 days ago
Comment by majorchord 4 days ago
The overwhelmingly vast majority of open-source code isn't actually looked at or audited. Yes it's there for all to see, but that doesn't mean it's doing any good at the moment, in this context.
Comment by manvillej 3 days ago
this isn't if-else statements, its a jumble of linear algebra and matrices.
I watched a demonstration where one AI was trained to be obsessed with penguins. they asked for a random set of numbers from it. They fed that set into another AI model to analyze and the new model started to become obsessed with penguins.
I dont think open source is the only one to worry about though. I don't know if we even have a way to guarantee any model is secure, open or closed.
Comment by SamInTheShell 4 days ago
Comment by petcat 4 days ago
Comment by slowin 4 days ago
I'm sure this will change (and I can't wait for it!) but as of today, open models might be fine for summarizing and writing docs, but you need SOTA to work on code if you want to be competitive.
Comment by nemomarx 4 days ago
Comment by tomashubelbauer 4 days ago
Comment by kbwal7 4 days ago
Comment by redox99 4 days ago
Comment by happycube 3 days ago
Comment by water-drummer 4 days ago
Comment by horsawlarway 4 days ago
I also work in software, and while I vaguely disagree that open models can't be used (they absolutely fit into productive niches here, and holy hell are the last generation [ex laguna s1, kimi k3, glm 5.3, etc] actually decent) - I will agree that SOTA are a better fit for software development, especially when used in conjunction with an already very expensive employee who's driving them.
But for "Corporate America"... no. You absolutely don't need SOTA. They're doing things like transcription, summarization, customer interaction, minor technical tasks like form creation in existing tools, report generation (ex - powerpoint, pdf, docs, etc) and other general "white collar tasks". Think about roles in business that are in the 60-85k compensation range.
It's mostly busy work that keeps existing processes flowing and the business on the rails. Important, but not research/novel.
And cheap ai... is a wonderful fit for a lot of this. No one wants to replace an employee making 80k with a less reliable AI that costs 45k a year in tokens (SOTA). But they're absolutely willing to drop 2-3k/year on AI (~100/month - right in the open model cost range) for that employee if they can get a 10% bump in productivity or happiness.
Comment by patja 4 days ago
Seems like if you ask 5 different people what "real coding" means you might get 5 different answers.
Not everyone is building the next framework or compiler.
Self-hosted Qwen 3.8 @Q4 on my RTX 3090 can produce beautiful functional CRUD pages and apps all day long. And that is 90% of the "real coding" being done in corporate settings.
The quote in the article about Mazda vs. Maserati captures this. Many might want the Maserati and drool over its specs and capabilities, but balk at the cost and how often are they really going to run it up to full performance limits on their daily commute to their cubicle?
Comment by ThrowawayR2 4 days ago
Comment by slowin 4 days ago
Comment by tinyplanets 4 days ago
Comment by spopejoy 2 days ago
Comment by transdev12 4 days ago
Comment by taf2 4 days ago
Comment by MisterMunchkin 3 days ago
Comment by taf2 3 days ago
Comment by simonw 4 days ago
> By May, open models accounted for 20 percent of AT&T’s A.I. use. That has since risen to 40 percent and may jump to 60 percent in the coming months, Mr. Markus said in an interview.
This is missing a crucial detail. We know they "help with customer service, call transcription and coding", but which of those have been upgrade to open models?
Call transcription is trivial to do with open models. I can run Whisper or Parakeet on a low-spec laptop.
"Customer service" could mean a lot of things, but it sounds feasible for open models too.
"Coding" - they might go to open models for that, but I expect the costs involved in paying for closed models for software developers within AT&T are a fraction of the costs involved in transcribing all of their calls or handling aspects of custom service for millions of customers.
From later in the story:
> AT&T researches Chinese models but is not using them, Mr. Markus said. Instead, it is working with popular alternatives made by American companies such as the Gemma A.I. model from Google and the Llama A.I. models from Meta.
Gemma 4 is great, but really, Llama, in 2026?
Comment by mewse-hn 4 days ago
I'd assume the author is just getting confused because of ollama and llama.cpp and all the other ecosystem "llama" that are still in use. Llama really did kick off the open models thing
Comment by _doctor_love 4 days ago
Adoption of open-source models to my mind is a similar step in that direction. In all cases, the goal is to become untethered from a mercurial vendor.
Comment by honr 4 days ago
Comment by hparadiz 4 days ago
Comment by wnmurphy 4 days ago
https://chatjimmy.ai/ blew my mind at how fast etched model weights can be.
For on-device LLMs, there's a point of diminishing returns, meaning you don't need to have the latest frontier model for most operations.
Comment by hparadiz 4 days ago
Comment by andriy_koval 4 days ago
Comment by falaki 4 days ago
Comment by golem14 4 days ago
Corporal, put him away.
Comment by overfeed 4 days ago
Comment by unrented7977 3 days ago
Comment by spopejoy 2 days ago
Comment by kittikitti 4 days ago
I highly recommend utilizing an open sourced embedding model instead of paying for a closed source one. It's vastly more reasonable to run an open sourced embedding model as a first step. They're much, much smaller and, due to the overhead of network latency, and running it locally has almost the same speed as through an API even on slow computers.
I would even go so far as to say that closed source embedding models have a high risk of data hostage. If a team doesn't have access to the embedding model, the embeddings become useless. A corporation like OpenAI could, say, hike the prices to that model by 1000x and everyone would have to pay up or forfeit any utility of the data.
I envision a future where open source embedding models are shipped with relevant technologies and implemented by currently under-utilized chips like NPU's. A startup developing cheap microprocessors that can run them is an idea I would pay cash for. Or perhaps they will be bundled with security tokens.
While it might be impractical for all corporate teams to run language models, it is very realistic for everyone to operate an open sourced embedding model, at least in their private cloud. Better yet, utilize transfer learning on an open sourced one to train your own, that way the embedding vector is more secure against competitors and trade secrets.
Comment by Zambyte 4 days ago
Comment by manyatoms 4 days ago
Comment by overfeed 4 days ago
Who's going to be the new Bill Gates, with a vision for "a GPU cluster in every home?"
Comment by anon373839 3 days ago
Comment by bfrog 4 days ago
I fully await my ai in a usb box. The models are plateauing and some clever company is secretly working on this already I’m sure of it.
Imagine baking in a model weight set in ROM with an analog computer doing what otherwise takes way too much power in digital form.
Comment by mudil 4 days ago
Comment by chasd00 4 days ago
Comment by cmiles8 4 days ago
Besides, in their present rather dire financial state there isn’t much to sue these companies for anyway cash wise. NYTimes is suing on IP grounds.
Comment by chasd00 4 days ago
Comment by therealdrag0 4 days ago
Comment by nozzlegear 3 days ago
Comment by verytrivial 4 days ago
Comment by motbus3 4 days ago
Comment by overfeed 4 days ago
Comment by happycube 3 days ago
Comment by iainctduncan 4 days ago
In my experience as a tech diligence assessor for PE firms for the last 7 years, investors really, really don't like companies being beholded to single entities that they don't control. Anthropic and OpenAI have demonstrated that they are not trustworthy, or predicatable, or finanically safe, or even capable of hitting three fucking nines. Investors know they need companies to be on the AI train, but they really don't like vendor lockin to the big AI companies. Every diligence I get asked "how easily can they change models?"
I think when open models reach 80% or 90% capability (or maybe even less!) a whole lot of companies are going to say "almost as good with way less risk is a better deal".
Comment by MrDresden 4 days ago
Has there been a sea change in how investors view these things in general, or is it only AI?
Comment by iainctduncan 3 days ago
On the other hand, nobody trusts the big AI companies not to pull shenanigans or dramatically raise prices... or even be in business in five years.
Comment by spopejoy 2 days ago
Comment by ozozozd 3 days ago
But in terms of reliability - uptime, product, legal - they are in a different league.
There were exactly 0 instances waking up to a product decision at AWS completely breaking your product or workflows.
Comment by solid_fuel 4 days ago
Comment by schopra909 4 days ago
This feels reminiscent of the big push to RAG a few years ago. And, more broadly the skunkworks projects that big companies tout in the press before they end up killing, when the operational overhead becomes too much for their liking.
Ultimately, the narrative is good for the consumer and the enterprise. It’ll mean OpenAI and anthropic will have to keep prices low. But ultimately, in the course of the next 10 years, I don’t see enterprises wanting to do this themselves. It’ll just be simpler (and eventually safer in their eyes) to send traffic to the big labs.
Comment by CodingJeebus 4 days ago
A) the uptime matrix between Github, OpenAI, and Anthropic means that we've faced multiple entire days of not being able to ship org-wide due to our reliance on automated code review and other tooling. Every cloud service baked into our CI pipeline becomes a point of failure. We can't live without AI anymore, but it too often either directly or indirectly gets interferes with our ability to ship, and I don't see this improving any time soon.
B) Locally hosted AI has serious advantages with regard to PII/sensitive data management, and there's not much the frontier models can do to overcome this. There are so many things I want to build and let loose in a sensitive data environment but can't due to data governance around frontier models.
C) Anthropic and OpenAI cannot keep prices low forever. They're still burning insane amounts of cash and at some point, they're going to have to transition from growth mode to profit mode. They're already juicing their sales pipelines to the max with introductory pricing and other things to get people in the door. But those are all short-term online marketing plays.
Comment by schopra909 4 days ago
I think ultimately the folks with the purses won’t care enough about a for it to be taken seriously, even if it’s an engineering bottleneck.
B definitely has scope but still smaller than I’d expect. When I was an intern at yelp, I was migrating us off internal credit card management to Braintree/stripe. No one would have imagined outsourcing that in early 2000s. There’s a long tail of stuff that you can’t sent to a 3rd party; but for most use cases it’ll suffice.
For c, true; but this is expensive for everyone, including the Chinese model companies that are trying to undercut Open Ai / anthropic. In the nth degree, i think the field will bring the cost down to the place where it’s manageable (see the existence of the cheap Chinese models). The real question is the $$ spent on pushing the research forward at scale.
Comment by lvl155 4 days ago
Talk to “AI” executives at large firms and 95% of them are clueless sales types that crawled their way to the top. Then again, it is basically a repeat of IBM, Microsoft, Oracle, etc. Same dumb executives making decision to not get fired and enjoy their place at corp.
Comment by DEF14A 4 days ago
Comment by c7b 4 days ago
Could someone clarify whether there are actual open source models that are competitive with the likes of Gemma (mentioned in the article), or is the headline just wrong?
Comment by Balgair 4 days ago
I use opensource models at work because my work is too cheap to spring for a $20/mo account for me. Since HuggingFace models can be run on my laptop now (still very slow though), nothing is leaving the 'secure environment' and so I can actually get work done (instead of the 'old' version of coding and writing - google).
Comment by pllbnk 4 days ago
Comment by Balgair 2 days ago
Comment by _superposition_ 4 days ago
Comment by ctkhn 4 days ago
Comment by therealdrag0 4 days ago
Comment by maxrev17 4 days ago
Comment by ungreased0675 3 days ago
Comment by arbuge 4 days ago
Which makes the Muse 1.3 launch this week particularly interesting, although to get the low cost version you do need to agree to share data with Meta.
Comment by AnotherGoodName 4 days ago
You can’t further train the closed models. The open models can be fine tuned for your company. Big companies fine tune models on all the internal systems and documentation, not just through .md files (you’d blow up the context trying it that way) but actual fine tuning of open weights models. A low tier but open weights model actually beats frontier models when you do this for a specific task.
I think the frontier providers need to have a way to isolate instances (bedrock style?) and allow fine tuning to compete. Big companies are absolutely fine tuning models right now and getting better results than even the best frontier models for their use cases.
Comment by hintymad 4 days ago
Take Anthropic for an example. Anthropic has successfully destroyed customer trust, at least for me. DHH in a recent interview mentioned that Claude refused to translate an article about immigration. Not summarize. Not editorialize. Translate! I think this reveals an unacceptable level of paternalism: Anthropic fundamentally believes that it possesses a moral authority superior to the people actually paying for the API. If such basic and mechanical translation is already too sensitive to touch, the goalposts have moved from safety into outright censorship. What prevents them from quietly deciding tomorrow that your proprietary business logic, financial data, or legal documents cross their invisible moral line?
Let alone how Anthropic treats Cursor and Figma - not that they are wrong as companies are free to compete legally, but nonetheless it shows that companies can't outsource their intelligence to a potential competitor.
Comment by eli 4 days ago
But I doubt this a major factor in the trend. I just don't think it's something most corporate users run into. My understanding is these guardrails are negotiable for enterprise customers anyway.
And, not for nothing, but if I owned a human-powered translation company I would've refused to translate it too.
Comment by biophysboy 4 days ago
Comment by thatmf 4 days ago
While I agree that Claude can be overly paternalistic at times, how should it respond to a request to translate, say, bomb-making instructions? It's reasonable to me that it might refuse this.
Comment by jimbokun 4 days ago
Also, you see zero distinction between hearing opinions on political topics you might find objectionable, and building a bomb to kill people?
Comment by eli 4 days ago
It was an incredibly racist post claiming “gypsies” are like invading wolves and that something more drastic must be done to get rid of them before they kill all the “sheep” in Copenhagen.
Comment by Animats 4 days ago
Comment by Quibblingeek 4 days ago
Nevertheless, it happened. The problem is that businesses demand stability. Building an application, service, or business process around an API that can be turned off whenever a government demands represents an unacceptable risk.
It’s not a mental exercise when it has already happened once before and likely will happen again. If you control the weights and invest in your own hardware to run them (or rent it from a cloud provider), that risk can be mitigated.
Comment by happycube 3 days ago
Comment by dominotw 4 days ago
Comment by MrResearcher 4 days ago
Comment by hparadiz 4 days ago
No one I know uses the built in VSCode extensions anymore. It's all TUIs now. You can use Opencode as a TUI now for local.
Comment by Zambyte 4 days ago
Comment by Quibblingeek 4 days ago
I’ve also had luck with copilot-cli, but find that to be more limiting and it’s only really worth using if you are paying for GitHub Copilot (or your company pays for it, as is my case).
Comment by gflh73 4 days ago
Have you been brought into line? Open source AI also violates copyrights.
Comment by ronsor 4 days ago
Comment by Der_Einzige 4 days ago
This is actual communism, and the fact that Bernie Sanders and every other member of the DSA isn't actively fighting for open source and is often fighting against all AI shows how fake their purported movements are and have always been.
Comment by add-sub-mul-div 4 days ago
Comment by tsimionescu 4 days ago
The fact they prioritize other fights more than OSS, and have a rather dim view of AI, is hardly proof that they are fake.
Comment by roarcher 4 days ago
Open source licenses are only enforceable because of copyright law. How are you going to enforce GPL3 when you have no legal authority to say what people are allowed to do with your code?
Comment by 2948154 4 days ago
Open source authors have always been protective of their copyright. There are numerous examples when drivers have been copied between BSD/Linux (I forget which direction) which led to huge flame wars.
The whole point of the GPL is that it uses copyright and copyright assignment to the FSF to protect what it calls software freedom.
BSD authors are very upset if the attribution clause isn't observed. And so on.
It is communism to exploit poor open source authors? I have to read Marx again.
Comment by fcarraldo 4 days ago
https://jacobin.com/2026/07/ai-nationalization-sanders-liber...
Comment by shimman 4 days ago
Comment by graemep 4 days ago
Comment by Supermancho 4 days ago
Comment by swader999 4 days ago
Comment by subarctic 4 days ago
Comment by biophysboy 4 days ago
It makes me skeptical that the flagship companies are sustainable. Every company is going to maximize “fuel efficiency” to save time and money.
Then again, maybe the cheaper models have more markup for them, in which case they are probably happy w this arrangement. I’d be curious to know how the money making varies by model.
Comment by woah 4 days ago
Comment by overfeed 4 days ago
Comment by biophysboy 4 days ago
Edit: I also have to read the methods anyway for scientific accountability/integrity anyways, so I may as well play that role at the outset.
Comment by mylies43 4 days ago
Plus there is something to say about being in the drivers seat, youll have a much better idea of how it works instead of needing to talk to claude and hope its correct. Since most LLMs also not very good at ideas even in my experience with better models its better to think for 10mins, youll get a much high quality result
Comment by biophysboy 4 days ago
Comment by timterim 4 days ago
Comment by happycube 3 days ago
Comment by optimalsolver 4 days ago
Comment by amelius 4 days ago
Comment by shevy-java 4 days ago
Comment by Kuyawa 4 days ago
The race is still on
Comment by Supermancho 2 days ago
Comment by siliconc0w 4 days ago
Just use Fable 5.1/Opus max for the hardest problems, GPT Sol high as your workhorse, and maybe terra for async batch stuff you don't really care about. Gemini 3.8 High also looks pretty good and is quite fast if you're already a GCP shop. You can basically benefit from open models without using them because they force the frontier models to be cheaper.
Comment by FuriouslyAdrift 4 days ago
We have our own on-premise inference server (quad MI300A) that runs Kimi 2.8 extremely well and we transitioned all heavy work to it since it's basically instantaneous for the whole team. It's a good enough solution and we will hit break even before the end of the year already.
Not everyone needs frontier models and availability is frequently much more important than a lot of companies realize.
Comment by bnchrch 4 days ago
Coding, maybe.
But for operationalized/repeatable tasks it definitely does.
For example I have a workflow that I was running in April that effectively would cost $30k in token spend for each full run.
However now, with GLM 5.3-flash, we've brought the cost down to $7k-9k with our evals showing we've had no loss in recall, precision etc..
Comment by cisrockandroll 4 days ago
Comment by sriniwasx 4 days ago
Comment by natebc 4 days ago
Comment by wewewedxfgdf 4 days ago
Comment by solid_fuel 4 days ago
At least with self-hosted models the people running them get full control and don’t need to worry about the model or guardrails changing under their feet.