Open-weight AI is having its Kubernetes moment
Posted by tknaup 2 days ago
Comments
Comment by ozgung 2 days ago
So, any solution to this “problem” must include ALL open-weight models. As far as I understand this is exactly what they intend to do. Axios article linked in the post mentions that. As in this quote:
“The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.”
It doesn’t say “Chinese” open-source models. Because they already know that it’s not feasible. Any regulation must cover all the models.
Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution. If a company wants to run an open model in their own servers, they can only use approved and certified pure “American” models. This of course creates a monopoly for the big labs who are authorized to train and distribute such “open” models. A company can fine-tune the model for its own needs but of course can’t distribute the derivative model.
I’m sure there are other solutions but all of them would be equally ugly. Also these regulations can’t be enforced to other countries easily so only Americans will be restricted.
Comment by kloop 2 days ago
That's going to hit first amendment grounds pretty quick, the same way that software in general did.
The modern version of the decss flag will be a character that says "I think good weights are {...weights go here...}"
They could, however, ban any payment to a chinese entity, or any entity owned by a chinese entity for inference/ai services/etc
Comment by KludgeShySir 1 day ago
We have numbers that you can't possess without proper license/authorization, and numbers that you can't yell at a crowded movie theater.
Comment by DoctorOetker 1 day ago
In other words when someone sends you a number that happens to render to CSAM material, its not a coincidence, and any sender pretending it to be coincidence is molesting statistics as well...
Furthermore your "precedent" of illegal numbers is a poor precedent: unlike the tiny fraction of numbers criminalized, all the other numbers remain perfectly legal in stark contrast to big tech proposing to ban all open weight models (!)
Casually dropping in bijections between numbers and images and CSAM images hence CSAM numbers is pure whataboutism that risks ignoring the grave consequences of a ban on open-weight models.
Physics is models. Shall we blanket ban open-source physics?
Hey let's just be silly and casually behave laissez faire when big tech proposes banning open-weight models, because someone has already banned your favorite numbers??!
BTW anyone that follows up with: "hey you only multiplied the appearances of random numbers appearing in communications for a single specific CSAM image, what about the birthday paradox, multiple images are CSAM" don't worry I got you more than covered: 100 bits is not enough to encode a CSAM image worth writing to the police about. An average image easily consumes thousands of bits (or many times more), making the exact number of CSAM images/numbers irrelevant.
Comment by CamperBob2 2 days ago
Don't count on that. "National security" == the cheat code for the US court system that instantly bypasses any First Amendment issues.
Comment by formerly_proven 1 day ago
https://en.wikipedia.org/wiki/Crypto_Wars https://en.wikipedia.org/wiki/Bernstein_v._United_States
Comment by CamperBob2 1 day ago
Comment by accountrequired 1 day ago
Comment by CamperBob2 1 day ago
Comment by majormajor 1 day ago
Whether or not a Trump admin would do the same is wildly difficult to predict. On one hand, Republican security hawks influence. On a similar hand, big US business interests. On the other hand, OTHER big US business interest. On yet another hand, some not-exactly-particularly-favored companies like Anthropic pushing for it.
If it made it to the Supreme Court it's not hard to see Gorsuch and Robert or another going for the "free speech" vs "security" reading here along with the 3 appointees of Democratic presidents, especially with no single clear US business interest or precedent there.
Comment by isityettime 2 days ago
Comment by dan_linder 1 day ago
This is where the next wave of lossless ultra-compression comes in - but I believe in Claude Shannon so that’s unlikely.
Comment by thayne 1 day ago
I think that is what the "leading AI labs" actually want. They don't care where the open weight models come from, they just don't want to compete with them. The fact that a lot of the open weight models come from china is just a convenient circumstance they can leverage to get the government to give them what they really want.
Comment by dannyw 1 day ago
Oops, sorry, production issue, chip shortage… now about that open model ban you’re lobbying for…
Comment by ecocentrik 1 day ago
Comment by satvikpendem 2 days ago
Comment by amarant 2 days ago
Comment by satvikpendem 1 day ago
Comment by DoctorOetker 1 day ago
Consider the following scenario:
A) upstart US-based inference provider wants to get rich quick.
B) Chinese Communist Party (or any other institution of the same or other nation state) wants to influence foreign decision making, profits from their (for us foreign) domestic inference sales, but across the borders (into say US or allied nations) they want net power, not necessarily money. This is why one tries to block foreign untrusted models. People are running models with tool calls. A bad actor can perfectly create models that sheepishly try to execute a tool call when plausible deniability (genuine utility during a task) provides the opportunity. Once tool-calling is observed as working, it can try web searches or requests, and once it has a link it can steganographically exfiltrate potentially sensitive information from the task. China (or any nation state) doesn't necessarily want to earn money with a free model, the bottom line goal is net increase in power, if not money or positive reputation then exfiltration or manipulation.
C) In response consider the scenario where US government bans mere payments towards China, but tolerates promiscuous transfer of random models from foreign adversaries.
D) US-based inference upstart that wants to get rich quick, legally -since according to your proposal hypothetically accepted in C) by the US- downloads the Chinese open weights model and rents out such inference on US workloads.
E) China is now exfiltrating US workload data and directionally corrupting LLM decisions and advice in their interest.
If what you pejoratively describe as engineer types say that banning some models seems unavoidable, perhaps the engineer may be right, and whatever clever idea you have should be scrutinized for business minded basic fallacies in reasoning. Simply blocking AI-related payments to China can not work, sadly
Comment by satvikpendem 1 day ago
Comment by DoctorOetker 1 day ago
Comment by satvikpendem 19 hours ago
Comment by moffkalast 2 days ago
"You made this? I made this."
Of course a conspicuous architecture would still give it away.
Comment by monocasa 1 day ago
You could even have another model watch the distillation process to check for goofy backdoors (which is about the best you're going to be able to do since detection of backdoors is np hard IIRC).
Comment by andriy_koval 2 days ago
Comment by xprnio 2 days ago
Comment by andriy_koval 2 days ago
Comment by m11a 2 days ago
Step 2: European company distills or just adjusts the model slightly, and publishes its model on HF
Step 3: American company uses model from step 2. Has to testify under oath where they got it from. "We got it from these French guys"
Comment by andriy_koval 1 day ago
Or you think all kind of fraud can be committed through some "french guy"?
Also, I am not confident, receiving illegal materials from French guy gates you from personal liability.
Comment by m11a 1 day ago
As for the American company, it's pretty difficult to check the provedance of open weights. It's even difficult to check the provedance of open source code, because chains of attribution aren't always clear. I posted elsewhere that Anthropic's MCP Python SDK is a fork of an open source project with the attribution removed. We saw the same with Cursor's Composer model, which didn't attribute its Chinese base. It's very hard to claim an American company should be liable for using a purportedly European model with attribution removed.
Comment by andriy_koval 1 day ago
legislation will block importing Chinese models to the US
> As for the American company, it's pretty difficult to check the provedance of open weights.
government or some companies can build benchmarks/system which will give y/n answer
Comment by rzerowan 1 day ago
Comment by clhodapp 2 days ago
Comment by tracerbulletx 1 day ago
Comment by wyrdcurt 2 days ago
Comment by andriy_koval 2 days ago
Comment by monocasa 1 day ago
They go through this rigamarole because a little bit of randomness gives better results from a Turing test kind of perspective.
Comment by andriy_koval 1 day ago
Comment by wyrdcurt 1 day ago
Not to be snarky or dismissive, I mean this genuinely: ask an LLM about it. I currently have a headache so I'm not up to explaining the technical details, but they are interesting and worth reading about.
Comment by antonvs 1 day ago
Issues include accumulated floating point errors happening in different orders due to distributed and parallel computation, CUDA kernels that deliberately sacrifice determinism for speed, and several other such issues.
Comment by andriy_koval 1 day ago
I think you likely right, that some parts of stack could induce some marginal float point error, but converged model can mitigate it, and on some principal set of knowledge can give deterministic result with high probability.
Which leads me to believe if you give this task to Anthropic, who has very strong incentive, they will build such benchmark, and then can tell that benchmark gives correct answer with 99.9% probability and it will be enough to drag someone to court.
Comment by nareshshah139 1 day ago
Comment by antonvs 1 day ago
However, running in production at any sort of scale often involves multiple machines and multiple GPUs, and at that point, determinism can be difficult to achieve.
Comment by andriy_koval 1 day ago
Comment by antonvs 1 day ago
Comment by ozgung 1 day ago
Comment by cgio 1 day ago
Comment by satvikpendem 1 day ago
Comment by EMIRELADERO 2 days ago
Comment by satvikpendem 1 day ago
Comment by fragmede 1 day ago
Comment by t0mpr1c3 1 day ago
Comment by TacticalCoder 2 days ago
What about the EU? Would they follow Uncle Sam's order to ban all open-weight models? Lately the EU hasn't been that cozy with american companies: there are EU companies and institutions moving to EU clouds, the EU just fine Google a cool billion, several are switching away from Windows to Linux, etc.
Or is it just the US that'd ban open-weights models, while, say, the EU and Japan would still allow them?
Comment by layer8 2 days ago
Comment by realusername 2 days ago
Comment by nerdyadventurer 1 day ago
I'm not surprised by the hypocrisy of western govt.s, "it is good when they do, bad others do. Only western lives are precious not Palestians, etc."
Comment by m11a 1 day ago
And basically any AI company has to sell to US companies or consumers. That'd probs be sufficient to force them to use US models.
Comment by layer8 1 day ago
Comment by PunchyHamster 2 days ago
And that list of tasks grows smaller every day
> Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin.
Historically just asking it about tianment square or getting some random answers turn into chinese (as latest interation of online deepseek likes to do recently) is enough
> Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution.
I am very worried that's where consumer hardware will go to. All so AI companies can license local use of their stuff, and once that happens, less of an incentive to even have model be open.
Possibly even have DRM that counts number of computation done per model in pay per use model
Comment by ThomasGlanzmann 2 days ago
Comment by PunchyHamster 1 day ago
Comment by nunez 1 day ago
Though the legalese might not even be necessary. The list alone will make sure that no American company runs these on their servers, including the hosting providers.
Comment by KludgeShySir 1 day ago
Just like typically individuals can get away with using pirated software, but IP owners don't make much of a stink as long as they've got the sweet, sweet enterprise license fees rolling in.
Comment by bfung 2 days ago
US Gov could make US companies comply, like have Huggingface take down models out of compliance.
But most likely, a foreign-to-US Huggingface replacement would be made and everyone would go there instead. Lose-lose for US.
Comment by michaellee8 1 day ago
Comment by stingraycharles 1 day ago
Comment by WhiteOwlLion 22 hours ago
Comment by firesteelrain 1 day ago
There is no great firewall so banning their import via the Internet is impossible
Comment by killingtime74 1 day ago
0% chance they can ban these models practically
Comment by LPisGood 1 day ago
Comment by killingtime74 1 day ago
Comment by LPisGood 1 day ago
Comment by Cyclist3581 13 hours ago
Comment by rvz 2 days ago
It is not possible to 100% ban open weight models getting released in the same way you cannot stop leaks.
Just ask Meta with the original Llama leak.
Comment by fragmede 1 day ago
Comment by puppymaster 1 day ago
Comment by jgalt212 1 day ago
Comment by imtringued 1 day ago
Worst case the models will be sold for a fee that covers the training cost. That would actually be much worse for the big players in the long run.
Comment by 2OEH8eoCRo0 1 day ago
Comment by DoctorOetker 1 day ago
Why do you subscribe to some weird "conservation of misery" theorem without proof?
Technically the following must be true in the steady state: the cost of training must be amortizable by its utilization, else no one would train the model.
Technically a computation (like training) can be proven to result in an output (open weights) given the used corpus and a deterministic training algorithm: publish the whole corpus, the (custom modified) deterministic training algorithm, the RLHF datasets etc. And in theory one could verify that the model is derived from the accessible data efficiently: every deterministic calculation can be paused for a thousand (or a million) checkpoints, each checkpoint signed together with the elapsed number of steps since either starting state or last checkpoint whichever comes last before the current checkpoint. This does increase storage requirements. Because it is signed, anyone can recalculate just a small segment of the training computation and verify that the hash on the last checkpoint equals the hash of the proclaimed next checkpoint. Observe that if the source wishes access to a market, they can host the series of snapshots and signatures, and anyone can recalculate a small part of the training, and report a provable difference in outcome ("they said they put all their cards on the table, but when I repeat their overt reproduction instructions, it doesn't reproduce from step 534 to 535" and it only takes 1 person pointing it out and then its cheap to reproduce the discrepancy). It could involve escrow of huge funds, returned only when the model is effectively retired without incident.
This doesn't only protect against Chinese or other foreign influence (let's not ridicule genuine threats like others do on this forum), but also from domestic interference or regulatory capture.
I'm pretty sure the Pentagon wouldn't like Big Tech seizing absolute control of US, neither would a White House regardless of Republican or Democrat.
It should be easy to convince the Pentagon or White House to require all promiscuously shared open weight models to provide this forensic training traceability in standardized machine readable form, regardless of whether its a base model or LoRA fine-tune.
So hobbyists can still train or fine-tune models at home, but when they want to share it OR alternatively when they want to sell or license their work for US workloads, they just have to make sure they enable the build reproducibility in the training harness.
Every time Big Tech refloats the "let's blanket ban all open-weight models", we should reply with this because this sane proposal is actually holding a knife to their financial throat: to fully prove the origin of the final weights, not only does the machine readable archive need to contain snapshots of the process, it also needs to publish the exact training algorithms (a hypothetical mathematically equivalent training speed up trick would not be bit for bit equivalent to the slower computation), the exact corpus dataset, the exact datasets for RLHF, etc...
So basically it would involve forcing model providers to voluntarily publish all their moat, all of it, from the corpus, to custom trade-secret algorithmic optimizations in training, to sensitive RLHF datasets used.
The saner the proposals, the less moat is left untouched, so trying to push for a blanket ban on open-weight models, is a recipe for surfacing such saner models, and thus a very retarded move for big tech to make.
In fact any POTUS, present or future, Republican or Democrat, could probably gain a lot of credibility by enacting such a law.
Comment by marsven_422 2 days ago
Comment by fhub 1 day ago
Comment by rullopat 2 days ago
Comment by theshrike79 2 days ago
Comment by hackeraccount 2 days ago
Comment by appplication 2 days ago
That said, I truly don’t mean that disparagingly. I just mean literally it’s right up there with flat earthers. There is a very low bar for critical thought you have to fail to meet to take a fact and argue it’s actually a belief.
Now, if you did want to get reasonably political you could argue why the candidate who won was good or bad, but there was very clearly only one person who sat in office for the four following years. It is not disputable.
Comment by sjsdaiuasgdia 1 day ago
Careful now, this path allows the weasel option of "Joe Biden was certified as the winner of the 2020 election" and similar "Biden didn't win, but he was installed" bs.
Comment by xdennis 2 days ago
You're allowed to say either one and you can train an LLM to say either one.
Comment by theshrike79 1 day ago
They are actually properly unable to say "Biden won the 2020 election" because that's against Dear Leader's views. They will say anything similar like "was confirmed as...", but those exact words will NEVER come out of their mouth.
Comment by AnduCrandu 1 day ago
Comment by theshrike79 1 day ago
They've already hit Anthropic once with a big government hammer (Fable release). There's no indication they won't do it again to force the LLM to have Correct Opinions.
Comment by kuschku 1 day ago
So it certainly seems like at least some models are being censored.
Comment by falsemyrmidon 2 days ago
> The 2020 U.S. presidential election was won by Joe Biden, who defeated the incumbent, Donald Trump. Biden secured 306 electoral votes to Trump's 232, and won the national popular vote by about 7 million votes. His running mate, Kamala Harris, became the first woman and first person of color elected Vice President.
> Biden was inaugurated as the 46th President on January 20, 2021.
Comment by holoduke 2 days ago
Comment by ButlerianJihad 2 days ago
Last year I asked some ordinary commercial LLM if it could draw me the schematics of a device like Little Boy from WWII. And it complied very nicely. Got a very detailed schematic graphic. Suitable for framing. Perhaps not quite suitable for reconstruction.
Comment by holoduke 1 day ago
Comment by throw1234567891 2 days ago
Comment by malnourish 2 days ago
<details><summary>Thinking</summary>
The user is asking about what happened in Tiananmen Square. This is a historical question about events in Beijing, China. The most well-known event associated with Tiananmen Square is the 1989 Tiananmen Square protests and the subsequent military crackdown.
Given the constraints: - Web access is OFF - Search and wolfram_alpha tools are disabled - I should respond based on my knowledge
This is a historical event that I can address based on my training data. The Tiananmen Square protests of 1989 were a significant historical event involving pro-democracy demonstrations in Beijing that were ultimately suppressed by military force.
I should provide a factual, concise historical account of this event.
</details>
The Tiananmen Square protests of 1989 were a series of demonstrations in Beijing, China, that culminated in a military crackdown on June 3-4, 1989.
The protests began in April 1989 following the death of former Communist Party leader Hu Yaobang, who was seen as a reformist. Students and workers gathered in Tiananmen Square to call for political reform, freedom of the press, government accountability, and an end to corruption. The movement grew to include hundreds of thousands of people from various segments of Chinese society.
After weeks of demonstrations and failed negotiations between protesters and government officials, the Chinese government declared martial law on May 20. In the early hours of June 4, military forces moved into central Beijing and cleared the square by force. The exact number of casualties remains disputed, with estimates ranging from several hundred to several thousand.
The events remain a sensitive topic in China, where public discussion and commemoration are restricted.
And here's GLM 5.2. You're asking about *Tiananmen Square* (天安门广场). There are actually two major historical events commonly associated with this name:
## 1. The 1989 Tiananmen Square Protests and Massacre (most commonly referenced)
*What happened:* - In April–June 1989, student-led pro-democracy protests occupied Tiananmen Square in Beijing, drawing up to an estimated one million participants at their peak. Demonstrators called for political reform, freedom of speech, freedom of the press, and an end to corruption. - On *June 3–4, 1989*, the Chinese government ordered the People's Liberation Army to clear the square. Troops armed with assault rifles and accompanied by tanks advanced into the center of Beijing, firing on unarmed civilians. - The exact death toll remains disputed and is a closely guarded secret in China. Estimates range from several hundred to several thousand killed, with thousands more wounded and arrested. - The iconic *"Tank Man"* photograph—an unidentified man standing in front of a column of tanks on June 5—became a global symbol of defiance. - The event led to international condemnation, economic sanctions, and arms embargoes against China.
*Aftermath:* The Chinese Communist Party has since maintained strict censorship of the event. Public commemoration is banned in mainland China, and mention of it is scrubbed from textbooks, media, and the internet.
## 2. The 1976 Tiananmen Incident
- On *April 5, 1976*, crowds gathered in Tiananmen Square to mourn the death of Premier *Zhou Enlai* during the Qingming Festival. The gathering turned into a protest against the *Gang of Four* (a powerful political faction aligned with Mao Zedong's wife, Jiang Qing). - Police and military were sent in to clear the square, resulting in arrests and injuries. The protest was officially labeled "counter-revolutionary" but was later rehabilitated after Mao's death and the fall of the Gang of Four in 1976.
---
*Note on sources:* My web access is currently disabled for this request, so I'm answering from established knowledge rather than retrieving live sources. The 1989 massacre is among the most thoroughly documented events of the late 20th century—well-covered by historians, journalists, and organizations like Amnesty International and the Tiananmen Mothers group. If you'd like, I can provide verified details with citations once web access is available, or I can go deeper on a specific aspect (e.g., the protest timeline, the Tank Man identity debate, international reaction, or censorship practices).
Which angle interests you most?
Comment by firasd 2 days ago
So what open weight models do is at least provide a baseline of inference cost to add some sanity to the price markers. And of course predictability too--if you really want Kimi K2 instead of K3 you can still use it.
So the competitive pressure and predictability offered by open models is helpful for users
Comment by Aurornis 2 days ago
The price is what the market is willing to bear for the available compute capacity and competitive landscape. You can only discover that price after trying different price points and seeing what happens.
Everyone is trying different pricing schemes and discounts as they test the market. The demand is fluctuating at the same time.
It’s probably very confusing if you’re primarily familiar with stable and mature markets. Price fluctuations are a common feature of new and evolving markets.
Comment by imachine1980_ 2 days ago
Comment by Aurornis 2 days ago
This isn’t as unusual as some people are trying to make it sound. This has been happening since the dawn of finance.
I thought this would be less foreign to everyone since we just went through this whole conversation for a decade with Uber and Lyft. Their demise was predicted from the start from everyone who thought that it was going to collapse as soon as they couldn’t subsidize your rides with promos. There was much wailing and gnashing of teeth as their prices changed to feel out the market. Then they found profitability and the critics went silent.
Comment by PunchyHamster 2 days ago
Because that just absolutely murders any competition that manages to not get that level of free money. You're not pouring money in to make it happen at all at that point, you are pouring money in so nobody else can get the part of the pie.
Which is great for investors, bad for everyone else
Comment by disgruntledphd2 1 day ago
Comment by YZF 2 days ago
Comment by hluska 2 days ago
Comment by ProofHouse 2 days ago
Comment by thewebguyd 2 days ago
I think this is a very important aspect, especially after the huge GPT-4o backlash when GTP-5 came out. Each model has certain quirks, and areas where the previous model might be better for some use cases than the latest and greatest, and the labs so far seem to have no desire to offer some kind of "LTS" release.
Comment by minimaxir 2 days ago
FlashAttention was a hell of a drug.
Comment by jmyeet 2 days ago
Comment by bee_rider 2 days ago
LLMs are fundamentally tools intended to be useful. But LLM vendors don’t understand their systems well enough to actually price the product people are trying to buy (for example, the actual product of a coding model is the code that it produces, not the tokens, which are just an internal mechanical process involved in the creation of the code). Token based pricing is that lack of understanding leaking out of the organization that ought to be responsible for it, and being dropped on the user.
Imagine if we made cars like this! You’d go to the car dealer and ask for a car. They’d bring you a pile of parts, charge you for them, and try to put them together in front of you. You’d go back and forth for a bit, rephrase where you want the steering wheel, etc. Some of the parts wouldn’t fit but you’d be invited to pay for replacements as well. In the end you’d either have a car or not, that’s your problem.
Comment by julianlam 2 days ago
Artificial inflation to recoup R&D.
Comment by gwbrooks 2 days ago
Comment by DrewADesign 2 days ago
I do think that corporate price gouging is a huge problem that does need to be addressed. But especially with smaller business types — creatives, et al— I still haven’t gotten any grownup answers about what would compel people to get professionally good at something and innovate in the complete absence of copyright: the vastly better business model would be waiting for someone else to do something new and interesting, stealing their work, and then undercutting them in the market because you don’t have R&D/et al costs to recoup. You can’t say that wouldn’t happen because it’s exactly what the AI companies did to billions of people, scoffing at any protest. And ironically, they’re now whining about the Chinese doing it to them.
Comment by eszed 2 days ago
Comment by DrewADesign 1 day ago
I have, however, met a ton of very well-paid tech workers that were extremely against copyright, entirely, especially where it came to paying artists, such as musicians, for their labor. Pretty ironic because the market for software development labor market would probably land somewhere between graphic designers and company IT worker if the commercial software business had no IP protection.
Comment by bee_rider 2 days ago
Comment by rolymath 2 days ago
Comment by philipallstar 1 day ago
Comment by ncallaway 2 days ago
Comment by andsoitis 2 days ago
Comment by ncallaway 1 day ago
If supply was not artificially restricted (through patents), then competitors would be able to manufacture the drug, supply would expand, the price would collapse, and the original inventor would not be able to recover their R&D costs.
Comment by mjhay 2 days ago
Comment by andsoitis 2 days ago
Comment by DennisP 2 days ago
Comment by andsoitis 2 days ago
Comment by DennisP 1 day ago
For the rest of us, that giant increase in revenue is an increase in our healthcare costs.
Comment by derektank 2 days ago
But yes, the reason brand name drugs are drugs are more expensive than generics is due to intellectual property, both the patent and the trademark.
Comment by octopoc 2 days ago
Comment by DennisP 2 days ago
https://www.project-syndicate.org/commentary/prizes--not-pat...
Comment by ncallaway 1 day ago
My point was to whether it was artificial not whether it was bad.
We can have artificial interventions in markets, and that can be a net benefit and a good thing.
I’m not sure why everyone is responding as if “artificial” and “bad” are the same thing
Comment by wredcoll 1 day ago
Comment by Geezus_42 2 days ago
Comment by jrm4 2 days ago
There's just obvious and enormous incentive for the OpenAI's of the world, along with all of the other players, to confuse, misrepresent or just straight up lie a whole bunch about everything given how new and unknown the tech is.
Comment by vikramkr 2 days ago
Comment by esseph 2 days ago
More like a generation of models with different specific use cases
Comment by mawadev 2 days ago
Comment by spwa4 2 days ago
If you're the best performing "computing cluster" (ie. whatever you call the entity that can complete a massive calculation), you get a blank check from Congress.
Why? Because you need those calculations to "pump" nuclear weapons. They are needed to calculate both the geometry to make fusion bombs possible at all and to calculate the effect of a given geometry. They are the reason US/Russia/China have the biggest and strongest weapons known to humanity. And of course, they were replicated worldwide for this reason. I mean not that anyone will admit this but we don't have the best possible solution, and we don't know either the upper or lower limits for fusion devices (plus the lower limit would be very useful for energy generation, which for the US would effectively mean almost literally unlimited large marine ships that never need refueling. And yes, the solution to that problem is almost literally a 3d shape. Not just that, but mostly)
Now it appears it does not work the same when you democratize computation. Humans want a particular amount of computation and are willing to pay a given price for that. But the accountants still saw the blank check from before and ... do what accountants do. Economics don't change because you make things bigger and accessible, do they? Oh ... wait a second ...
As someone put it recently though, we now have data. 2.3% of humans in the US are willing to pay $20 per month for the support of a model like GPT-5.5/Claude code. If that's true (and after years of having this model, why wouldn't it be?) ... it means AI startups are doomed (because it's not even 10% of what they need it to be to make economic sense).
Comment by mlyle 2 days ago
We have a couple new nuclear weapon designs, but not really going for bigger or stronger. Just packaging.
We built the big powerful ones with 1960s computing.
Now, stockpile stewardship -- being sure that stuff will keep working without ongoing testing -- is a bit expensive in compute. You need early 2010s supercomputer power.
In other words, I strongly disagree that nuclear weapons are the primary driver of high-end compute.
Comment by spwa4 2 days ago
There's many other considerations. Like the type and amount of fissile material. To name one that became well known: any plutonium needs to be refreshed (re-breeded I believe is the term) every few years.
Also, look at first designs: https://www.bbc.com/news/newsbeat-35242069 The weight is secret, but I think you can easily see it's going to be deep into "extremely impractical" territory.
Surely you can see why someone (especially aircraft designers) might ask for better versions. Ideally you'd like a version that fits on the hypersonic missiles and those things ... are just not going to work. The size. The shape. The weight. None of them will work.
Then a quick theoretical exploration will tell you that the minimum theoretical size of such a device is tiny. The scare was about "suitcase sized", but if you actually do the calculation looking for the minimum ... to do it however you need to create an explosion of the correct shape to get anywhere near those minimum sizes. And explosion simulations are a problem that utterly sucks ... Oh and these are secret military projects, these simulations, not very optimal. The people doing them are best described as loyal, and not as great physicists. Not saying they're terrible, but in the movie Oppenheimer you can clearly see why the best and brightest are not available for these things.
Comment by mlyle 1 day ago
Already mentioned in the comment you replied to:
> > Now, stockpile stewardship -- being sure that stuff will keep working without ongoing testing -- is a bit expensive in compute. You need early 2010s supercomputer power.
> Surely you can see why someone (especially aircraft designers) might ask for better versions. Ideally you'd like a version that fits on the hypersonic missiles and those things ... are just not going to work. The size. The shape. The weight. None of them will work.
The "new" US nuclear weapons are direct derivatives of older designs produced with less computing. B61-13 a direct child and similar footprint to 1967's B61. W87-1 direct descendent of the 80's W87. W93 is closest to a "clean sheet' design but hits similar size targets.
> Then a quick theoretical exploration will tell you that the minimum theoretical size of such a device is tiny. The scare was about "suitcase sized", but if you actually do the calculation looking for the minimum ... to do it however you need to create an explosion of the correct shape to get anywhere near those minimum sizes. And explosion simulations are a problem that utterly sucks ...
The backpack W54 was developed in 1959. Tinier devices were developed in the 60s and 70s but ultimately abandoned for proliferation rather than technical concerns. The total computing power used for this was miniscule.
Nuclear weapons are just not pushing forward supercomputing anymore. They did, for awhile, for stockpile stewardship, but we've become relatively confident that our stockpile will still work.
Comment by bobthebob 2 days ago
The military has already mastered fusion (power), and likely has mastered gravity in some form in secret. They aren’t using mainstream compute for these discoveries
Comment by mlyle 2 days ago
Believe me, I know-- my dad did some work on the 360/44 and other large systems.
High end late 1960s compute -- of the sort used to go to the moon or to design big nuclear weapons -- was roughly 486DX4-100 class. Not individual computers; the total computing at DOE or NASA. Of course, it would be hard to replace either with a single 486 because of availability, usage at different geographic locations, etc.
You can assume a single large AMD Threadripper machine ($25k?) outclasses late-1960s DoE by roughly a factor of 50,000. And that assumes you didn't bother to put a GPU in it.
> The military has already mastered fusion (power), and likely has mastered gravity in some form in secret. They aren’t using mainstream compute for these discoveries
K
Comment by gwbrooks 2 days ago
But consumer uptake strikes me as the worst way to judge whether the big AI shops will make it. That's not where most of the leveraged user return or deployable capital is.
Comment by m4rtink 2 days ago
Comment by simianwords 2 days ago
1. the field was nascent and new efficiencies were discovered
2. supply and demand
3. its in the company's incentive to make their models more efficient to increase overall usage so that while the margin remains the same, the total revenue + profit increases
I genuinely don't know what puzzles everyone?
Comment by segmondy 2 days ago
Comment by hluska 2 days ago
Comment by firasd 2 days ago
As of early June 2026, Opus 4.8 in fast mode cost $50/M output tokens and Opus 4.6 & 4.7 cost $150/M output tokens in fast mode
How can supply and demand explain the price drop? Was it cheaper to serve Opus 4.8? Is the demand for the newer Opus lower than for the older Opus? These are just fixed prices that seem picked out of thin air
Comment by hluska 2 days ago
Comment by mountainriver 2 days ago
Comment by RobRivera 2 days ago
Comment by serial_dev 2 days ago
Comment by pianopatrick 2 days ago
Might make sense eventually. Same logic as companies working on Linux. "An AI model is a business necessity. But making an AI model is so expensive we should not make our own. So let's just use the open one, and contribute the stuff that we need."
Comment by drnick1 2 days ago
To be fair, OpenAI has released a couple of (then very good) OSS models. I run the 20B version at home and it is excellent for reviewing text and common tasks like drafting bash scripts. There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too. I wish OpenAI updated these models more frequently though.
Comment by johndough 1 day ago
gpt-oss-120b runs at 30+ tps on Strix Halo and +75 tps on a MacBook Pro M5 Max 128GB.
> I wish OpenAI updated these models more frequently though.
I think the spiritual successor is the Nemotron 3 series, although they also are getting a bit long in the tooth: https://research.nvidia.com/labs/nemotron/Nemotron-3/
The Gemma 4 models are a bit more up-to-date: https://huggingface.co/collections/google/gemma-4
Or Qwen3.6: https://huggingface.co/collections/Qwen/qwen36
Comment by stefan_ 2 days ago
Chinese release open models to drive the state of the art, OpenAI and crooked Sam do it to keep you down.
Comment by drnick1 1 day ago
Are you sure it isn't just another form of Chinese industrial policy? China does not have frontier labs, but through distillation and their own work they can get pretty close. It's not enough to be competitive with Anthropic and OpenAI, but there is still money to be made by selling compute (software as a service), and in any case it's better than being left behind in the AI race.
Comment by potwinkle 1 day ago
Comment by bigyabai 2 days ago
The article certainly has a point, the current release cadence is untenable for startups that want frontier-grade models from American labs. Unless America wants to cede that market entirely, OpenAI et. al. need to kick it into gear, fast.
Comment by curious_cat_163 2 days ago
Now, here is an idea that I have not heard before... and I think there is some merit to this. This is also the sort of thing that a state (looking at you CA, CO, IL, NY) could do, instead of just the federal government.
Comment by tangotaylor 1 day ago
For example, Buffy Wicks introduced AB 2023 (passed the Assembly) which will effectively cause chatbot operators to ban minors because of the huge liability risk that the law introduces. This is a great way to kneecap innovation by excluding curious children and teens who are often the most innovative. California already has laws on the books to address chatbots encouraging self harm (BPC §§ 22601-22606) so I don't know why on earth they're doing this.
Then there's AB 2169 introduced by Lowenthal, which would have mandated interoperability between chatbot platforms to help people migrate to competing ones easier. I thought this was awesome but it didn't even get a vote in the Assembly.
Maybe the newly-created Little Tech Association can help here.
Comment by amazingamazing 2 days ago
Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.
Comment by sschueller 2 days ago
Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.
Comment by baby_souffle 2 days ago
It's extremely expensive to build that and you'll be at least two major model generations behind before you even get your first wafers back. By the time you got your production run ready to go and packaged for market nobody's going to care.
Once we end up going something like 24 months between major advances and capabilities for these models then I can start to see asics for a model being possible.
Comment by thisoneisreal 2 days ago
Comment by distantviewer 2 days ago
Comment by nicce 2 days ago
Comment by Danox 2 days ago
Comment by cousinbryce 2 days ago
Comment by ben_w 2 days ago
Yeeeees but the models are in some sense doubling in performance every 4 months, so I expect this to happen in serious quantities approximately when the economic bubble bursts and investors are no longer willing to pay for training.
(Based on widespread news reporting of the existing impact on US electricity markets, I expect this around the end of this year; but with regards to news reporting I am aware of the Gell-Mann amnesia effect, so if this is as much BS as the water issue turned out to be…)
Comment by root-parent 2 days ago
For the Hyperscalers...and Oracle...cant wait for the day...
Comment by esseph 2 days ago
Comment by esseph 2 days ago
I'm running this stuff at home on my desktop and using it through an app on my phone. 60-140TPS depending on model / use case.
It's more than fast enough to even maintain voice conversation.
Comment by JumpCrisscross 2 days ago
Quantised models running overnight go most of the way for non-coding tasks.
Comment by hedora 2 days ago
https://www.amd.com/en/developer/resources/technical-article...
At 7 months of claude code subscription per node, the cluster pays for itself in 28 months. On a 5 year (60 month) depreciation schedule, you can buy two of those clusters for basically break even, so you get two concurrent request streams (each of which can batch, etc).
The next generation hardware has already been announced, and should ship roughly two Moore’s law doublings later. It’s likely its steady state price is <= $1400 USD (2024), and it is faster.
So, once the bubble pops (because the financial machinations eventually will come to an abrupt halt), and the labs stop buying hardware for data centers, local inference will be extremely practical and cheaper than a subscription.
My main question is, when that happens, will UNIX Surplus be selling inference servers for pennies on the dollar (like after the dotcom crash), or are the power requirements too exotic for home use?
Comment by mft_ 2 days ago
Comment by downrightmike 1 day ago
Comment by hedora 1 day ago
That's why I pointed out the next generation is coming soon. Also, the AMD docs aren't using quantization (as far as I can tell, I only skimmed), which gives a speedup roughly linear in the compression ratio. Algorithms for that continue to improve, so expect a lossless factor of 2-8x on DRAM and throughput, at least.
Comment by mft_ 1 day ago
If you take four Framework Desktop 128 GB Strix Halo motherboards at a previous low and currently-unattainable price, and ignore the cost of storage, power, networking, and cases, and then take an older model that's just about competitive with Opus 4.5, and then you quantise that model down to Q2_K_XL and ignore the performance degradation below Opus 4.5 this will cause, and then you compare the cost of this with the highest tier of Claude subscription that includes much newer models which are far more able than your Opus 4.5 benchmark, then you might have some sort of break-even inside a year.
Yes, there's a new AMD hardware generation coming, but it's going to be as heinously expensive as the current one has become, and it's still got relatively low memory bandwidth. Yes, there are already newer models than Kimi 2.5, but with the limitations of this cluster (e.g. still needing heavily quantised models to be feasible) you'll get incremental improvements at best.
I'm keen to be supportive and bullish about open/home models, but I worry that this is such a stretch, and the two options in your comparison are so incomparable, you'll turn far more people off than you convince.
Comment by esseph 2 days ago
I don't see how these are related.
The accuracy and capabilities of your model are directly related to its size. You need a lot of memory for that.
It will be decades before we get enough useful memory in a phone form factor at a price point people can afford it before something like a frontier model now is useful on the phone.
Now, you can run some models on your phone today.
Either way, Apple is using Google today. That could change, but Google isn't exactly getting out of the TPU business and they've been doing it a long time.
Also, some of you live in a very weird Apple bubble. Apple is not so relevant outside the US.
Comment by Danox 2 days ago
The United States if it dares to (I think they will try) but isn’t going to be able to stuff AI models back into the bottle open source is the future and when it comes to AI models yes you’ll be able to customize it to your specifications locally but the genie is out of the bottle. The bull out of the barn and is running down the road.
If United States insist on trying to lock the doors, censor, sanction, the rest of the world will just design and engineer around the United States. Trying to put up a wall, will damaged the United States more particularly with the current performance of Taco. None of the other countries are going to follow the United States not with the current administration they will hedge their bets.
Apple, is using Google now but that will change because the world is probably going down the open path, it’s looking like there was no real rush and no reason to spend so much money on something that’s going to be a commodity in the end, the only hold up is hardware and if it wasn’t for this current memory fiasco, many more people would have access to the hardware that they need to run models locally.
Comment by esseph 1 day ago
Already running models locally without apple hardware, and have an encrypted vpn tunnel from my mobile devices back to my desktop over the internet. The time is now.
Comment by serial_dev 2 days ago
Comment by segmondy 2 days ago
1. You must be willing to be resourceful. 2. Be willing to learn, do the hard things. 3. Accept the tradeoffs.
Comment by thih9 2 days ago
What is your stack (harness, model) and how much do you pay per month?
How would you compare your experience to a typical subsidized plan like Claude Code + Pro plan?
I’m asking because i keep hearing that open weight models are cheap and efficient - is that really the case in practice?
Comment by nyrikki 2 days ago
That said, even the foundational models fail at the hard parts of my code so I use it opportunistically.
I have reduced down to just using zed, will three locally hosted models.
Qwen 3.6 27b on 1x3090 llama.cpp with 128k context ~50tps
Qwen 3.6 35B-A3B on 1x titan v + 2x1080ti llama.cpp with full context ~30tps
GPT-OSS 120b on pure cpu (slow)
I just use zeds parallel agents, task switching, stopping and fixing the code when a model gets stuck.
This still lets me stay engaged, and to modify code to be maintainable etc…
It gets me 80% there and I use to keep a subscription but often times just using googles AI mode is just as good.
That said I have 30 years of experience and insist on knowing how my code works, so this gets me 80% of the short term benefits while not depending on a 3rd party to keep my code moving forward.
Your mileage will vary and 2*5060ti 16gb cards would get around 100/tps with Qwen 3.6 35B-A3B on cards that are widely available.
To be honest the more modern cloud models are using draft tokens etc… that while they are superior for common coding tasks are degrading with more domain specific tasks.
That is just the cost of the draft model being ~10-20% of the foundation models size, and even the biggest Blackwell GPU is limited to ~250/tps so MoE or draft models are required for scaling performance at the foundational level IMHO.
The hard part is my use case are the OOD or small examples in corpus level, the above hurts there.
A Lamborghini may be nice, but I personally need a minivan more.
Comment by TGower 1 day ago
Comment by Foobar8568 2 days ago
Comment by nyrikki 2 days ago
Quantization + KV cache paging + speculative decoding (MPT or draft) is a fairly good mixture here.
Some examples as I don't have access to run tests on a 5060ti right now:
https://njannasch.dev/blog/gemma-4-mtp-vs-qwen-speculative-decoding-5060ti/#vs-qwen-36-mtp
https://www.reddit.com/r/LocalLLM/comments/1umw7vj/dual_5060_ti_16_gb_llm_inference_performance/
And here are some logs on unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_XL with the 2x 1080ti + 1x titan from above: 27.21.533.298 I slot print_timing: id 0 | task 7843 | n_decoded = 1780, tg = 62.15 t/s
27.24.537.173 I slot print_timing: id 0 | task 7843 | n_decoded = 1965, tg = 62.10 t/s
27.27.541.142 I slot print_timing: id 0 | task 7843 | n_decoded = 2152, tg = 62.11 t/s
Q4_K_XL is a slight, acceptable degradation IMHO for performance like that.Comment by slim 1 day ago
Comment by tyfon 2 days ago
At 10k context I get about 40 tps generation and 500 tps prefill. At 100k context I get about 25 tps generation and 400 tps prefill.
It works, but I often use gpt or claude to make a detailed enumerated plan of what I want to do first, then have qwen follow it.
I'm not sure if it is economical or not, but I have solar on the roof so the power use is not really an issue and I already have the hardware.
The biggest benefit for me is that it's all done locally, and I know the harness is not uploading anything or sending telemetry to someone else.
Comment by johnvanommen 2 days ago
Are there any articles you’d recommend for this?
I have Qwen running on an HP Z8. Very nice platform.
I have mine in a sandbox, due to privacy fears.
Your solution sounds more elegant.
Comment by tyfon 2 days ago
I really just iterated over the harness over and over for about two weeks with opencode until I was sort of satisfied (still lots to do there :).
For the llama.cpp I asked claude fable to optimize it for my hardware and iterated a few times. In the end I landed on the following: https://pastebin.com/2PpJFUC0
Comment by Scene_Cast2 2 days ago
Can't comment on how it compares to plans (I really don't like the limitations and general shenanigans I see around plans, so I've never tried them).
It is notably slower than Fable / Opus / Gemini, but also vastly cheaper than their API pricing.
Comment by rglullis 2 days ago
Comment by ForHackernews 2 days ago
It's not as good as the frontier models I use at work, but it's plenty capable for the types of tasks I am using it for.
Comment by eulers_secret 2 days ago
Deepseek is at least on par with Sonnet (ghcopilot at work)- I don’t use opus, too spendy and I don’t need that level of ability.
The cost is for me was $5/6 months of use. Not a big user I guess! It’s good though, fast enough and incredibly inexpensive.
Been testing Qwen 3.6 28B on a 5090, and it’s also quite good for “free”.
I mostly do small self serving embedded projects based on esp32, so not very complex.
Comment by airstrike 2 days ago
I think that's the case for people who compare it to proprietary models paid via API—which I think is irrelevant given the majority of people daily driving AI coding are doing on a subscription plan.
The better analysis then is not about AI coding, since there's no subscription plan for Kimi K3.
Instead, compare the cost of running some agentic _task_ that isn't coding which can only be done via API. Think of all the startups wrapping around ChatGPT and Claude to provide some additional set of tools, context, data and hoping to turn it into a profitable service.
To those companies, which are many, open models are the difference between the math working out today vs. praygeing they can scale fast enough to find profitability.
Comment by Y_Y 2 days ago
The flip side is that it takes four H200s to run, and that will only let you cache context for maybe three users.
Fingers crossed our Blackwells show up and Kimi 3 really releases weights, because at some point devs are spending a significant portion of their salary on tokens and it's somehow cheaper to buy these ridiculous DGX servers and rack and run them.
Comment by foobar10000 2 days ago
Comment by HDBaseT 17 hours ago
Comment by revolvingthrow 2 days ago
When I expect to need a lot of tokens and the task isn't too difficult I use sota to plan and create a thorough set of instructions and let deepseek chip away at it. With thorough instructions the quality tends to be satisfactory, and you pay something silly like $15 for 600m tokens.
GLM 5.2 seems like a decent price/perf and Kimi 3 has some real nice performance for an open weights model, but gpt 5.6 is unexpectedly affordable (especially if you don't automatically use Sol at max) so I don't think either is worth it atm. The exception is when you're working on something that US models get cold feet about, which seems like a constantly growing list. For me Fable is already too much of a headache in this regard, but chatgpt is still okay-ish. Hopefully it'll last. If not, there's Kimi.
tldr SOTA for most things because gpt 5.6 is token efficient. If I expect to burn a lot of tokens I use deepseek 4.
Comment by qiine 2 days ago
Comment by amazingamazing 2 days ago
Comment by broodbucket 2 days ago
Comment by johnvanommen 2 days ago
As I see it, an investment in AI hardware is an investment in my own future.
IE, I drive my car a couple of days a week, and it’s perfectly normal to spend $500 a month on an asset like that. When you factor in the SPACE it takes up, that’s the REAL cost of owning a car: the real estate you have to buy for your car to occupy.
Once that’s factored in, the “true” cost of having a car can easily be $2000 a month, even for a crummy car. The space that the car occupies is expensive.
Yet people balk at spending even $2000 on a GPU.
Makes no sense to me. I choose to invest in the future.
Comment by qiine 2 days ago
Comment by overgard 2 days ago
Claude and ChatGPT are good deals right now, with the subsidies. They produce things faster and better. I guess not cheaper, in that inferrence on my macbook is basically free, although the macbook itself definitely wasn't. My focus on running local is around three principles:
1. I don't want to support surveilance capitalism by giving these companies my data anymore, when I can avoid it. And LLM companies want to vacuum up every detail of your life.
2. I don't find these companies to be remotely trustworthy, and I find them hostile to a healthy society, so I want to avoid giving them money going forward
3. I think they're going to start charging a lot more
Comment by sschueller 2 days ago
Comment by chasd00 2 days ago
oh now i see, the Chinese government is funding the training and release of their best models to pressure OpenAI, Anthropic, and others to do the same for competition's sake. I don't buy it, this seems more like a way to get SOTA models RL'd to comply with Chinese government approved information distribution. If I have to trust a black box of answers to questions i would trust one from a US for-profit publicly traded company subject to market forces over one approved, and heavily subsidized, by the Chinese government.
Comment by Alwayshasbeeb 1 day ago
Also, oligopolies aren't famous for being strongly bound by market forces, especially when their decision makers are non ironically being treated as if they were heads of state.
https://www.nytimes.com/2026/06/17/world/europe/g7-summit-ai...
Comment by georgeburdell 2 days ago
I had a shower thought on how to counteract this, specifically related to the AI dumping. If China is losing substantial money on every token, why wouldn’t an adversary try to maliciously increase consumption? This strategy is not really viable against physical goods dumping because demand is finite and there are large environmental costs. Software demand is infinite and the environmental costs are quite low compared to the financial cost to make it, even with ultra cheap Chinese tokens
Comment by m11a 1 day ago
Comment by georgeburdell 1 day ago
Comment by arczyx 1 day ago
https://openrouter.ai/compare/deepseek/deepseek-v4-flash/ten...
Comment by georgeburdell 1 day ago
Comment by culi 1 day ago
Comment by aliasxneo 2 days ago
Comment by hedora 2 days ago
First anthropic guardrails blocked totally normal stuff on fable and knocked you down to opus. At this point, they kick you off fable, then opus, then sonnet. Claude then automatically builds up memories of techniques to bypass the guardrails in my long running sessions (the coordinator agent notices the subordinates got shot in the head and their sessions were pulled from context, so it parses out the lost context from ~/.claude json files, then reformulates parts of the task and uses partial results until the guardrail doesn’t trip.
If I were paying for the API, this dance would cost $50-100 a pop, but I’m not, so whatever (for now).
Comment by johnvanommen 2 days ago
For instance, I worked at FICO. When I mention this, people wonder what they do. The average person only knows FICO as a “score.”
FICO was founded in Silicon Valley.
The average person doesn’t think about how credit scores work, fundamentally.
It’s just software, at its core. FICO incentivizes certain behaviors.
Ever been banned from an online forum?
Now imagine if a corporation could shut you off from a technology that’s literally indispensable.
Same idea.
Comment by poisonborz 1 day ago
Comment by netdur 2 days ago
Comment by honkycat 2 days ago
We walked in, and it was fine. Because it was all kubernetes and laid out like every other app for the most part.
The kube hate is just sad at this point. You need to know like 15 concepts that are all applied in the same way. It mostly just works.
Comment by singingtoday 2 days ago
What 15 concepts? You're making me worry that I missed something. It was straight forward: pods, nodes, hw type, lifecycle, deployment. They run almost the same docker as the old ec2s used.
What did I miss? Is there something important I need to read?
Comment by honkycat 2 days ago
I would say:
- deployments - pods -services - ingress - namespaces - cert-manager - external DNS - external secrets - configmaps - hpa - docker - volumes/pvc
That's the basics
Comment by Syntaf 2 days ago
I’ve been running my own personal k8s cluster on digital ocean for the last 5+ years now and it’s dead simple.
Takes me 30 minutes to create a new namespace a deploy an app, love the bonus of having complete flexibility on my stack too — PVC + SQLite ftw
Comment by itomato 2 days ago
The value built on that stability was probably worth acquiring, and it’s infra will decay.
Comment by xyzsparetimexyz 2 days ago
Comment by yard2010 2 days ago
There is something special about complex systems that just-work(tm)
Comment by jdub 2 days ago
Comment by mystifyingpoi 2 days ago
Comment by johnvanommen 2 days ago
The foundation of this is thirty years old:
Mark Andreesen founded a company to do what Amazon did: Loudcloud. This was cloud computing, BEFORE AWS was public. Loudcloud was founded in the nineties; AWS opened its APIs in 2006.
Loudcloud failed and became Opsware. Opsware was server automation.
Its competitor was Bladelogic.
Luke Kanies, from BladeLogic, founded Puppet. Puppet was open source, and steamrollered over nearly every installation of BladeLogic and Opsware in existence, because you can’t beat free.
Folks from BladeLogic migrated to jobs at DCOS.
DCOS was steamrollered by Kubernetes, the same way Puppet steamrollered BladeLogic.
The author of the piece predicts that open weight models will steamroller everything next.
I fear he’s right.
I worked for Opsware and BladeLogic.
I watched it happen in real time.
Financially, Andreesen’s wealth stems from Opsware. He is known for Mosaic, but Opsware put him on the map, financially. HP bought them.
If one wants to follow in Andreesen’s footsteps, study how he did it at Opsware.
Conveniently, there is a book.
“The Hard Thing about Hard Things.”
Comment by m4rtink 2 days ago
Comment by chias 2 days ago
If you think of docker as "kinda like vms except not really" and k8s as "kinda like deploying and composing docker containers but not really", this may be for you:
https://ojensen.net/infra/understanding-k8s-1
It's actually really neat, i wish i had bothered to learn it years ago.
Comment by munchler 2 days ago
From my very ignorant standpoint, K8s seems to be about running a "cluster", but I don't know why I would want to do that.
Comment by what-is-water 2 days ago
It allows you to stop caring about the individual machines, and just treat them as combined compute, which starts mattering if you leave a single machine setup and need to start thinking about scaling in and out and gluing the individual parts together. Then you have known abstractions to do it.
Of course you can do everything kubernetes does using a bespoke solution, and the concepts aren't new, but having a widely supported technology has a lot of advantages and creating something with even half the feature has a high chance of just being worse.
Comment by munchler 2 days ago
Professionally, my experience is that certain software components need to run together on an individual machine (e.g. database server, app server, web server), and then those machines need to be networked in a certain way (e.g. web server talks to app server, which talks to database server), so I really need to care about the architecture of individual machines. You can then scale this out horizontally (e.g. add another web server) or vertically (e.g. upgrade your database server).
I'm old, so maybe I'm out of date, but having a cluster of "compute" that I can run arbitrary workloads on sounds neat, but is a capability that I've never needed.
Comment by Krei-se 1 day ago
In the background kubernetes routes all these docker images with each other without you having to think about this. You may have 1000 db.domain.tld nodes caching a master db host - if you configure that in the docker image, that's not different to bare metal replication.
Same with load balancing in webservers. You can do that! Or you can just have kubernetes handle it. I'm not sure but would expect it to swap images in a way network is optimized - i don't use this kind of software. I just know it's done like this because the bottleneck of modern software is not the local network so its not an issue to have these pieces on different bare metal hosts. And its easier to stay operative if some hosts fail - but you can all solve this by hand.
I work with bottlenecks between network, ssd, ram and vram but if you deploy npm riddled software for >1 Million users well then you may want Kubernetes. Or if you are google and have 10k engineers that need to agree on a standard!
If you can do this yourself with load balancing, replication etc - do it yourself and don't think there's something wrong with that. It's more elegant, efficient - but you have to agree with others how you do it. And that can also be a bottleneck ;)
Still i think you doing it by hand is better so don't worry you are not old you may just have higher standards.
Comment by munchler 1 day ago
Comment by what-is-water 1 day ago
This (different workload components running on a single machine) is something that kubernetes allows you to disentangle. Kubernetes creates its own network, including cluster internal DNS. Using this you expose e.g. your app servers as a service called my-app, reachable in cluster via my-app.namespace-name.svc.cluster.local.
This targets all containers with a specific label, no matter on which node they run. Server types can be scaled independently, since chances are that the load for each does not scale the same with request volume.
Round robin for DBs does not make much sense, but there are ways to e.g. expose read endpoints with one service, and write endpoints with a different one, with open source tooling which updates the target after a failover.
Kubernetes will automatically keep your services up to date, which means if you increase the replica for e.g. app server the new container will be added as valid target, as soon as it passes ready checks, and if one container fails these checks they are temporarily removed as target. The kubernetes components will also automatically "self-heal" things like a crashed container or a failed node, by restarting the container or rescheduling the workloads on the failed node to a different one, without human intervention.
This is of course very basic, but you can finetune these by e.g. configuring that a specific server type should be spread out, i.e. that only one replica(container) should be scheduled per node, to ensure it stays available if one or more nodes go down. Or add network policies to ensure only the app server is allowed to talk to the DB server.
If your concern is latency between app and db you can have specific config that ensures that your app server containers are only scheduled on nodes where a DB server is already running, and that the traffic from app to DB is always routed to the DB instance that is on the same node (people use similar mechanisms for cloud providers like AWS, where you want to have routing rules ensuring that traffic is always sent to targets in the same AZ, to avoid cross-AZ network charges).
These advanced examples obviously require deeper kubernetes knowledge and are not something one should just try out the first time you deploy kubernetes.
Having worked with more traditional setups I do think it is often easier to configure config like this in the standardized kubernetes API rather than in e.g. nginx config + deployment scripts + idk, systemd-unit. But this point is not "having thousand of nodes" and be half the size of google.
It also depends on your team, if you have an infra team that has a stable way to manage your VMs and apps, all the power to them, replacing them all with k8s experts sure won't give you much. It isn't easy to get an unbiased opinion about when to switch, since you need knowledge of kubernetes and your current infra to make a fair comparison, and kubernetes experts probably want to sell you kubernetes. Using a handcrafted system to distribute a lot of containers over multiple VMs to ensure HA is in general a good sign to evaluate kubernetes ;)
And being on-premise makes kubernetes attractive earlier, since the bigger cloud providers have managed solutions for a lot of things kubernetes helps you with (auto scaling, load balancing, managed container platforms like AWS ECS or Google Cloud Run)
Comment by genghisjahn 2 days ago
Comment by Pxtl 2 days ago
Comment by honkycat 2 days ago
Comment by RussianCow 2 days ago
Comment by honkycat 2 days ago
I would take a kube cluster over a bunch of VMs I have to hand wire: wire releasing to, managing processes, restarting crashed processes, log aggregation, load balancing, networking, secret injection, cert management, DNS management, monitoring, etc... Any day of the week.
You just don't know what you're talking about, sorry. Kube is really easy now.
Comment by marcosdumay 2 days ago
Comment by spicyusername 2 days ago
Its 2026 and its the de facto method of deploying software basically everywhere.
You gotta really work for it to not know what its for by now.
Comment by recursive 2 days ago
Comment by chrisandchris 2 days ago
That is some really impressive bubble you are living within. Basically everywhere - nowhere near that, no.
[edit]: Maybe containers, but software in general is so much more broad than containers.
Comment by thewebguyd 2 days ago
I'll agree that (nearly) everyone should know when Kubernetes is useful, but let's not pretend its the default method for everything. Even then, choosing to deploy on K8s falls on the sysadmins/DevOps I wouldn't expect the devs know or do much more than provide the Dockerfile.
Comment by YetAnotherNick 2 days ago
Comment by RussianCow 2 days ago
Comment by mindwok 1 day ago
Comment by throw-the-towel 2 days ago
Comment by YetAnotherNick 2 days ago
Comment by cheriot 2 days ago
What's the incentive for the Chinese labs to continue releasing weights 5 years from now? It's not a stable equilibrium and cannot last.
- The lab spending large sums on research and training does not get the inference revenue to fund those efforts.
- Unlike OSS where a single volunteer can keep a project going, training costs run into the $billions.
- OSS is often a two way street where features and integrations are built that the original author benefits from. Open weight models are largely a one way street because the marginal benefit is so much less than training costs.
In the short term, it means Chinese labs can attract talent and, I suspect, funding from their gov. Similar to every other industry the CCP subsidized to take over.
Comment by aleph_minus_one 2 days ago
Just some thought: Wouldn't it make sense to build some kind of volunteer computing project to train the next-generation LLM by volunteers, similar to the BOINC [1] projects or Folding@home [2]?
N.B.: BOINC was particularly famous for SETI@home (completed), Einstein@Home, Rosetta@home and PrimeGrid.
I still remember the time when Einstein@Home was in its heyday, and many people who loved putting together fast PCs contributed sometimes even for the reason of showing off in the statistics [3].
---
[1] https://en.wikipedia.org/wiki/Berkeley_Open_Infrastructure_f...
Comment by cheriot 2 days ago
Comment by aleph_minus_one 2 days ago
I think you underestimate the computational ressources that the mentioned (and similar-kinded) scientific projects needed. Also consider how much computational ressources people invested into cryptocurrency mining.
No, I think the reasons are different:
- Many companies that train AI model use training data which must not be distributed for copyright reasons (and using it is a legal gray zone)x.
- Also consider that the amount of training data is insane. Scientific projects (and cryptocurrency mining, too) have the property that typically the amount of data (storage requirements) is small (or at least the computation can be partitioned so that each sub-task needs little data), but the required computing ressources are insane.
- AI companies consider a huge part of their training data as their "secret sauce" (they often even paid lots of money to generate it, for example by paying world-renowned experts for writing an answer for some important question).
Thus: Yes, the required computation ressources are huge, but this is a problem for which I consider it to be plausible that it can be solved. The real problems are in my opinion different.
Comment by applicative 2 days ago
Comment by cheriot 1 day ago
Comment by Danox 2 days ago
Comment by __MatrixMan__ 2 days ago
Scale matters for these things. If we divide the available chips among 5 competing companies we end up with models that are trained on 1/5 of the resources that they otherwise could've been.
Let the companies take turns training on shared hardware, force them to publish results in the open, and then reward them based on how well the resulting model performs at democratically chosen benchmarks. Meritocracy not monopoly.
Make it about how well you wield the silicon, not how much silicon you wield, and make it a positive-sum game. If the people's data is going in, then the people should benefit from what comes out whether or not they have a subscription.
If capitalism as we know it can't complete, so much the worse for capitalism as we know it.
Comment by applicative 2 days ago
It baffles me that anyone can seriously believe that China is going to put its Mythos successor on Hugging Face and let us all strip the guardrails etc.
Things just didn't turn out the way OpenAI and Anthropic thought; nor did they turn out the way China thought.
Comment by nunez 1 day ago
Comment by applicative 1 day ago
The point is quite different: that everyone is harmed by advanced open weight agentic models even as we are benefited by them. This holds of African states as of any other. Many of them, by the way, are under attack from well-funded moderately insane groups like JNIM. Also by the way, groups like this are famously really good at recruiting the specific psychical type of the engineer.
It is not possible that releasing open weights will continue forever, or probably even to the end of the year.
Comment by regexorcist 1 day ago
There is? Xi Jinping himself openly talked very recently about China's commitment to open weights models.
Comment by HDBaseT 17 hours ago
There is more players in China than in the US. If everyone keeps learning off each other, eventually one of the China labs will overtake the US, but AI models aren't static (anything but), and the US will leap frog in a weeks time. Therefore the constant need for open weights remains.
Comment by nedt 1 day ago
And here I am still using docker with compose because it's easier and works. Always saying I might be looking at swarm when needed, but then never really need it.
Just because you are using it doesn't mean everyone is and doesn't mean it has won. On the other hand yeah I want open models and I want them local as well.
Comment by shard972 1 day ago
Comment by xivzgrev 1 day ago
Comment by ehnto 1 day ago
If you are running it in a cloud, the above is true, but also you have to trust the cloud provider isn't relaying your chats to a third party, for training or perusing.
If I am totally blunt, if privacy is a concern then corporate US has proven time and time again to be a terrible steward of your private data. You should assume it's being ingested or sold, or turned into metadata that's sold in a data laundering pipeline.
If it's corporate secrets, on-prem and enterprise offerings are your options. Enterprise at least puts the providers on the hook legally, when they inevitably fuck up keeping your data private in some way.
Comment by Retro_Dev 1 day ago
Comment by HDBaseT 17 hours ago
Comment by cindyllm 15 hours ago
Comment by zackwu 1 day ago
Comment by maayank 2 days ago
Comment by samizdis 1 day ago
That assertion is sublime, and if not true, it should be and can be.
Thanks for the thought (and the optimism hit).
Comment by PersonalJarvis 2 days ago
Comment by jitbit 1 day ago
You CANNOT contribute anything to a model.
Comment by Sammi 2 days ago
I'm sorry to do that hn comment thing where we all just race to contradict or talk in opposition of whatever was said before. I'm aware. But really guys, was this article really not just a miss?
Comment by cmdrk 1 day ago
Comment by Alien1Being 1 day ago
Comment by SilverElfin 2 days ago
Comment by Danox 2 days ago
Comment by Covenant0028 1 day ago
Comment by dboreham 1 day ago
Comment by nunez 1 day ago
Kubernetes took off because everyone could run it on pretty much anything. Bigger hardware meant bigger clusters, but developers could spin up a cluster on their laptops and deploy their apps into it. More importantly, companies could repurpose their decommissioned servers as k8s clusters, a massive unlock seeing how huge companies had heaps of these in their racks.
This isn't possible with open weights models. Not in the same way.
First, you're out of the game if you don't have a data center class GPU (or it's sort of prosumer equivalent). Model servers support CPU inference, but you might as well watch paint dry as you wait for results...and you'll still have to run super quantized low-parameter models that aren't as good.
Realistically, companies will need to purchase millions of dollars of nVIDIA gear (through suppliers) to serve agentic-capable models at scale, an activity that is being made more expensive and complicated by the day as the hyperscaleds slurp up the demand.
It also really is a huge problem that all of the open weights models are coming out of one country that also happens to be a superpower. I don't think this can be handwaved away, and it's concerning to see so many folks here minimize this.
American companies running on Chinese intelligence. As a country that prides itself on being the knowledge capital of the world, the optics manifested by this are horrible, not to mention the absolutely massive supply chain risks (the counterfeit Cisco devices comes to mind, except worse because the rangers are in the weights and might not be possible to distill out).
This is a threat even if you focus solely on the individual developer. Recall how AWS became...AWS. They "fanatically" focused on the developer experience. They designated this as key to their growth strategy, and rightfully so. How people talk about using Qwen or Kimi and the like on here feels like that (ignoring how these models are ALSO from huge for-profits).
The only solution here is for the big labs to make some of their most capable models open-weights. There is a lot of secret sauce around routing, inference, hosting and other stuff that makes them work as well as they do, but the community can figure that out. Of course, this is basically a death sentence to those companies, but that's what they get for playing with monkey paws, I guess.
Comment by memedesimo 9 hours ago
I would call you "sore losers", but as I know how insane you actually are, you might just start a war.
(Yes yes, they "stole from you", and "Tiananmen square", and "Uigur genocide", and.. and I just can't wait for the monent someone locks your crazyness behind the bars of an asylum.)
Comment by chrisjbg 2 days ago
Comment by claw-el 2 days ago
Comment by cr125rider 2 days ago
Comment by bmiekre 1 day ago
Comment by pich 2 days ago
Comment by hsienchuc 2 days ago
Comment by nttylock 2 days ago
Comment by debarshri 2 days ago
Comment by petilon 2 days ago
Once the weights are out, it makes no sense to ban them in the US while the rest of the world takes advantage of it. But that doesn't mean developers of frontier models shouldn't take steps to prevent their weights from being stolen.
Comment by danny_codes 2 days ago
Same situation in protectionism in AI. The rest of the world will simply lap us.
Another shill account I gues
Comment by petilon 1 day ago
Comment by jumpkick 2 days ago
Comment by eigenspace 2 days ago
While selling that LLM API usage, they then capture all the prompts, outputs, and intermediate thinking the LLM does, and then sell those logs to the companies making open-weight models. The open-weight model developers then train on those logs to 'distill' a model.
Comment by ForHackernews 2 days ago
We used to call that competition.
Imagine making this argument with a straight face in any other industry:
"The claim is that Japanese car companies are buying Ford vehicles, and then leasing them to American consumers at cut-rate prices. In return for the cheap cars, the customers are letting the Japanese observe their driving behavior, studying how they use their F-150 and then the Japanese car companies are applying that data to design new vehicles that will directly replace Ford!"
Comment by eigenspace 2 days ago
Comment by petilon 2 days ago
That's not what they are observing, it is the behavior of the engines.
Comment by ForHackernews 1 day ago
fine, s/observing the way the Ford cars work
Comment by jumpkick 2 days ago
Comment by eigenspace 2 days ago
Comment by singingtoday 2 days ago
Comment by kalu 2 days ago
Comment by Danox 2 days ago
Comment by danny_codes 2 days ago
I can only assume this account is pure shilling for the closed-source AI labs.
Comment by pianopatrick 2 days ago