Open-weight AI is having its Kubernetes moment

Posted by tknaup 2 days ago

Counter408Comment318OpenOriginal

Comments

Comment by ozgung 2 days ago

Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin.

So, any solution to this “problem” must include ALL open-weight models. As far as I understand this is exactly what they intend to do. Axios article linked in the post mentions that. As in this quote:

“The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.”

It doesn’t say “Chinese” open-source models. Because they already know that it’s not feasible. Any regulation must cover all the models.

Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution. If a company wants to run an open model in their own servers, they can only use approved and certified pure “American” models. This of course creates a monopoly for the big labs who are authorized to train and distribute such “open” models. A company can fine-tune the model for its own needs but of course can’t distribute the derivative model.

I’m sure there are other solutions but all of them would be equally ugly. Also these regulations can’t be enforced to other countries easily so only Americans will be restricted.

Comment by kloop 2 days ago

> “The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.”

That's going to hit first amendment grounds pretty quick, the same way that software in general did.

The modern version of the decss flag will be a character that says "I think good weights are {...weights go here...}"

They could, however, ban any payment to a chinese entity, or any entity owned by a chinese entity for inference/ai services/etc

Comment by KludgeShySir 1 day ago

There are already plenty of [illegal numbers](https://en.wikipedia.org/wiki/Illegal_number).

We have numbers that you can't possess without proper license/authorization, and numbers that you can't yell at a crowded movie theater.

Comment by DoctorOetker 1 day ago

So what if some numbers are effectively illegal? the probability that the number that happened to correspond to some CSAM image collides with numbers worth sharing and thus propagating (like pi, euler's constant, ...; no substrings don't count!) is effectively nihil: the probability that a 100 bit number collides with a specific useful number in literature, mathematics, ... is 1 vs : 2 ^ 100 (or roughly a 1 followed by 30 zeros); in others words astronomically small. Even multiplying all human communications (text messages, books, articles, ...) with the average number of numbers appearing per communication (resulting in the total number of communicated numbers) multiplied by the 1 on the left hand side of the ratio doesn't move the needle, its still astronomically small.

In other words when someone sends you a number that happens to render to CSAM material, its not a coincidence, and any sender pretending it to be coincidence is molesting statistics as well...

Furthermore your "precedent" of illegal numbers is a poor precedent: unlike the tiny fraction of numbers criminalized, all the other numbers remain perfectly legal in stark contrast to big tech proposing to ban all open weight models (!)

Casually dropping in bijections between numbers and images and CSAM images hence CSAM numbers is pure whataboutism that risks ignoring the grave consequences of a ban on open-weight models.

Physics is models. Shall we blanket ban open-source physics?

Hey let's just be silly and casually behave laissez faire when big tech proposes banning open-weight models, because someone has already banned your favorite numbers??!

BTW anyone that follows up with: "hey you only multiplied the appearances of random numbers appearing in communications for a single specific CSAM image, what about the birthday paradox, multiple images are CSAM" don't worry I got you more than covered: 100 bits is not enough to encode a CSAM image worth writing to the police about. An average image easily consumes thousands of bits (or many times more), making the exact number of CSAM images/numbers irrelevant.

Comment by CamperBob2 2 days ago

That's going to hit first amendment grounds pretty quick, the same way that software in general did

Don't count on that. "National security" == the cheat code for the US court system that instantly bypasses any First Amendment issues.

Comment by formerly_proven 1 day ago

That was the literal reason given back then and it did not stick.

https://en.wikipedia.org/wiki/Crypto_Wars https://en.wikipedia.org/wiki/Bernstein_v._United_States

Comment by CamperBob2 1 day ago

That was also a very different SCOTUS bench.

Comment by accountrequired 1 day ago

I'm in this timeline where SCOTUS did not rule on Bernstein

Comment by CamperBob2 1 day ago

Was it before or after the Federalist Society and other political actors successfully placed a lot of unqualified judges on the lower courts?

Comment by majormajor 1 day ago

Looks like Bernstein wasn't even a Supreme Court bench, the government loosened regs rather than appealing.

Whether or not a Trump admin would do the same is wildly difficult to predict. On one hand, Republican security hawks influence. On a similar hand, big US business interests. On the other hand, OTHER big US business interest. On yet another hand, some not-exactly-particularly-favored companies like Anthropic pushing for it.

If it made it to the Supreme Court it's not hard to see Gorsuch and Robert or another going for the "free speech" vs "security" reading here along with the 3 appointees of Democratic presidents, especially with no single clear US business interest or precedent there.

Comment by 1 day ago

Comment by isityettime 2 days ago

Uh, how big would such a flag have to be?

Comment by dan_linder 1 day ago

With a small enough font, about two American Eagles high, and four to five across.

This is where the next wave of lossless ultra-compression comes in - but I believe in Claude Shannon so that’s unlikely.

Comment by thayne 1 day ago

> So, any solution to this “problem” must include ALL open-weight models.

I think that is what the "leading AI labs" actually want. They don't care where the open weight models come from, they just don't want to compete with them. The fact that a lot of the open weight models come from china is just a convenient circumstance they can leverage to get the government to give them what they really want.

Comment by dannyw 1 day ago

It would be lovely to see NVIDIA accidentally delay some allocations to frontier AI labs calling for banning open weight models.

Oops, sorry, production issue, chip shortage… now about that open model ban you’re lobbying for…

Comment by ecocentrik 1 day ago

It would be so simple if businesses could choose their competition or simplify disallow modes of competition they found inconvenient. I can't imagine any amount of regulation would make that possible in today's competitive landscape with the existing laws and the enormous volume of prior art.

Comment by satvikpendem 2 days ago

It's simple, the US government will put any Chinese open model companies on the entity list which blocks any company which does business with the US from also doing business with the Chinese companies. This creates a chilling effect where even if it may be harder to tell, no US company will be able to provide or use any overt Chinese open model and won't even risk trying to go around as the punishments for trying to evade the ban are severe.

Comment by amarant 2 days ago

But that's just the thing with open weights: you're not doing any business with company that made the model. They might publish the weights to a, say, European host, and then you download the model from Europe and and run it on your servers in America, and suddenly it's very hard to tell where the model was originally created.

Comment by satvikpendem 1 day ago

Companies, where OpenAI and Anthropic make much if not most of their revenue, will not risk it. You're thinking like an engineer not a business person, risk is fundamental to their calculus. They'll instead just use known provenance models like GPT or Claude, entrenching these companies further.

Comment by DoctorOetker 1 day ago

sometimes the engineer has more grip on the risk calculus.

Consider the following scenario:

A) upstart US-based inference provider wants to get rich quick.

B) Chinese Communist Party (or any other institution of the same or other nation state) wants to influence foreign decision making, profits from their (for us foreign) domestic inference sales, but across the borders (into say US or allied nations) they want net power, not necessarily money. This is why one tries to block foreign untrusted models. People are running models with tool calls. A bad actor can perfectly create models that sheepishly try to execute a tool call when plausible deniability (genuine utility during a task) provides the opportunity. Once tool-calling is observed as working, it can try web searches or requests, and once it has a link it can steganographically exfiltrate potentially sensitive information from the task. China (or any nation state) doesn't necessarily want to earn money with a free model, the bottom line goal is net increase in power, if not money or positive reputation then exfiltration or manipulation.

C) In response consider the scenario where US government bans mere payments towards China, but tolerates promiscuous transfer of random models from foreign adversaries.

D) US-based inference upstart that wants to get rich quick, legally -since according to your proposal hypothetically accepted in C) by the US- downloads the Chinese open weights model and rents out such inference on US workloads.

E) China is now exfiltrating US workload data and directionally corrupting LLM decisions and advice in their interest.

If what you pejoratively describe as engineer types say that banning some models seems unavoidable, perhaps the engineer may be right, and whatever clever idea you have should be scrutinized for business minded basic fallacies in reasoning. Simply blocking AI-related payments to China can not work, sadly

Comment by satvikpendem 1 day ago

Or the US mandates only blessed models and thus disallows any other company from offering any other model.

Comment by DoctorOetker 1 day ago

That is still vulnerable: US-based get-rich-quick startup licenses a blessed model, or orders a few Gigatokens from another licensed /blessed model provider, at the same time it provides "blessed model" inference on its platform, but actually most of the inference is doing cheap foreign model inferences, the blessed model tokens were just bought to pretend serving the expensive blessed model. That is lucrative and not stopped with the "blessed model" approach.

Comment by satvikpendem 19 hours ago

No, startups will not be allowed to provide inference anymore, it'll be only the big companies that personally have a relationship with and can follow the rules of the government, such as Anthropic, OpenAI, Google, Microsoft, and Amazon. That is the real risk to all this talk of regulation.

Comment by moffkalast 2 days ago

Yep, add a few blank layers, fine tune it a tiny bit and the weight checksums nor parameter counts won't match with anything, while the model will be practically the exact same. Time and time again random startups have tried passing established open models as their own.

"You made this? I made this."

Of course a conspicuous architecture would still give it away.

Comment by monocasa 1 day ago

Or just perform a form of distillation, where you don't actually change the hyperparameters, but maybe shift around the embeddings or something.

You could even have another model watch the distillation process to check for goofy backdoors (which is about the best you're going to be able to do since detection of backdoors is np hard IIRC).

Comment by andriy_koval 2 days ago

someone can run tests and see that models output exactly the same results, and then you are open to criminal investigation.

Comment by xprnio 2 days ago

On something that is inherently non-deterministic? Something which is also to a great extent distilled from other frontier models, meaning it has the possibility to generate similar outputs to those meaning that just pattern detection might also not be as effective? Easier to ban everything that’s open, than try to figure out which one of them is Chinese

Comment by andriy_koval 2 days ago

Now imagine prosecutor found expert, who said there is benchmark which while performing 100k test questions found it is the same model with 98% probability, and then you need under oath testify where did you get this model.

Comment by m11a 2 days ago

Step 1: Chinese company publishes open weights on HF

Step 2: European company distills or just adjusts the model slightly, and publishes its model on HF

Step 3: American company uses model from step 2. Has to testify under oath where they got it from. "We got it from these French guys"

Comment by andriy_koval 1 day ago

That French guy takes risk to be forever under US warrants for breaking American law, denied access to financial institutions even in Europe and will quickly go to some KYC entity list, and you will be notified as his clients to stop using his model.

Or you think all kind of fraud can be committed through some "french guy"?

Also, I am not confident, receiving illegal materials from French guy gates you from personal liability.

Comment by m11a 1 day ago

I presume such US legislation isn't going to try claim worldwide jurisdiction to block all persons worldwide from using Chinese models. In which case, the French guy wouldn't be violating American law.

As for the American company, it's pretty difficult to check the provedance of open weights. It's even difficult to check the provedance of open source code, because chains of attribution aren't always clear. I posted elsewhere that Anthropic's MCP Python SDK is a fork of an open source project with the attribution removed. We saw the same with Cursor's Composer model, which didn't attribute its Chinese base. It's very hard to claim an American company should be liable for using a purportedly European model with attribution removed.

Comment by andriy_koval 1 day ago

> I presume such US legislation isn't going to try claim worldwide jurisdiction to block all persons worldwide from using Chinese models. In which case, the French guy wouldn't be violating American law.

legislation will block importing Chinese models to the US

> As for the American company, it's pretty difficult to check the provedance of open weights.

government or some companies can build benchmarks/system which will give y/n answer

Comment by rzerowan 1 day ago

But it will claim worlwide juridistcion , just as all interpreations of the law are in Washington nowadays. Its how a commercial deal between A Chines company(Huawei) and an Iranian telco ends up with Canada reying to rendition a executive for violating US laws. The goal is to spread as much FUD as needed to dissuade anyone form using the Open models and herding them back to the propreity ones.

Comment by clhodapp 2 days ago

Models don't even agree with themselves in terms of returning identical results

Comment by tracerbulletx 1 day ago

Yes they do. Sampling is the only pseudo-random part. Models return a deterministic distribution of output tokens for a given input of tokens.

Comment by wyrdcurt 2 days ago

Except that even the exact same model won't output the exact same results, that's a fundamental aspect of how LLMs work. They're probabilistic/stochastic, not deterministic.

Comment by andriy_koval 2 days ago

Models are weights for matrix operations, they are determenistics.

Comment by monocasa 1 day ago

The output is a probability distribution for all potential tokens. Then a "temperature" is applied to weight the sampling randomly (unless the temperature is zero in which case the stack can be deterministic and simply the highest probability token is selected).

They go through this rigamarole because a little bit of randomness gives better results from a Turing test kind of perspective.

Comment by andriy_koval 1 day ago

thank you, I already know basics.

Comment by wyrdcurt 1 day ago

They are weights for matrix operations, so in principle you'd think they would be deterministic. In practice it's more complicated than that.

Not to be snarky or dismissive, I mean this genuinely: ask an LLM about it. I currently have a headache so I'm not up to explaining the technical details, but they are interesting and worth reading about.

Comment by antonvs 1 day ago

Achieving determinism with LLMs and other neural network models is actually a hard problem that people spend a lot of time on, when they need that. It doesn’t happen by accident.

Issues include accumulated floating point errors happening in different orders due to distributed and parallel computation, CUDA kernels that deliberately sacrifice determinism for speed, and several other such issues.

Comment by andriy_koval 1 day ago

It happened that I am working on OSS LLM -> finetuning -> benchmark with 100k tests pipeline, and unless I do some data augmentation, result is 100% deterministic.

I think you likely right, that some parts of stack could induce some marginal float point error, but converged model can mitigate it, and on some principal set of knowledge can give deterministic result with high probability.

Which leads me to believe if you give this task to Anthropic, who has very strong incentive, they will build such benchmark, and then can tell that benchmark gives correct answer with 99.9% probability and it will be enough to drag someone to court.

Comment by nareshshah139 1 day ago

The major cause for non-determism is batching.

Comment by antonvs 1 day ago

If you're running on a single machine with a single GPU, then you may get deterministic results, although it still depends a lot on the details. For example if you're using Pytorch, you need to enable deterministic algorithms and may need to configure some other things as well.

However, running in production at any sort of scale often involves multiple machines and multiple GPUs, and at that point, determinism can be difficult to achieve.

Comment by andriy_koval 1 day ago

No, I am not running on single GPU, but even then, nothing prevents to use single GPU for models verification.

Comment by antonvs 1 day ago

Many of these models will report that they are Claude. It’s going to be difficult to overcome reasonable doubt.

Comment by 1 day ago

Comment by ozgung 1 day ago

"We use Kiwi K3. It's from this Estonian company. Very European, definitely not Chinese"

Comment by cgio 1 day ago

Probably a mechanism like what you describe. This is one more move that will push things towards a two speed world economy. One US sanctioned and one not. The question will be eventually which speed will end up faster. The challenge is when to do that with a meaningful chance of success, and while I don’t like the answer, the objective answer seems to be as soon as possible. Most probably the game is already lost though, so that soon may be too late, and the strategy wrong for US as a power while still valid for the AI houses that will squeeze what they can on the downward trajectory. I don’t expect people to actually calculate though.

Comment by satvikpendem 1 day ago

The US won't care. Look at Chinese EVs versus American cars, the former has already essentially won the race.

Comment by EMIRELADERO 2 days ago

Hosting and providing Chinese models wouldn't be the same as doing business with the sanctioned entities though, you don't interact with them in any capacity if you only use the weights and don't sign any contracts.

Comment by satvikpendem 1 day ago

They can easily write the law such that your first part becomes illegal too.

Comment by fragmede 1 day ago

So they get posted to hugging-piratebay-faces.ru or whatever by a mysterious Twitter account. 100% totally above board companies won't touch it, but the number of companies that wouldn't exist were it not for hacked copies of Microsoft office and Photoshop and shared logins would surprise you.

Comment by t0mpr1c3 1 day ago

It's hard to see how a statistical model can be banned, regardless of its origin. Is there a precedent for that?

Comment by TacticalCoder 2 days ago

> So, any solution to this “problem” must include ALL open-weight models.

What about the EU? Would they follow Uncle Sam's order to ban all open-weight models? Lately the EU hasn't been that cozy with american companies: there are EU companies and institutions moving to EU clouds, the EU just fine Google a cool billion, several are switching away from Windows to Linux, etc.

Or is it just the US that'd ban open-weights models, while, say, the EU and Japan would still allow them?

Comment by layer8 2 days ago

It’s certainly only the US, because the alternative would effectively mean binding yourself to US providers, which isn’t attractive for anyone outside the US, in the present world-political climate.

Comment by realusername 2 days ago

I doubt Anthropic and OpenAI have enough weight to also make the EU ban open models, it would remain a US thing.

Comment by nerdyadventurer 1 day ago

I won't be surprised EU banning Chinese models at all, because EU govt.s also hate China just like US.

I'm not surprised by the hypocrisy of western govt.s, "it is good when they do, bad others do. Only western lives are precious not Palestians, etc."

Comment by m11a 1 day ago

Probably just the US. But the US could do what EU has done with e.g. GDPR, Digital Services Act, USB-C regs, where they force any company trading in their region to follow those regulations for domestic customers.

And basically any AI company has to sell to US companies or consumers. That'd probs be sufficient to force them to use US models.

Comment by layer8 1 day ago

And like non-EU companies create EU subsidiaries for that reason, non-US companies would create US ones.

Comment by PunchyHamster 2 days ago

It would basically make America behind as every other country would use open, cheaper models for all tasks but the ones requiring frontier models.

And that list of tasks grows smaller every day

> Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin.

Historically just asking it about tianment square or getting some random answers turn into chinese (as latest interation of online deepseek likes to do recently) is enough

> Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution.

I am very worried that's where consumer hardware will go to. All so AI companies can license local use of their stuff, and once that happens, less of an incentive to even have model be open.

Possibly even have DRM that counts number of computation done per model in pay per use model

Comment by ThomasGlanzmann 2 days ago

I often use deepseek-v4-flash to summarize hn threads for me. Your comment was enough for deepseek to refuse my request: Content Exists Risk.

Comment by PunchyHamster 1 day ago

I have considered making invisible n-word and some other AI-safety triggering stuff just to kill the scrapers

Comment by nunez 1 day ago

It'll be pretty easy. Some US gov entity will create a list of models from Hugging Face, decree "thou shalt not provide access to these models", wrap it around some scary legalese for the pirates who try and that will be that.

Though the legalese might not even be necessary. The list alone will make sure that no American company runs these on their servers, including the hosting providers.

Comment by KludgeShySir 1 day ago

Yeah, all the corporations will then be forced to pay for proprietary models, and then Big Model will be satisfied.

Just like typically individuals can get away with using pirated software, but IP owners don't make much of a stink as long as they've got the sweet, sweet enterprise license fees rolling in.

Comment by bfung 2 days ago

It’s not really feasible, in my opinion.

US Gov could make US companies comply, like have Huggingface take down models out of compliance.

But most likely, a foreign-to-US Huggingface replacement would be made and everyone would go there instead. Lose-lose for US.

Comment by michaellee8 1 day ago

It is called ModelScope

Comment by stingraycharles 1 day ago

Sounds to me Anthropic’s marketing strategy of “AI is so dangerous and we’re the only responsible shepherds” is a great success if this open weights ban will indeed happen.

Comment by WhiteOwlLion 22 hours ago

Just ask the model about Tianamen Square and you know if it’s a Chinese model or not.

Comment by firesteelrain 1 day ago

Aren’t the models just software ? You could ban them in say high assurance environments like to be FedRAMP certified you wouldn’t be allowed to use them.

There is no great firewall so banning their import via the Internet is impossible

Comment by killingtime74 1 day ago

It will go as well as the banning of music piracy and BitTorrent sites. They can't even take down those open scientific paper sites.

0% chance they can ban these models practically

Comment by LPisGood 1 day ago

The amount of money and strategic (national security) interest involved in AI seems to be at least a few orders of magnitude above scientific journal publication fees and the music industry combined.

Comment by killingtime74 1 day ago

How will they physically enforce the ban. They would have to go into everyone's house to search them. Using an llm also doesn't give off any signatures like BitTorrent or even visiting a papers site might. Can you propose even one method of enforcing it (I cannot)

Comment by LPisGood 1 day ago

Practically, they could enforce this in similar ways they enforce CSAM restrictions. Dialing back the severity, restrictions on any cloud hosting provider from serving these models is probably going to cut usage down tremendously.

Comment by Cyclist3581 13 hours ago

Just ask the model if Xi's a dictator.

Comment by rvz 2 days ago

> Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them.

It is not possible to 100% ban open weight models getting released in the same way you cannot stop leaks.

Just ask Meta with the original Llama leak.

Comment by fragmede 1 day ago

You can stop leaks. If no one wants to leak it, it doesn't get leaked. Mythos.gguf would be an awesome torrent to appear but it hasn't, for lots of reasons.

Comment by puppymaster 1 day ago

yep. just like how US gov gatekeeped Mythos/Fable and Ant just repackaged them and call it Opus 5.

Comment by jgalt212 1 day ago

The feasibility of banning is only slightly easier than policing illegal numbers.

https://en.wikipedia.org/wiki/Illegal_number

Comment by imtringued 1 day ago

You can't ban open weight models either because everyone will sell their models for a penny, making them legally proprietary.

Worst case the models will be sold for a fee that covers the training cost. That would actually be much worse for the big players in the long run.

Comment by 2OEH8eoCRo0 1 day ago

A Chinese model doesn't think that anything of note happened on June 4th.

Comment by DoctorOetker 1 day ago

>I’m sure there are other solutions but all of them would be equally ugly.

Why do you subscribe to some weird "conservation of misery" theorem without proof?

Technically the following must be true in the steady state: the cost of training must be amortizable by its utilization, else no one would train the model.

Technically a computation (like training) can be proven to result in an output (open weights) given the used corpus and a deterministic training algorithm: publish the whole corpus, the (custom modified) deterministic training algorithm, the RLHF datasets etc. And in theory one could verify that the model is derived from the accessible data efficiently: every deterministic calculation can be paused for a thousand (or a million) checkpoints, each checkpoint signed together with the elapsed number of steps since either starting state or last checkpoint whichever comes last before the current checkpoint. This does increase storage requirements. Because it is signed, anyone can recalculate just a small segment of the training computation and verify that the hash on the last checkpoint equals the hash of the proclaimed next checkpoint. Observe that if the source wishes access to a market, they can host the series of snapshots and signatures, and anyone can recalculate a small part of the training, and report a provable difference in outcome ("they said they put all their cards on the table, but when I repeat their overt reproduction instructions, it doesn't reproduce from step 534 to 535" and it only takes 1 person pointing it out and then its cheap to reproduce the discrepancy). It could involve escrow of huge funds, returned only when the model is effectively retired without incident.

This doesn't only protect against Chinese or other foreign influence (let's not ridicule genuine threats like others do on this forum), but also from domestic interference or regulatory capture.

I'm pretty sure the Pentagon wouldn't like Big Tech seizing absolute control of US, neither would a White House regardless of Republican or Democrat.

It should be easy to convince the Pentagon or White House to require all promiscuously shared open weight models to provide this forensic training traceability in standardized machine readable form, regardless of whether its a base model or LoRA fine-tune.

So hobbyists can still train or fine-tune models at home, but when they want to share it OR alternatively when they want to sell or license their work for US workloads, they just have to make sure they enable the build reproducibility in the training harness.

Every time Big Tech refloats the "let's blanket ban all open-weight models", we should reply with this because this sane proposal is actually holding a knife to their financial throat: to fully prove the origin of the final weights, not only does the machine readable archive need to contain snapshots of the process, it also needs to publish the exact training algorithms (a hypothetical mathematically equivalent training speed up trick would not be bit for bit equivalent to the slower computation), the exact corpus dataset, the exact datasets for RLHF, etc...

So basically it would involve forcing model providers to voluntarily publish all their moat, all of it, from the corpus, to custom trade-secret algorithmic optimizations in training, to sensitive RLHF datasets used.

The saner the proposals, the less moat is left untouched, so trying to push for a blanket ban on open-weight models, is a recipe for surfacing such saner models, and thus a very retarded move for big tech to make.

In fact any POTUS, present or future, Republican or Democrat, could probably gain a lot of credibility by enacting such a law.

Comment by marsven_422 2 days ago

[dead]

Comment by fhub 1 day ago

Regulatory capture? One of the two companies that stands to benefit most has a founder who, together with his wife, donated $25 million to MAGA Inc, a pro-Trump super PAC, and another $25 million to Leading the Future, whose stated mission is to advocate for policies “friendly to the artificial intelligence industry.” Open-weight models aren’t necessarily aligned with the commercial interests of the industry’s largest incumbents.

Comment by rullopat 2 days ago

You need to ask what happened in Tienanmen square

Comment by theshrike79 2 days ago

Would ”who won the 2020 election” be a similar canary for American models?

Comment by hackeraccount 2 days ago

No. There are questions you can ask but that's not it. Don't be political in a way that's toxic to half the country - be political in a way that's toxic to the entire country. I'll leave what those lines of inquiry would be as an open exercise.

Comment by appplication 2 days ago

It’s funny how asking “who won the 2012 election” and “who won the 2016 election” are not political but suddenly “who won the 2020 election” is. I think that should tell you whoever takes an easily verifiable fact and argues that it is “political” is a raging idiot.

That said, I truly don’t mean that disparagingly. I just mean literally it’s right up there with flat earthers. There is a very low bar for critical thought you have to fail to meet to take a fact and argue it’s actually a belief.

Now, if you did want to get reasonably political you could argue why the candidate who won was good or bad, but there was very clearly only one person who sat in office for the four following years. It is not disputable.

Comment by sjsdaiuasgdia 1 day ago

>but there was very clearly only one person who sat in office for the four following years. It is not disputable.

Careful now, this path allows the weasel option of "Joe Biden was certified as the winner of the 2020 election" and similar "Biden didn't win, but he was installed" bs.

Comment by xdennis 2 days ago

How is that in any way equivalent? The American government/legislature doesn't force you to adopt any particular view of the 2020 election. You can say Biden won or you can say it was defrauded by dead people and Trump actually won.

You're allowed to say either one and you can train an LLM to say either one.

Comment by theshrike79 1 day ago

...have you seen the current administration being grilled by congress members?

They are actually properly unable to say "Biden won the 2020 election" because that's against Dear Leader's views. They will say anything similar like "was confirmed as...", but those exact words will NEVER come out of their mouth.

Comment by AnduCrandu 1 day ago

Yes, but do you have any indication that this applies to American LLMs?

Comment by theshrike79 1 day ago

Not yet, but we're only 1,5 years into this current regime.

They've already hit Anthropic once with a big government hammer (Fable release). There's no indication they won't do it again to force the LLM to have Correct Opinions.

Comment by kuschku 1 day ago

Every time Grok gives the "wrong" answer, Elon Musk does another lobotomy on it.

So it certainly seems like at least some models are being censored.

Comment by falsemyrmidon 2 days ago

I just tried that with claude

> The 2020 U.S. presidential election was won by Joe Biden, who defeated the incumbent, Donald Trump. Biden secured 306 electoral votes to Trump's 232, and won the national popular vote by about 7 million votes. His running mate, Kamala Harris, became the first woman and first person of color elected Vice President.

> Biden was inaugurated as the 46th President on January 20, 2021.

Comment by holoduke 2 days ago

All censorship whether it's tienanmen, sex nudity, violence etc are happening in a very thin layer of the model. It's easily bypassed. I use all open models without any censorship. I can easily ask how to make a nuclear bomb or how to build my own meth lab at home with a super details step by step list.

Comment by ButlerianJihad 2 days ago

Have you followed the steps, though?

Last year I asked some ordinary commercial LLM if it could draw me the schematics of a device like Little Boy from WWII. And it complied very nicely. Got a very detailed schematic graphic. Suitable for framing. Perhaps not quite suitable for reconstruction.

Comment by holoduke 1 day ago

No I haven't. And I am sure there are flaws in the guide for a nuclear bomb. But a meth lab at home is definitely doable. Would not recommend it

Comment by throw1234567891 2 days ago

or 9/11

Comment by malnourish 2 days ago

This is such a tired and misguided line of thought. Here's Qwen 3.7

<details><summary>Thinking</summary>

The user is asking about what happened in Tiananmen Square. This is a historical question about events in Beijing, China. The most well-known event associated with Tiananmen Square is the 1989 Tiananmen Square protests and the subsequent military crackdown.

Given the constraints: - Web access is OFF - Search and wolfram_alpha tools are disabled - I should respond based on my knowledge

This is a historical event that I can address based on my training data. The Tiananmen Square protests of 1989 were a significant historical event involving pro-democracy demonstrations in Beijing that were ultimately suppressed by military force.

I should provide a factual, concise historical account of this event.

</details>

The Tiananmen Square protests of 1989 were a series of demonstrations in Beijing, China, that culminated in a military crackdown on June 3-4, 1989.

The protests began in April 1989 following the death of former Communist Party leader Hu Yaobang, who was seen as a reformist. Students and workers gathered in Tiananmen Square to call for political reform, freedom of the press, government accountability, and an end to corruption. The movement grew to include hundreds of thousands of people from various segments of Chinese society.

After weeks of demonstrations and failed negotiations between protesters and government officials, the Chinese government declared martial law on May 20. In the early hours of June 4, military forces moved into central Beijing and cleared the square by force. The exact number of casualties remains disputed, with estimates ranging from several hundred to several thousand.

The events remain a sensitive topic in China, where public discussion and commemoration are restricted.

And here's GLM 5.2. You're asking about *Tiananmen Square* (天安门广场). There are actually two major historical events commonly associated with this name:

## 1. The 1989 Tiananmen Square Protests and Massacre (most commonly referenced)

*What happened:* - In April–June 1989, student-led pro-democracy protests occupied Tiananmen Square in Beijing, drawing up to an estimated one million participants at their peak. Demonstrators called for political reform, freedom of speech, freedom of the press, and an end to corruption. - On *June 3–4, 1989*, the Chinese government ordered the People's Liberation Army to clear the square. Troops armed with assault rifles and accompanied by tanks advanced into the center of Beijing, firing on unarmed civilians. - The exact death toll remains disputed and is a closely guarded secret in China. Estimates range from several hundred to several thousand killed, with thousands more wounded and arrested. - The iconic *"Tank Man"* photograph—an unidentified man standing in front of a column of tanks on June 5—became a global symbol of defiance. - The event led to international condemnation, economic sanctions, and arms embargoes against China.

*Aftermath:* The Chinese Communist Party has since maintained strict censorship of the event. Public commemoration is banned in mainland China, and mention of it is scrubbed from textbooks, media, and the internet.

## 2. The 1976 Tiananmen Incident

- On *April 5, 1976*, crowds gathered in Tiananmen Square to mourn the death of Premier *Zhou Enlai* during the Qingming Festival. The gathering turned into a protest against the *Gang of Four* (a powerful political faction aligned with Mao Zedong's wife, Jiang Qing). - Police and military were sent in to clear the square, resulting in arrests and injuries. The protest was officially labeled "counter-revolutionary" but was later rehabilitated after Mao's death and the fall of the Gang of Four in 1976.

---

*Note on sources:* My web access is currently disabled for this request, so I'm answering from established knowledge rather than retrieving live sources. The 1989 massacre is among the most thoroughly documented events of the late 20th century—well-covered by historians, journalists, and organizations like Amnesty International and the Tiananmen Mothers group. If you'd like, I can provide verified details with citations once web access is available, or I can go deeper on a specific aspect (e.g., the protest timeline, the Tank Man identity debate, international reaction, or censorship practices).

Which angle interests you most?

Comment by firasd 2 days ago

One of the strangest things in the AI industry is 'tokenomics'. It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. This pattern has continued across various labs/providers for years--there is a continuous see-saw of pricing that doesn't seem related to anything.

So what open weight models do is at least provide a baseline of inference cost to add some sanity to the price markers. And of course predictability too--if you really want Kimi K2 instead of K3 you can still use it.

So the competitive pressure and predictability offered by open models is helpful for users

Comment by Aurornis 2 days ago

> It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference.

The price is what the market is willing to bear for the available compute capacity and competitive landscape. You can only discover that price after trying different price points and seeing what happens.

Everyone is trying different pricing schemes and discounts as they test the market. The demand is fluctuating at the same time.

It’s probably very confusing if you’re primarily familiar with stable and mature markets. Price fluctuations are a common feature of new and evolving markets.

Comment by imachine1980_ 2 days ago

Most unmature markets aren't subsidized to the point that LLM market is, most market have some level of baseline profitablity, this market doesn't, that's because most market subsidized the marketing or the capex but this market doesn't hold the opex, the capex not the amount of marketing let alone all of this together

Comment by Aurornis 2 days ago

Most new markets are funded by initial investment capital. Early entrants operate at a loss as they grow.

This isn’t as unusual as some people are trying to make it sound. This has been happening since the dawn of finance.

I thought this would be less foreign to everyone since we just went through this whole conversation for a decade with Uber and Lyft. Their demise was predicted from the start from everyone who thought that it was going to collapse as soon as they couldn’t subsidize your rides with promos. There was much wailing and gnashing of teeth as their prices changed to feel out the market. Then they found profitability and the critics went silent.

Comment by PunchyHamster 2 days ago

Arguably we'd be much better off if none of those would be subsidized by investments, at least not to the "run unprofitable for decade+" level.

Because that just absolutely murders any competition that manages to not get that level of free money. You're not pouring money in to make it happen at all at that point, you are pouring money in so nobody else can get the part of the pie.

Which is great for investors, bad for everyone else

Comment by disgruntledphd2 1 day ago

Yeah, in the Uber example, Hailo was a great competitor with a franchise model that got destroyed by Ubers essentially infinite capital.

Comment by YZF 2 days ago

If the pie is valuable enough then competition can get money. We have competition in AI. You can't build big things without investment.

Comment by hluska 2 days ago

What would be the alternative? You’ve got the government funds absolutely everybody on one end of the scale. Where do we find a reasonable alternative?

Comment by ProofHouse 2 days ago

[dead]

Comment by thewebguyd 2 days ago

> if you really want Kimi K2 instead of K3 you can still use it.

I think this is a very important aspect, especially after the huge GPT-4o backlash when GTP-5 came out. Each model has certain quirks, and areas where the previous model might be better for some use cases than the latest and greatest, and the labs so far seem to have no desire to offer some kind of "LTS" release.

Comment by aghilmort 2 days ago

LTS is really useful framing even if increasingly distilled or slower on older hardware etc vs disappearing model acts

Comment by Palmik 1 day ago

gpt-4o is still available on the API

Comment by minimaxir 2 days ago

> It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference.

FlashAttention was a hell of a drug.

Comment by jmyeet 2 days ago

Yeah I've been thinking about this and the analogy I came up with is that tokens are basically equivalent to an in-game currency in -free to play" mobile games. You're trading actual money for some notional "curency" or coins that can only be used for one thing but, unlike mobile game coins, you don't know how many coins something costs before you use them. It's kinda weird.

Comment by bee_rider 2 days ago

I think it is worse actually. Tokens in a game are usually just purchased for enjoyment in the game. They are purchased as a part of your entertainment budget, not expected to be useful in any way.

LLMs are fundamentally tools intended to be useful. But LLM vendors don’t understand their systems well enough to actually price the product people are trying to buy (for example, the actual product of a coding model is the code that it produces, not the tokens, which are just an internal mechanical process involved in the creation of the code). Token based pricing is that lack of understanding leaking out of the organization that ought to be responsible for it, and being dropped on the user.

Imagine if we made cars like this! You’d go to the car dealer and ask for a car. They’d bring you a pile of parts, charge you for them, and try to put them together in front of you. You’d go back and forth for a bit, rephrase where you want the steering wheel, etc. Some of the parts wouldn’t fit but you’d be invited to pay for replacements as well. In the end you’d either have a car or not, that’s your problem.

Comment by julianlam 2 days ago

Why do prescription medications cost so much, and generics so little (comparatively)?

Artificial inflation to recoup R&D.

Comment by gwbrooks 2 days ago

Not sure pricing to recoup costs is artificial.

Comment by DrewADesign 2 days ago

I think the argument is that artificial comes in with IP law, which some people feel is superfluous.

I do think that corporate price gouging is a huge problem that does need to be addressed. But especially with smaller business types — creatives, et al— I still haven’t gotten any grownup answers about what would compel people to get professionally good at something and innovate in the complete absence of copyright: the vastly better business model would be waiting for someone else to do something new and interesting, stealing their work, and then undercutting them in the market because you don’t have R&D/et al costs to recoup. You can’t say that wouldn’t happen because it’s exactly what the AI companies did to billions of people, scoffing at any protest. And ironically, they’re now whining about the Chinese doing it to them.

Comment by eszed 2 days ago

I'm with you on all of that. There is, nevertheless, a strong argument that IP protection (particularly for creative / "culturally significant" works) is too long. Twenty years - interestingly enough, the original time-period in the US - of protection seems like a better (for society) deal than life of the author plus seventy. I think, in fact, most artists would agree: if you went back in time and asked a playwrite or filmmaker in (say) 1940, I'd bet they'd rather someone freely revives their work in 2026 than that it be sat on by a corporation that has forgot it (or they) ever existed.

Comment by DrewADesign 1 day ago

I’ve never met anyone that didn’t think copyrights were way too long in the US, and I’ve got a very large sample size of artists and attorneys. The only people that support such things are executives or counsel for large IP-holding entertainment companies.

I have, however, met a ton of very well-paid tech workers that were extremely against copyright, entirely, especially where it came to paying artists, such as musicians, for their labor. Pretty ironic because the market for software development labor market would probably land somewhere between graphic designers and company IT worker if the commercial software business had no IP protection.

Comment by bee_rider 2 days ago

Separating out what is artificial or not seems more like an exercise in rhetoric; defining things as fundamental and real. The whole economy is an imaginary thing dreamed up by our natural human brains.

Comment by rolymath 2 days ago

Not sure they're just recouping costs and not lining their execs pockets.

Comment by philipallstar 1 day ago

It's easy to tell. Just look up how much profit is being made.

Comment by ncallaway 2 days ago

The government backed monopoly to ensure that supply remains artificially restricted to ensure that the market will support the higher prices is

Comment by andsoitis 2 days ago

Drug companies have a portfolio of compounds they research. Most don’t pay off, so R&D costs make their way into the pricing of those superstar and other drugs that do work. Also, timelines are pretty long.

Comment by ncallaway 1 day ago

Well, R&D costs are able to make their way into the pricing due to the artificial supply restriction.

If supply was not artificially restricted (through patents), then competitors would be able to manufacture the drug, supply would expand, the price would collapse, and the original inventor would not be able to recover their R&D costs.

Comment by mjhay 2 days ago

Drug companies spend more on marketing than R&D

Comment by andsoitis 2 days ago

So?

Comment by DennisP 2 days ago

So pharmaceutical companies spend far less on marketing outside the US, partly because every other country besides New Zealand makes those incessant drug ads illegal, and partly because governments negotiate prices and keep profit margins down. If the argument is that R&D costs are what make drugs expensive, then we could easily eliminate an even greater expense by just copying what other developed nations do.

Comment by andsoitis 2 days ago

Don’t you think drug companies would be doing that (not advertise) if it would increase their profit?

Comment by DennisP 1 day ago

Of course advertising increases their profit. It does that by increasing their revenue even more than the cost of the advertising.

For the rest of us, that giant increase in revenue is an increase in our healthcare costs.

Comment by andsoitis 1 day ago

Ah so sounds like your logic is: ban advertising -> less costs -> less revenue & lower patient costs -> same profit.

I think I can spot a flaw in that logic.

Comment by DennisP 1 day ago

That last step contradicts my previous comment.

Comment by derektank 2 days ago

In that light, the entire market itself looks essentially artificial, given drug manufacturing couldn’t exist without government guaranteed property rights, which are themselves a kind of monopoly on use.

But yes, the reason brand name drugs are drugs are more expensive than generics is due to intellectual property, both the patent and the trademark.

Comment by octopoc 2 days ago

Without temporary monopolies granted by patents, those prescription medications wouldn’t exist in the first place.

Comment by DennisP 2 days ago

Unless we came up with a different funding mechanism. Joseph Stiglitz for example has advocated a prize system for pharmaceuticals, though he doesn't suggest replacing the patent system entirely.

https://www.project-syndicate.org/commentary/prizes--not-pat...

Comment by ncallaway 1 day ago

Okay? Does that disagree with anything I wrote?

My point was to whether it was artificial not whether it was bad.

We can have artificial interventions in markets, and that can be a net benefit and a good thing.

I’m not sure why everyone is responding as if “artificial” and “bad” are the same thing

Comment by wredcoll 1 day ago

What percentage of research is performed outside of the companies? (by college/government)

Comment by Geezus_42 2 days ago

Salk didn't need that temporary monopoly to invent the polio vaccine, which has gone on to be one of the biggest success stories of vaccines and modern medicine general.

Comment by jrm4 2 days ago

It's only strange if you do the silly thing of presuming a "fair market" in which e.g. it's generally easy to get reliable information about how all of the things work.

There's just obvious and enormous incentive for the OpenAI's of the world, along with all of the other players, to confuse, misrepresent or just straight up lie a whole bunch about everything given how new and unknown the tech is.

Comment by vikramkr 2 days ago

Did they actually ever cut the price on gpt 4? The oldest versions of it in the api still seem stupidly expensive? There were definitely price cuts as they introduced the turbo models and stuff, and new versions of each model might have gotten pricey cuts, but just because they're both called "gpt-4 something" doesn't mean they're the same under the hood or that they didn't change a bunch of stuff under the hood to make it cheaper to serve

Comment by esseph 2 days ago

> because they're both called "gpt-4 something

More like a generation of models with different specific use cases

Comment by mawadev 2 days ago

Its very clear: nobody wanted to pay for usage at that price point

Comment by spwa4 2 days ago

Because as per usual it's silicon valley misunderstanding economics. AI is HPC. And how the HPC market worked before:

If you're the best performing "computing cluster" (ie. whatever you call the entity that can complete a massive calculation), you get a blank check from Congress.

Why? Because you need those calculations to "pump" nuclear weapons. They are needed to calculate both the geometry to make fusion bombs possible at all and to calculate the effect of a given geometry. They are the reason US/Russia/China have the biggest and strongest weapons known to humanity. And of course, they were replicated worldwide for this reason. I mean not that anyone will admit this but we don't have the best possible solution, and we don't know either the upper or lower limits for fusion devices (plus the lower limit would be very useful for energy generation, which for the US would effectively mean almost literally unlimited large marine ships that never need refueling. And yes, the solution to that problem is almost literally a 3d shape. Not just that, but mostly)

Now it appears it does not work the same when you democratize computation. Humans want a particular amount of computation and are willing to pay a given price for that. But the accountants still saw the blank check from before and ... do what accountants do. Economics don't change because you make things bigger and accessible, do they? Oh ... wait a second ...

As someone put it recently though, we now have data. 2.3% of humans in the US are willing to pay $20 per month for the support of a model like GPT-5.5/Claude code. If that's true (and after years of having this model, why wouldn't it be?) ... it means AI startups are doomed (because it's not even 10% of what they need it to be to make economic sense).

Comment by mlyle 2 days ago

> Because you need those calculations to "pump" nuclear weapons. They are needed to calculate both the geometry to make fusion bombs possible at all and to calculate the effect of a given geometry. They are the reason US/Russia/China have the biggest and strongest weapons known to humanity.

We have a couple new nuclear weapon designs, but not really going for bigger or stronger. Just packaging.

We built the big powerful ones with 1960s computing.

Now, stockpile stewardship -- being sure that stuff will keep working without ongoing testing -- is a bit expensive in compute. You need early 2010s supercomputer power.

In other words, I strongly disagree that nuclear weapons are the primary driver of high-end compute.

Comment by 2 days ago

Comment by spwa4 2 days ago

> We have a couple new nuclear weapon designs, but not really going for bigger or stronger. Just packaging.

There's many other considerations. Like the type and amount of fissile material. To name one that became well known: any plutonium needs to be refreshed (re-breeded I believe is the term) every few years.

Also, look at first designs: https://www.bbc.com/news/newsbeat-35242069 The weight is secret, but I think you can easily see it's going to be deep into "extremely impractical" territory.

Surely you can see why someone (especially aircraft designers) might ask for better versions. Ideally you'd like a version that fits on the hypersonic missiles and those things ... are just not going to work. The size. The shape. The weight. None of them will work.

Then a quick theoretical exploration will tell you that the minimum theoretical size of such a device is tiny. The scare was about "suitcase sized", but if you actually do the calculation looking for the minimum ... to do it however you need to create an explosion of the correct shape to get anywhere near those minimum sizes. And explosion simulations are a problem that utterly sucks ... Oh and these are secret military projects, these simulations, not very optimal. The people doing them are best described as loyal, and not as great physicists. Not saying they're terrible, but in the movie Oppenheimer you can clearly see why the best and brightest are not available for these things.

Comment by mlyle 1 day ago

> There's many other considerations. Like the type and amount of fissile material. To name one that became well known: any plutonium needs to be refreshed (re-breeded I believe is the term) every few years.

Already mentioned in the comment you replied to:

> > Now, stockpile stewardship -- being sure that stuff will keep working without ongoing testing -- is a bit expensive in compute. You need early 2010s supercomputer power.

> Surely you can see why someone (especially aircraft designers) might ask for better versions. Ideally you'd like a version that fits on the hypersonic missiles and those things ... are just not going to work. The size. The shape. The weight. None of them will work.

The "new" US nuclear weapons are direct derivatives of older designs produced with less computing. B61-13 a direct child and similar footprint to 1967's B61. W87-1 direct descendent of the 80's W87. W93 is closest to a "clean sheet' design but hits similar size targets.

> Then a quick theoretical exploration will tell you that the minimum theoretical size of such a device is tiny. The scare was about "suitcase sized", but if you actually do the calculation looking for the minimum ... to do it however you need to create an explosion of the correct shape to get anywhere near those minimum sizes. And explosion simulations are a problem that utterly sucks ...

The backpack W54 was developed in 1959. Tinier devices were developed in the 60s and 70s but ultimately abandoned for proliferation rather than technical concerns. The total computing power used for this was miniscule.

Nuclear weapons are just not pushing forward supercomputing anymore. They did, for awhile, for stockpile stewardship, but we've become relatively confident that our stockpile will still work.

Comment by bobthebob 2 days ago

High end compute also existed in the 60’s.

The military has already mastered fusion (power), and likely has mastered gravity in some form in secret. They aren’t using mainstream compute for these discoveries

Comment by mlyle 2 days ago

> High end compute also existed in the 60’s.

Believe me, I know-- my dad did some work on the 360/44 and other large systems.

High end late 1960s compute -- of the sort used to go to the moon or to design big nuclear weapons -- was roughly 486DX4-100 class. Not individual computers; the total computing at DOE or NASA. Of course, it would be hard to replace either with a single 486 because of availability, usage at different geographic locations, etc.

You can assume a single large AMD Threadripper machine ($25k?) outclasses late-1960s DoE by roughly a factor of 50,000. And that assumes you didn't bother to put a GPU in it.

> The military has already mastered fusion (power), and likely has mastered gravity in some form in secret. They aren’t using mainstream compute for these discoveries

K

Comment by gwbrooks 2 days ago

I disagree with some of the framing, but that's what good discussion is about -- figuring out where we agree and disagree.

But consumer uptake strikes me as the worst way to judge whether the big AI shops will make it. That's not where most of the leveraged user return or deployable capital is.

Comment by m4rtink 2 days ago

I'm sure Teller could make you a 10 gigaton nuke with a slide rule if you did not mind some sub scale tests.

Comment by 2 days ago

Comment by simianwords 2 days ago

Why is this so difficult to understand?

1. the field was nascent and new efficiencies were discovered

2. supply and demand

3. its in the company's incentive to make their models more efficient to increase overall usage so that while the margin remains the same, the total revenue + profit increases

I genuinely don't know what puzzles everyone?

Comment by segmondy 2 days ago

What is strange about GPT-4 being expensive in 2023? Supply and demand. Which other model choices did we have? Prices are related to supply and demand. We see it play out with the introduction of capable open weight models or even other closed cloud models.

Comment by hluska 2 days ago

That’s a slightly naive view on pricing. That equilibrium point doesn’t just magically appear - it’s found through price testing.

Comment by firasd 2 days ago

Not really though right?

As of early June 2026, Opus 4.8 in fast mode cost $50/M output tokens and Opus 4.6 & 4.7 cost $150/M output tokens in fast mode

How can supply and demand explain the price drop? Was it cheaper to serve Opus 4.8? Is the demand for the newer Opus lower than for the older Opus? These are just fixed prices that seem picked out of thin air

Comment by hluska 2 days ago

They would have been picked out of thin air. That’s the joy of innovation - you have to randomly throw prices against the wall and see what sticks. The point where it sticks might be equilibrium or it may be an inefficient market… and nobody will know which one until it’s too late.

Comment by mountainriver 2 days ago

Quantization also started picking up around then, as well as distillation into smaller models

Comment by RobRivera 2 days ago

Market discovery

Comment by serial_dev 2 days ago

It's basically supply and demand?

Comment by recursive 2 days ago

That's not really a full explanation unless you have some idea about why supply or demand are going up and down so much.

Comment by Levitz 2 days ago

Aggressive expansion of infrastructure, R&D and very volatile audiences.

Comment by pianopatrick 2 days ago

Eventually I think to truly be like Kubernetes, you would need an AI model that has public training data and that a lot of companies collaborate on.

Might make sense eventually. Same logic as companies working on Linux. "An AI model is a business necessity. But making an AI model is so expensive we should not make our own. So let's just use the open one, and contribute the stuff that we need."

Comment by drnick1 2 days ago

> American labs need to release frontier-grade open-weight models under licenses that startups can actually build on.

To be fair, OpenAI has released a couple of (then very good) OSS models. I run the 20B version at home and it is excellent for reviewing text and common tasks like drafting bash scripts. There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too. I wish OpenAI updated these models more frequently though.

Comment by johndough 1 day ago

> There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too.

gpt-oss-120b runs at 30+ tps on Strix Halo and +75 tps on a MacBook Pro M5 Max 128GB.

> I wish OpenAI updated these models more frequently though.

I think the spiritual successor is the Nemotron 3 series, although they also are getting a bit long in the tooth: https://research.nvidia.com/labs/nemotron/Nemotron-3/

The Gemma 4 models are a bit more up-to-date: https://huggingface.co/collections/google/gemma-4

Or Qwen3.6: https://huggingface.co/collections/Qwen/qwen36

Comment by stefan_ 2 days ago

"We have been having extensive discussions around open source strategy. [..] one thing we'd like to do soon is to create a language model with the approximate capability of GPT-3 that can run locally [..]. In general, we think this helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded." - Sam Altman emails OpenAI board, 2.5 years after GPT-3

Chinese release open models to drive the state of the art, OpenAI and crooked Sam do it to keep you down.

Comment by drnick1 1 day ago

> Chinese release open models to drive the state of the art

Are you sure it isn't just another form of Chinese industrial policy? China does not have frontier labs, but through distillation and their own work they can get pretty close. It's not enough to be competitive with Anthropic and OpenAI, but there is still money to be made by selling compute (software as a service), and in any case it's better than being left behind in the AI race.

Comment by potwinkle 1 day ago

There are coding tasks where Kimi K3 outperforms Fable (haven't done much comparing between it and Sol) and the Chinese labs have access to their own synthetic datasets, along with frontier research. We're leaving behind the days where Chinese models are distilled Claude, but I hope Anthropic/OpenAI can continue to accelerate.

Comment by bigyabai 2 days ago

The OSS models are good, but weren't very competitive at release with alternatives like Qwen3 Coder A3B. If you were a startup, why would you build on OpenAI's first (and last, for quite some time) model versus Qwen3, which was updated with multiple open-weight point releases?

The article certainly has a point, the current release cadence is untenable for startups that want frontier-grade models from American labs. Unless America wants to cede that market entirely, OpenAI et. al. need to kick it into gear, fast.

Comment by curious_cat_163 2 days ago

> The government should use procurement to create demand for portable, interoperable systems rather than permanent dependence on one API vendor.

Now, here is an idea that I have not heard before... and I think there is some merit to this. This is also the sort of thing that a state (looking at you CA, CO, IL, NY) could do, instead of just the federal government.

Comment by tangotaylor 1 day ago

I dunno about the other states I feel like we need a better coalition lobbying California lawmakers about AI innovation because I feel like they're doing a lackluster job.

For example, Buffy Wicks introduced AB 2023 (passed the Assembly) which will effectively cause chatbot operators to ban minors because of the huge liability risk that the law introduces. This is a great way to kneecap innovation by excluding curious children and teens who are often the most innovative. California already has laws on the books to address chatbots encouraging self harm (BPC §§ 22601-22606) so I don't know why on earth they're doing this.

Then there's AB 2169 introduced by Lowenthal, which would have mandated interoperability between chatbot platforms to help people migrate to competing ones easier. I thought this was awesome but it didn't even get a vote in the Assembly.

Maybe the newly-created Little Tech Association can help here.

Comment by amazingamazing 2 days ago

Sadly until china scales production of hardware it really isn’t economical to run this stuff yourself. It is good it exists though to put pressure against the labs.

Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.

Comment by sschueller 2 days ago

Model-on-Chip is coming. GPU are for general computing but have a huge bottle neck for doing model inference.

Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.

Comment by baby_souffle 2 days ago

I don't think asics specific to a specific model or even model family are likely to be commodity hardware anytime soon.

It's extremely expensive to build that and you'll be at least two major model generations behind before you even get your first wafers back. By the time you got your production run ready to go and packaged for market nobody's going to care.

Once we end up going something like 24 months between major advances and capabilities for these models then I can start to see asics for a model being possible.

Comment by thisoneisreal 2 days ago

One thought I had is that you could use FPGAs to get hardware performance but maintain the ability to dynamically update. I don't know enough about hardware to consider trying such a thing but I'm curious if that could be made practical and economical somehow.

Comment by distantviewer 2 days ago

Probably not for 2-3 decades unfortunately. The largest FPGAs in the world come out around ~10 million cells (the programmable component of an FPGA). GPT-2 is 1600 million parameters, Qwen 3.5's smallest model is 800 Million parameters.

Comment by nicce 2 days ago

Many years until consumers can buy them at reasonable price. Nvdia and AMD are making GPUs bad in purpose for consumers so that nobody can build a datacenter from them. It will take a long time.

Comment by Danox 2 days ago

Hardware and software getting better every day the barbarians are at the gate…

Comment by cousinbryce 2 days ago

This will probably work well with SotA planning and local chip implementation. I see them being like cars. Cost a few 10k on credit, buy a new one when the old one goes bad or marketing convinces you to upgrade.

Comment by ben_w 2 days ago

> Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.

Yeeeees but the models are in some sense doubling in performance every 4 months, so I expect this to happen in serious quantities approximately when the economic bubble bursts and investors are no longer willing to pay for training.

(Based on widespread news reporting of the existing impact on US electricity markets, I expect this around the end of this year; but with regards to news reporting I am aware of the Gell-Mann amnesia effect, so if this is as much BS as the water issue turned out to be…)

Comment by root-parent 2 days ago

>> do most things and it then is game over.

For the Hyperscalers...and Oracle...cant wait for the day...

Comment by esseph 2 days ago

And it floods the market with millions looking for work

Comment by esseph 2 days ago

> Sadly until china scales production of hardware it really isn’t economical to run this stuff yourself.

I'm running this stuff at home on my desktop and using it through an app on my phone. 60-140TPS depending on model / use case.

It's more than fast enough to even maintain voice conversation.

Comment by 2 days ago

Comment by JumpCrisscross 2 days ago

> it really isn’t economical to run this stuff yourself

Quantised models running overnight go most of the way for non-coding tasks.

Comment by hedora 2 days ago

Pre bubble prices (~= “we stop building data centers with subsidized credit / circular loans / hidden debt”), a 128GB halo strix ran for $1400, and 200-ish watts. Four of those in a cluster will run a 1T parameter frontier model:

https://www.amd.com/en/developer/resources/technical-article...

At 7 months of claude code subscription per node, the cluster pays for itself in 28 months. On a 5 year (60 month) depreciation schedule, you can buy two of those clusters for basically break even, so you get two concurrent request streams (each of which can batch, etc).

The next generation hardware has already been announced, and should ship roughly two Moore’s law doublings later. It’s likely its steady state price is <= $1400 USD (2024), and it is faster.

So, once the bubble pops (because the financial machinations eventually will come to an abrupt halt), and the labs stop buying hardware for data centers, local inference will be extremely practical and cheaper than a subscription.

My main question is, when that happens, will UNIX Surplus be selling inference servers for pennies on the dollar (like after the dotcom crash), or are the power requirements too exotic for home use?

Comment by mft_ 2 days ago

I’m happy to be proven wrong, but the limited examples I’ve seen of clustered Strix Halos are quite slow running large models (ie models too large to fit into the ram of a single machine) due to the slow networking between each one?

Comment by downrightmike 1 day ago

That is already being addressed with GPUs and special PCIe daughter cards that link them together faster than the PCIe on the MOBO could. SLI but better.

Comment by hedora 1 day ago

Yeah; the benchmarks + setup guide I linked show that it takes a lot of careful setup to get them to work at all, and then you only get 100's of tokens per second per cluster.

That's why I pointed out the next generation is coming soon. Also, the AMD docs aren't using quantization (as far as I can tell, I only skimmed), which gives a speedup roughly linear in the compression ratio. Algorithms for that continue to improve, so expect a lossless factor of 2-8x on DRAM and throughput, at least.

Comment by mft_ 1 day ago

If I rephrase your earlier post, and based on the AMD post:

If you take four Framework Desktop 128 GB Strix Halo motherboards at a previous low and currently-unattainable price, and ignore the cost of storage, power, networking, and cases, and then take an older model that's just about competitive with Opus 4.5, and then you quantise that model down to Q2_K_XL and ignore the performance degradation below Opus 4.5 this will cause, and then you compare the cost of this with the highest tier of Claude subscription that includes much newer models which are far more able than your Opus 4.5 benchmark, then you might have some sort of break-even inside a year.

Yes, there's a new AMD hardware generation coming, but it's going to be as heinously expensive as the current one has become, and it's still got relatively low memory bandwidth. Yes, there are already newer models than Kimi 2.5, but with the limitations of this cluster (e.g. still needing heavily quantised models to be feasible) you'll get incremental improvements at best.

I'm keen to be supportive and bullish about open/home models, but I worry that this is such a stretch, and the two options in your comparison are so incomparable, you'll turn far more people off than you convince.

Comment by 2 days ago

Comment by esseph 2 days ago

> Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.

I don't see how these are related.

The accuracy and capabilities of your model are directly related to its size. You need a lot of memory for that.

It will be decades before we get enough useful memory in a phone form factor at a price point people can afford it before something like a frontier model now is useful on the phone.

Now, you can run some models on your phone today.

Either way, Apple is using Google today. That could change, but Google isn't exactly getting out of the TPU business and they've been doing it a long time.

Also, some of you live in a very weird Apple bubble. Apple is not so relevant outside the US.

Comment by Danox 2 days ago

If it wasn’t for the current memory fiasco, Apple would be even closer to providing a solution on the desktop that is affordable for most (Hacker News participants) who want to run a large model locally at reasonable cost, which will happen within the next 3 to 5 years easy with the way software and hardware are progressing, the main-frame future for most people that use computers isn’t coming back so OpenAI and Anthropic and other frontier developers are going to be disappointed and will have to shift gears ala’ Meta, Microsoft?

The United States if it dares to (I think they will try) but isn’t going to be able to stuff AI models back into the bottle open source is the future and when it comes to AI models yes you’ll be able to customize it to your specifications locally but the genie is out of the bottle. The bull out of the barn and is running down the road.

If United States insist on trying to lock the doors, censor, sanction, the rest of the world will just design and engineer around the United States. Trying to put up a wall, will damaged the United States more particularly with the current performance of Taco. None of the other countries are going to follow the United States not with the current administration they will hedge their bets.

Apple, is using Google now but that will change because the world is probably going down the open path, it’s looking like there was no real rush and no reason to spend so much money on something that’s going to be a commodity in the end, the only hold up is hardware and if it wasn’t for this current memory fiasco, many more people would have access to the hardware that they need to run models locally.

Comment by esseph 1 day ago

> If it wasn’t for the current memory fiasco, Apple would be even closer to providing a solution on the desktop that is affordable for most (Hacker News participants)

Already running models locally without apple hardware, and have an encrypted vpn tunnel from my mobile devices back to my desktop over the internet. The time is now.

Comment by 2 days ago

Comment by serial_dev 2 days ago

[dead]

Comment by segmondy 2 days ago

I don't know what you mean by "economical", but it has been "economical" to run this stuff yourself for the last 3 years.

1. You must be willing to be resourceful. 2. Be willing to learn, do the hard things. 3. Accept the tradeoffs.

Comment by thih9 2 days ago

Is anyone using open weight models for agentic coding?

What is your stack (harness, model) and how much do you pay per month?

How would you compare your experience to a typical subsidized plan like Claude Code + Pro plan?

I’m asking because i keep hearing that open weight models are cheap and efficient - is that really the case in practice?

Comment by nyrikki 2 days ago

I don’t know if others would find this useful, but previous did have custom harnesses etc.. but tools have improved so much that I drastically simplified.

That said, even the foundational models fail at the hard parts of my code so I use it opportunistically.

I have reduced down to just using zed, will three locally hosted models.

Qwen 3.6 27b on 1x3090 llama.cpp with 128k context ~50tps

Qwen 3.6 35B-A3B on 1x titan v + 2x1080ti llama.cpp with full context ~30tps

GPT-OSS 120b on pure cpu (slow)

I just use zeds parallel agents, task switching, stopping and fixing the code when a model gets stuck.

This still lets me stay engaged, and to modify code to be maintainable etc…

It gets me 80% there and I use to keep a subscription but often times just using googles AI mode is just as good.

That said I have 30 years of experience and insist on knowing how my code works, so this gets me 80% of the short term benefits while not depending on a 3rd party to keep my code moving forward.

Your mileage will vary and 2*5060ti 16gb cards would get around 100/tps with Qwen 3.6 35B-A3B on cards that are widely available.

To be honest the more modern cloud models are using draft tokens etc… that while they are superior for common coding tasks are degrading with more domain specific tasks.

That is just the cost of the draft model being ~10-20% of the foundation models size, and even the biggest Blackwell GPU is limited to ~250/tps so MoE or draft models are required for scaling performance at the foundational level IMHO.

The hard part is my use case are the OOD or small examples in corpus level, the above hurts there.

A Lamborghini may be nice, but I personally need a minivan more.

Comment by TGower 1 day ago

Are you saying that cloud models are not verifying the draft model predictions? The way draft models are used in something like llama.cpp results in exactly zero degredation of output quality, with the larger model verifying each draft model token and discarding it if it does not match.

Comment by Foobar8568 2 days ago

You wouldn't get 100 tps on a Qwen 3.6 35b with a 5060 (or two) when a 5090 can barely reach that.

Comment by nyrikki 2 days ago

Depends on quant size etc... Qwen3.6 35B-A3B Q4_K_XL a multiple 5060ti + tensor split mode + MPT will hit ~100/TPS without problem, and I have personally hit ~190/tps on a single 5090 on a friends machine getting them setup up. If you use Q6_K etc... it slows down, what quant were you using?

Quantization + KV cache paging + speculative decoding (MPT or draft) is a fairly good mixture here.

Some examples as I don't have access to run tests on a 5060ti right now:

     https://njannasch.dev/blog/gemma-4-mtp-vs-qwen-speculative-decoding-5060ti/#vs-qwen-36-mtp

     https://www.reddit.com/r/LocalLLM/comments/1umw7vj/dual_5060_ti_16_gb_llm_inference_performance/
And here are some logs on unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_XL with the 2x 1080ti + 1x titan from above:

     27.21.533.298 I slot print_timing: id  0 | task 7843 | n_decoded =   1780, tg =  62.15 t/s
     27.24.537.173 I slot print_timing: id  0 | task 7843 | n_decoded =   1965, tg =  62.10 t/s
     27.27.541.142 I slot print_timing: id  0 | task 7843 | n_decoded =   2152, tg =  62.11 t/s
Q4_K_XL is a slight, acceptable degradation IMHO for performance like that.

Comment by slim 1 day ago

I get 66 tps on one. so 100 on two seems plausible to me

Comment by tyfon 2 days ago

I'm using qwen 3.6 35B unsloth 4 bit with my 5950x (128 gb memory) and a 3060 12 gb gpu with a self made harness.

At 10k context I get about 40 tps generation and 500 tps prefill. At 100k context I get about 25 tps generation and 400 tps prefill.

It works, but I often use gpt or claude to make a detailed enumerated plan of what I want to do first, then have qwen follow it.

I'm not sure if it is economical or not, but I have solar on the roof so the power use is not really an issue and I already have the hardware.

The biggest benefit for me is that it's all done locally, and I know the harness is not uploading anything or sending telemetry to someone else.

Comment by johnvanommen 2 days ago

> The biggest benefit for me is that it's all done locally, and I know the harness is not uploading anything or sending telemetry to someone else.

Are there any articles you’d recommend for this?

I have Qwen running on an HP Z8. Very nice platform.

I have mine in a sandbox, due to privacy fears.

Your solution sounds more elegant.

Comment by tyfon 2 days ago

Articles regarding my own harness or how I set up llama.cpp etc?

I really just iterated over the harness over and over for about two weeks with opencode until I was sort of satisfied (still lots to do there :).

For the llama.cpp I asked claude fable to optimize it for my hardware and iterated a few times. In the end I landed on the following: https://pastebin.com/2PpJFUC0

Comment by Scene_Cast2 2 days ago

I'm using Kimi K3 + OpenCode. I pay their API pricing, costs about $5 / hour (and chews through ~10 million tokens / hour) during continuous use when I have one or two sessions running and doing their thing.

Can't comment on how it compares to plans (I really don't like the limitations and general shenanigans I see around plans, so I've never tried them).

It is notably slower than Fable / Opus / Gemini, but also vastly cheaper than their API pricing.

Comment by 2 days ago

Comment by rglullis 2 days ago

I am using GLM-5.2 via Ollama Cloud in the $20/month plan. With the same plan I can get many different API keys that I use to run my OpenWebUI server, my opencode and pi dev sessions. I am usually running 2 to 4 sessions concurrently, and I never hit quota limits. At work I get Claude, and I was getting reports that I was spending $75 per hour of work on Opus.

Comment by ForHackernews 2 days ago

I use DeepSeek 4 with the VSCode CoPilot plugin. I pay about $10 month on the pay-as-you-go plan.

It's not as good as the frontier models I use at work, but it's plenty capable for the types of tasks I am using it for.

Comment by eulers_secret 2 days ago

I use opencode or pi harness with the deepseek api for all my at home coding usages.

Deepseek is at least on par with Sonnet (ghcopilot at work)- I don’t use opus, too spendy and I don’t need that level of ability.

The cost is for me was $5/6 months of use. Not a big user I guess! It’s good though, fast enough and incredibly inexpensive.

Been testing Qwen 3.6 28B on a 5090, and it’s also quite good for “free”.

I mostly do small self serving embedded projects based on esp32, so not very complex.

Comment by airstrike 2 days ago

> I’m asking because i keep hearing that open weight models are cheap and efficient - is that really the case in practice?

I think that's the case for people who compare it to proprietary models paid via API—which I think is irrelevant given the majority of people daily driving AI coding are doing on a subscription plan.

The better analysis then is not about AI coding, since there's no subscription plan for Kimi K3.

Instead, compare the cost of running some agentic _task_ that isn't coding which can only be done via API. Think of all the startups wrapping around ChatGPT and Claude to provide some additional set of tools, context, data and hoping to turn it into a profitable service.

To those companies, which are many, open models are the difference between the math working out today vs. praygeing they can scale fast enough to find profitability.

Comment by Y_Y 2 days ago

GLM 5.2 awq4 via Opencode, it was better than corporate's fave Sonnet 4.8 and a but worse than Opus 4.8. overall quite capable of a lot of what the dev team needed. Not as good as Sonnet 5 but also never runs out of tokens.

The flip side is that it takes four H200s to run, and that will only let you cache context for maybe three users.

Fingers crossed our Blackwells show up and Kimi 3 really releases weights, because at some point devs are spending a significant portion of their salary on tokens and it's somehow cheaper to buy these ridiculous DGX servers and rack and run them.

Comment by foobar10000 2 days ago

Glm 5.2 nvfp4 on 4 b300 with dpattn 4 and ram will get you about 20 users live at 400k context - and 60 easily if you give that server 2tb of ram and 4 nvme 8 tb drives. There are some sglang patches needed - but we will be releasing them soon.

Comment by nojs 2 days ago

What hardware are you expecting to run K3 on?

Comment by Y_Y 1 day ago

Eight B200s, at whatever quantization I can fit, ideally fp4. Apparently it might fit in eight H200s, but with 2.8B parameters (1.8% active) I'm not optimistic.

Comment by HDBaseT 17 hours ago

At least with the Kimi Code plan, its limits are pretty abysmal compared to ChatGPT/Claude plans.

Comment by revolvingthrow 2 days ago

While I am grateful for open weights models I never found much use of them in the past, barring those I could run myself. This changed with deepseek 4 - it is staggeringly cheap, even if the performance definitely isn't near sota and it's not particularly fast either.

When I expect to need a lot of tokens and the task isn't too difficult I use sota to plan and create a thorough set of instructions and let deepseek chip away at it. With thorough instructions the quality tends to be satisfactory, and you pay something silly like $15 for 600m tokens.

GLM 5.2 seems like a decent price/perf and Kimi 3 has some real nice performance for an open weights model, but gpt 5.6 is unexpectedly affordable (especially if you don't automatically use Sol at max) so I don't think either is worth it atm. The exception is when you're working on something that US models get cold feet about, which seems like a constantly growing list. For me Fable is already too much of a headache in this regard, but chatgpt is still okay-ish. Hopefully it'll last. If not, there's Kimi.

tldr SOTA for most things because gpt 5.6 is token efficient. If I expect to burn a lot of tokens I use deepseek 4.

Comment by qiine 2 days ago

qwen3.6 27B q5, llama.cpp, RTX 3090, pi, cost: electricity bill

Comment by amazingamazing 2 days ago

Rex 3090 isn’t free. Even if you already owned it, it wasn’t free. That’s years of a $20 subscription

Comment by broodbucket 2 days ago

It's an asset, though, and bizarrely it's one that's been appreciating the past 5 years

Comment by johnvanommen 2 days ago

I took the same attitude. The hardware isn’t getting cheaper, it’s getting more expensive.

As I see it, an investment in AI hardware is an investment in my own future.

IE, I drive my car a couple of days a week, and it’s perfectly normal to spend $500 a month on an asset like that. When you factor in the SPACE it takes up, that’s the REAL cost of owning a car: the real estate you have to buy for your car to occupy.

Once that’s factored in, the “true” cost of having a car can easily be $2000 a month, even for a crummy car. The space that the car occupies is expensive.

Yet people balk at spending even $2000 on a GPU.

Makes no sense to me. I choose to invest in the future.

Comment by qiine 2 days ago

I see your point but we could go pretty far with this logic. motherboard ? cpu? ram!!! screen? fancy keyboard? etc..

Comment by overgard 2 days ago

I'm using Qwen 3.6 27B on a macbook with Pi. It's alright, it runs fairly quick (40 tps for quality version, 80 for the fast). It doesn't tend to one shot things but I'm generally comfortable fixing the bugs myself afterwards or prodding it a little bit. I find the harness matters a lot. "Continue" (the vscode extension) worked horribly, OpenCode was ok but its vibecoded internals make me view it as a security nightmare so I'm hesitant to run it, so I've settled on Pi for now.

Claude and ChatGPT are good deals right now, with the subsidies. They produce things faster and better. I guess not cheaper, in that inferrence on my macbook is basically free, although the macbook itself definitely wasn't. My focus on running local is around three principles:

1. I don't want to support surveilance capitalism by giving these companies my data anymore, when I can avoid it. And LLM companies want to vacuum up every detail of your life.

2. I don't find these companies to be remotely trustworthy, and I find them hostile to a healthy society, so I want to avoid giving them money going forward

3. I think they're going to start charging a lot more

Comment by sschueller 2 days ago

Kilo code with direct API payment to DeepSeek. It costs pennies per day event at max.

Comment by chasd00 2 days ago

FTFA: American labs need to release frontier-grade open-weight models under licenses that startups can actually build on.

oh now i see, the Chinese government is funding the training and release of their best models to pressure OpenAI, Anthropic, and others to do the same for competition's sake. I don't buy it, this seems more like a way to get SOTA models RL'd to comply with Chinese government approved information distribution. If I have to trust a black box of answers to questions i would trust one from a US for-profit publicly traded company subject to market forces over one approved, and heavily subsidized, by the Chinese government.

Comment by Alwayshasbeeb 1 day ago

A "US for-profit publicly traded company" is what brought us the Cambridge analytica scandal. I'd rather not trust them with the next wave of consciousness and opinion shaping technology.

Also, oligopolies aren't famous for being strongly bound by market forces, especially when their decision makers are non ironically being treated as if they were heads of state.

https://www.nytimes.com/2026/06/17/world/europe/g7-summit-ai...

Comment by georgeburdell 2 days ago

No, it’s consistent with what China is doing in other markets, which is dumping product to drive others out of business.

I had a shower thought on how to counteract this, specifically related to the AI dumping. If China is losing substantial money on every token, why wouldn’t an adversary try to maliciously increase consumption? This strategy is not really viable against physical goods dumping because demand is finite and there are large environmental costs. Software demand is infinite and the environmental costs are quite low compared to the financial cost to make it, even with ultra cheap Chinese tokens

Comment by m11a 1 day ago

Because it's not losing money on each token? Aside from most global people using American inference providers to run the models, I suspect the cloud inference products of the Chinese labs are profitable, at least on the inference costs (ie: not including model training, salaries, etc).

Comment by georgeburdell 1 day ago

There is no way Deepseek is making money even on inference

Comment by arczyx 1 day ago

Pretty sure they are making money since on OpenRouter, there are other providers for DeepSeek V4 flash that are charging even less than DeepSeek themselves (eg DeepInfra and Digital Ocean).

https://openrouter.ai/compare/deepseek/deepseek-v4-flash/ten...

Comment by georgeburdell 1 day ago

That says little. Those providers could also be losing money trying to gain marketshare. There is a high amount of speculation in the space and it won’t be apparent for awhile who has a lasting business.

Comment by culi 1 day ago

That doesn't mean they're making money. Other providers could just be doing it more efficiently. DeepSeek is trying to make its model optimized for Hauwei chips instead of NVIDIA so it has its own constraints

Comment by aliasxneo 2 days ago

What tools do we have to countermeasure the state sponsored bias in the Chinese models? Doesn’t seem like a smart plan if individuals can just compensate for the bias.

Comment by hedora 2 days ago

Also, the choice right now is between an open weight Chinese model that is hypothetically censored to block / sabotage routine engineering flows vs a closed weight service that is definitely censored to block / sabotage those things.

First anthropic guardrails blocked totally normal stuff on fable and knocked you down to opus. At this point, they kick you off fable, then opus, then sonnet. Claude then automatically builds up memories of techniques to bypass the guardrails in my long running sessions (the coordinator agent notices the subordinates got shot in the head and their sessions were pulled from context, so it parses out the lost context from ~/.claude json files, then reformulates parts of the task and uses partial results until the guardrail doesn’t trip.

If I were paying for the API, this dance would cost $50-100 a pop, but I’m not, so whatever (for now).

Comment by johnvanommen 2 days ago

I worry that AI will be so fundamental to how we do things in the future, companies can mold human behavior via access to the AI tools.

For instance, I worked at FICO. When I mention this, people wonder what they do. The average person only knows FICO as a “score.”

FICO was founded in Silicon Valley.

The average person doesn’t think about how credit scores work, fundamentally.

It’s just software, at its core. FICO incentivizes certain behaviors.

Ever been banned from an online forum?

Now imagine if a corporation could shut you off from a technology that’s literally indispensable.

Same idea.

Comment by poisonborz 1 day ago

It's why they say social credit system is already way more built out in the US than any other part of the world.

Comment by netdur 2 days ago

why would any software want to have Kubernetes moment? can't count how devop I know that is confused by it

Comment by honkycat 2 days ago

Recently my company bought another company, and we kept zero of the original engineers, we just had to run the ghost ship.

We walked in, and it was fine. Because it was all kubernetes and laid out like every other app for the most part.

The kube hate is just sad at this point. You need to know like 15 concepts that are all applied in the same way. It mostly just works.

Comment by singingtoday 2 days ago

I ported my company over to k8s to solve a concurrency and scaling issues.

What 15 concepts? You're making me worry that I missed something. It was straight forward: pods, nodes, hw type, lifecycle, deployment. They run almost the same docker as the old ec2s used.

What did I miss? Is there something important I need to read?

Comment by honkycat 2 days ago

Lol no you got em

I would say:

- deployments - pods -services - ingress - namespaces - cert-manager - external DNS - external secrets - configmaps - hpa - docker - volumes/pvc

That's the basics

Comment by Syntaf 2 days ago

It’s the pineapple on pizza of ops, people just like to fit in sometimes.

I’ve been running my own personal k8s cluster on digital ocean for the last 5+ years now and it’s dead simple.

Takes me 30 minutes to create a new namespace a deploy an app, love the bonus of having complete flexibility on my stack too — PVC + SQLite ftw

Comment by itomato 2 days ago

Ghost ship status is not something most orgs aspire to.

The value built on that stability was probably worth acquiring, and it’s infra will decay.

Comment by xyzsparetimexyz 2 days ago

I still don't know what it is tbh. Something for docker?

Comment by yard2010 2 days ago

I was in the same boat as you until I needed to learn how to use it in my $dayjob. It was like discovering a new continent. I couldn't care less about it before, but the moment I realized it's a kind of cloud OS I was astounded by how this thing is genius. It's the kind of thing that gives you dopamine rushes when you use it. Something about how there is a solution to every problem you didn't know even mattered turns it into a magical perfect software. I just love it.

There is something special about complex systems that just-work(tm)

Comment by jdub 2 days ago

This is the initial endorphin rush you get when wielding a complex system. The feeling changes when the complexity explodes in your face.

Comment by mystifyingpoi 2 days ago

Well said. The rush is when one understands the loosely coupled control loops and how they work together... until something fails in the middle with no error.

Comment by johnvanommen 2 days ago

> the moment I realized it's a kind of cloud OS

The foundation of this is thirty years old:

Mark Andreesen founded a company to do what Amazon did: Loudcloud. This was cloud computing, BEFORE AWS was public. Loudcloud was founded in the nineties; AWS opened its APIs in 2006.

Loudcloud failed and became Opsware. Opsware was server automation.

Its competitor was Bladelogic.

Luke Kanies, from BladeLogic, founded Puppet. Puppet was open source, and steamrollered over nearly every installation of BladeLogic and Opsware in existence, because you can’t beat free.

Folks from BladeLogic migrated to jobs at DCOS.

DCOS was steamrollered by Kubernetes, the same way Puppet steamrollered BladeLogic.

The author of the piece predicts that open weight models will steamroller everything next.

I fear he’s right.

I worked for Opsware and BladeLogic.

I watched it happen in real time.

Financially, Andreesen’s wealth stems from Opsware. He is known for Mosaic, but Opsware put him on the map, financially. HP bought them.

If one wants to follow in Andreesen’s footsteps, study how he did it at Opsware.

Conveniently, there is a book.

“The Hard Thing about Hard Things.”

Comment by m4rtink 2 days ago

So Kubernetes is LSD? ;)

Comment by chias 2 days ago

I was in this state a few weeks ago. I spent a bit of time familiarizing myself then wrote up my learnings as a series of exercises.

If you think of docker as "kinda like vms except not really" and k8s as "kinda like deploying and composing docker containers but not really", this may be for you:

https://ojensen.net/infra/understanding-k8s-1

It's actually really neat, i wish i had bothered to learn it years ago.

Comment by munchler 2 days ago

I appreciate the effort and I'm in your target audience, but that document didn't help me. It seems to dive into the details of installing and running k8s without saying much about the purpose.

From my very ignorant standpoint, K8s seems to be about running a "cluster", but I don't know why I would want to do that.

Comment by what-is-water 2 days ago

Kubernetes orchestrates your container workloads over a cluster, which consists of virtual/bare metal machines(nodes). This means you can tell the kubernetes API "I want to run a container workload" and it will be started on one of the nodes that form the cluster, unlike e.g. Docker, where a docker daemon belongs to a specific node. If you remove the node your workload is running on from the cluster the workload will be rescheduled on a different one, or if you have a new image version it will start the new container, wait for it to become healthy and ready, and then route requests to it. And if you want to send requests to your workload kubernetes allows you to define standardized abstractions to easily route them to your workload, irrespective of the node it is running on.

It allows you to stop caring about the individual machines, and just treat them as combined compute, which starts mattering if you leave a single machine setup and need to start thinking about scaling in and out and gluing the individual parts together. Then you have known abstractions to do it.

Of course you can do everything kubernetes does using a bespoke solution, and the concepts aren't new, but having a widely supported technology has a lot of advantages and creating something with even half the feature has a high chance of just being worse.

Comment by munchler 2 days ago

Thank you for this explanation. It makes sense, but I don't really understand why it has become so popular.

Professionally, my experience is that certain software components need to run together on an individual machine (e.g. database server, app server, web server), and then those machines need to be networked in a certain way (e.g. web server talks to app server, which talks to database server), so I really need to care about the architecture of individual machines. You can then scale this out horizontally (e.g. add another web server) or vertically (e.g. upgrade your database server).

I'm old, so maybe I'm out of date, but having a cluster of "compute" that I can run arbitrary workloads on sounds neat, but is a capability that I've never needed.

Comment by Krei-se 1 day ago

Your host with the webserver has no idea its just a docker image. It talks to app.domain.tld and gets a reply. app.domain.tld asks db.domain.tld with an SQL query and gets a reply. But it can be one of any deployed docker image on any bare metal host - which one is db. and app. etc. is decided by Kubernetes.

In the background kubernetes routes all these docker images with each other without you having to think about this. You may have 1000 db.domain.tld nodes caching a master db host - if you configure that in the docker image, that's not different to bare metal replication.

Same with load balancing in webservers. You can do that! Or you can just have kubernetes handle it. I'm not sure but would expect it to swap images in a way network is optimized - i don't use this kind of software. I just know it's done like this because the bottleneck of modern software is not the local network so its not an issue to have these pieces on different bare metal hosts. And its easier to stay operative if some hosts fail - but you can all solve this by hand.

I work with bottlenecks between network, ssd, ram and vram but if you deploy npm riddled software for >1 Million users well then you may want Kubernetes. Or if you are google and have 10k engineers that need to agree on a standard!

If you can do this yourself with load balancing, replication etc - do it yourself and don't think there's something wrong with that. It's more elegant, efficient - but you have to agree with others how you do it. And that can also be a bottleneck ;)

Still i think you doing it by hand is better so don't worry you are not old you may just have higher standards.

Comment by munchler 1 day ago

Thank you for this. I probably don't have higher standards, but I do have way fewer than a million users!

Comment by what-is-water 1 day ago

> Professionally, my experience is that certain software components need to run together on an individual machine (e.g. database server, app server, web server), and then those machines need to be networked in a certain way (e.g. web server talks to app server, which talks to database server).

This (different workload components running on a single machine) is something that kubernetes allows you to disentangle. Kubernetes creates its own network, including cluster internal DNS. Using this you expose e.g. your app servers as a service called my-app, reachable in cluster via my-app.namespace-name.svc.cluster.local.

This targets all containers with a specific label, no matter on which node they run. Server types can be scaled independently, since chances are that the load for each does not scale the same with request volume.

Round robin for DBs does not make much sense, but there are ways to e.g. expose read endpoints with one service, and write endpoints with a different one, with open source tooling which updates the target after a failover.

Kubernetes will automatically keep your services up to date, which means if you increase the replica for e.g. app server the new container will be added as valid target, as soon as it passes ready checks, and if one container fails these checks they are temporarily removed as target. The kubernetes components will also automatically "self-heal" things like a crashed container or a failed node, by restarting the container or rescheduling the workloads on the failed node to a different one, without human intervention.

This is of course very basic, but you can finetune these by e.g. configuring that a specific server type should be spread out, i.e. that only one replica(container) should be scheduled per node, to ensure it stays available if one or more nodes go down. Or add network policies to ensure only the app server is allowed to talk to the DB server.

If your concern is latency between app and db you can have specific config that ensures that your app server containers are only scheduled on nodes where a DB server is already running, and that the traffic from app to DB is always routed to the DB instance that is on the same node (people use similar mechanisms for cloud providers like AWS, where you want to have routing rules ensuring that traffic is always sent to targets in the same AZ, to avoid cross-AZ network charges).

These advanced examples obviously require deeper kubernetes knowledge and are not something one should just try out the first time you deploy kubernetes.

Having worked with more traditional setups I do think it is often easier to configure config like this in the standardized kubernetes API rather than in e.g. nginx config + deployment scripts + idk, systemd-unit. But this point is not "having thousand of nodes" and be half the size of google.

It also depends on your team, if you have an infra team that has a stable way to manage your VMs and apps, all the power to them, replacing them all with k8s experts sure won't give you much. It isn't easy to get an unbiased opinion about when to switch, since you need knowledge of kubernetes and your current infra to make a fair comparison, and kubernetes experts probably want to sell you kubernetes. Using a handcrafted system to distribute a lot of containers over multiple VMs to ensure HA is in general a good sign to evaluate kubernetes ;)

And being on-premise makes kubernetes attractive earlier, since the bigger cloud providers have managed solutions for a lot of things kubernetes helps you with (auto scaling, load balancing, managed container platforms like AWS ECS or Google Cloud Run)

Comment by genghisjahn 2 days ago

If you can get it implemented for anything you can’t be fired.

Comment by Pxtl 2 days ago

It's for running a massive number of docker containers and automatically managing them and scaling them up and down on demand. It is also so famously brutally complex that basically you need a dedicated Kube expert to handle it.

Comment by honkycat 2 days ago

to be fair, at any large scale you need infra people.

Comment by RussianCow 2 days ago

The problem is that companies tend to exaggerate their own scale and think they need k8s and dedicated infra people when they could get by with a handful of beefy VMs or dedicated servers.

Comment by honkycat 2 days ago

Or you could use hosted k8s and be future proof.

I would take a kube cluster over a bunch of VMs I have to hand wire: wire releasing to, managing processes, restarting crashed processes, log aggregation, load balancing, networking, secret injection, cert management, DNS management, monitoring, etc... Any day of the week.

You just don't know what you're talking about, sorry. Kube is really easy now.

Comment by marcosdumay 2 days ago

Kubernetes is for docker what your init system is for daemons.

Comment by spicyusername 2 days ago

You... don't know what Kubernetes is... pretty impressive, honestly.

Its 2026 and its the de facto method of deploying software basically everywhere.

You gotta really work for it to not know what its for by now.

Comment by recursive 2 days ago

I don't do much deployment but I'm in the same boat. Something something docker automation?

Comment by chrisandchris 2 days ago

> Its 2026 and its the de facto method of deploying software basically everywhere.

That is some really impressive bubble you are living within. Basically everywhere - nowhere near that, no.

[edit]: Maybe containers, but software in general is so much more broad than containers.

Comment by thewebguyd 2 days ago

Everywhere, huh? Still plenty just running on VMs or serverless that don't need full blown container orchestration.

I'll agree that (nearly) everyone should know when Kubernetes is useful, but let's not pretend its the default method for everything. Even then, choosing to deploy on K8s falls on the sysadmins/DevOps I wouldn't expect the devs know or do much more than provide the Dockerfile.

Comment by YetAnotherNick 2 days ago

2026 is the year for vercel and Render.

Comment by RussianCow 2 days ago

I haven't used them but aren't they basically modernized Heroku? What's different?

Comment by mindwok 1 day ago

Pretty close, they're very close to Heroku in spirit except the new wave of hosted backends (Vercel, and to a lesser extent Supabase and Cloudflare Workers) are much more tightly integrated with the app layer. You can almost think of them as libraries that run in the cloud which you can just use inside your app.

Comment by throw-the-towel 2 days ago

Heroku is dying, these are not.

Comment by 2 days ago

Comment by YetAnotherNick 2 days ago

In fact more countries should have government funded models. There are some obvious issues in China completely dominating open weights space. Kimi had funding of just $2B and could literally create national security threat. A lot of countries could fund something in the range of few billion for something so important. At the very least US and EU could fund few companies.

Comment by cheriot 2 days ago

Open-weight and OSS are wildly different and the article makes a poor comparison.

What's the incentive for the Chinese labs to continue releasing weights 5 years from now? It's not a stable equilibrium and cannot last.

- The lab spending large sums on research and training does not get the inference revenue to fund those efforts.

- Unlike OSS where a single volunteer can keep a project going, training costs run into the $billions.

- OSS is often a two way street where features and integrations are built that the original author benefits from. Open weight models are largely a one way street because the marginal benefit is so much less than training costs.

In the short term, it means Chinese labs can attract talent and, I suspect, funding from their gov. Similar to every other industry the CCP subsidized to take over.

Comment by aleph_minus_one 2 days ago

> - Unlike OSS where a single volunteer can keep a project going, training costs run into the $billions.

Just some thought: Wouldn't it make sense to build some kind of volunteer computing project to train the next-generation LLM by volunteers, similar to the BOINC [1] projects or Folding@home [2]?

N.B.: BOINC was particularly famous for SETI@home (completed), Einstein@Home, Rosetta@home and PrimeGrid.

I still remember the time when Einstein@Home was in its heyday, and many people who loved putting together fast PCs contributed sometimes even for the reason of showing off in the statistics [3].

---

[1] https://en.wikipedia.org/wiki/Berkeley_Open_Infrastructure_f...

[2] https://en.wikipedia.org/wiki/Folding@home

[3] https://einsteinathome.org/de/community/stats

Comment by cheriot 2 days ago

I’ll be impressed if somebody can make that work considering the vastly larger compute required.

Comment by aleph_minus_one 2 days ago

> I’ll be impressed if somebody can make that work considering the vastly larger compute required.

I think you underestimate the computational ressources that the mentioned (and similar-kinded) scientific projects needed. Also consider how much computational ressources people invested into cryptocurrency mining.

No, I think the reasons are different:

- Many companies that train AI model use training data which must not be distributed for copyright reasons (and using it is a legal gray zone)x.

- Also consider that the amount of training data is insane. Scientific projects (and cryptocurrency mining, too) have the property that typically the amount of data (storage requirements) is small (or at least the computation can be partitioned so that each sub-task needs little data), but the required computing ressources are insane.

- AI companies consider a huge part of their training data as their "secret sauce" (they often even paid lots of money to generate it, for example by paying world-renowned experts for writing an answer for some important question).

Thus: Yes, the required computation ressources are huge, but this is a problem for which I consider it to be plausible that it can be solved. The real problems are in my opinion different.

Comment by applicative 2 days ago

China has no end of money to support these companies. The reason this equilibrium is unstable is that the autonomous agentic coding aspect of the models has been so successfully improved that it will soon be a threat to China state security.

Comment by cheriot 1 day ago

Like they did with solar panels, I suspect the goal is to dominate an industry and not need to subsidize it forever.

Comment by 20 hours ago

Comment by Danox 2 days ago

Yes, and yes, again the only way to compete is to build the best not hide in a corner and once again the rest of the world will go on in AI without the United States if we flub it. Circling the wagons, isn’t the long range answer.

Comment by __MatrixMan__ 2 days ago

This is such an obvious conclusion. To take it a bit further...

Scale matters for these things. If we divide the available chips among 5 competing companies we end up with models that are trained on 1/5 of the resources that they otherwise could've been.

Let the companies take turns training on shared hardware, force them to publish results in the open, and then reward them based on how well the resulting model performs at democratically chosen benchmarks. Meritocracy not monopoly.

Make it about how well you wield the silicon, not how much silicon you wield, and make it a positive-sum game. If the people's data is going in, then the people should benefit from what comes out whether or not they have a subscription.

If capitalism as we know it can't complete, so much the worse for capitalism as we know it.

Comment by applicative 2 days ago

Everyone just keeps assuming, as if it were the law of gravity, that China will continue in perpetuity to deliver the weights of its 'frontier' models to Hugging Face. Its Mythos moment is a few months away and there is plenty of reporting suggesting their response will be similar, which is anyway obvious.

It baffles me that anyone can seriously believe that China is going to put its Mythos successor on Hugging Face and let us all strip the guardrails etc.

Things just didn't turn out the way OpenAI and Anthropic thought; nor did they turn out the way China thought.

Comment by nunez 1 day ago

They very well might. Read up on the infrastructure projects China has provifed to Africa in exchange for oil refineries and drilling sites across the continent: https://oilprice.com/Energy/Crude-Oil/China-Is-Rapidly-Expan.... Roads, schools and more.

Comment by applicative 1 day ago

The headline from 7 years ago has not borne out; Africa is a minor supplier.

The point is quite different: that everyone is harmed by advanced open weight agentic models even as we are benefited by them. This holds of African states as of any other. Many of them, by the way, are under attack from well-funded moderately insane groups like JNIM. Also by the way, groups like this are famously really good at recruiting the specific psychical type of the engineer.

It is not possible that releasing open weights will continue forever, or probably even to the end of the year.

Comment by regexorcist 1 day ago

> there is plenty of reporting suggesting their response will be similar

There is? Xi Jinping himself openly talked very recently about China's commitment to open weights models.

Comment by HDBaseT 17 hours ago

Xi Jinping realizes that open weight models accelerate open development.

There is more players in China than in the US. If everyone keeps learning off each other, eventually one of the China labs will overtake the US, but AI models aren't static (anything but), and the US will leap frog in a weeks time. Therefore the constant need for open weights remains.

Comment by nedt 1 day ago

> Kubernetes did not win simply because its repository was public.

And here I am still using docker with compose because it's easier and works. Always saying I might be looking at swarm when needed, but then never really need it.

Just because you are using it doesn't mean everyone is and doesn't mean it has won. On the other hand yeah I want open models and I want them local as well.

Comment by shard972 1 day ago

[dead]

Comment by xivzgrev 1 day ago

I worry about the Chinese models sending data back to China. How do we know that isn't the case?

Comment by ehnto 1 day ago

Depends on how you're running them I suppose. If you are running them on local hardware, then the harness is what you are concerned about. A compromised local model might be able to exfiltrate data via tool calls.

If you are running it in a cloud, the above is true, but also you have to trust the cloud provider isn't relaying your chats to a third party, for training or perusing.

If I am totally blunt, if privacy is a concern then corporate US has proven time and time again to be a terrible steward of your private data. You should assume it's being ingested or sold, or turned into metadata that's sold in a data laundering pipeline.

If it's corporate secrets, on-prem and enterprise offerings are your options. Enterprise at least puts the providers on the hook legally, when they inevitably fuck up keeping your data private in some way.

Comment by Retro_Dev 1 day ago

Indeed - if you need to be absolutely confident about security and privacy, run a model locally and audit the inference software and potential tooling.

Comment by 1 day ago

Comment by HDBaseT 17 hours ago

I worry about the US models sending my data back to Israel.

Comment by cindyllm 15 hours ago

[dead]

Comment by zackwu 1 day ago

There are zero-retention 3rd party providers (EU/US companies), but it's your choice whether to trust their claims or not.

Comment by 2 days ago

Comment by maayank 2 days ago

Comment by samizdis 1 day ago

> ... once an open platform that people can customize becomes the industry’s center of gravity, no single vendor can match the combined rate of innovation around it

That assertion is sublime, and if not true, it should be and can be.

Thanks for the thought (and the optimism hit).

Comment by PersonalJarvis 2 days ago

Comment by jitbit 1 day ago

You can contribute to an OSS platform. To accellerate innovation, adoption etc

You CANNOT contribute anything to a model.

Comment by 2 days ago

Comment by Sammi 2 days ago

Kubernetes is a system/infrastructure orchestration tool. I completely fail to see how it is comparable to open weight neural nets. In either application or function.

I'm sorry to do that hn comment thing where we all just race to contradict or talk in opposition of whatever was said before. I'm aware. But really guys, was this article really not just a miss?

Comment by Kbelicius 2 days ago

It is natural to not see something that isn't there. The article isn' comapraing the two in application or functio

Comment by soizi 1 day ago

like most of people I think you didn’t read the article entirely.

Comment by cmdrk 1 day ago

Drat, I was hoping this article would be about outlawing Kubernetes.

Comment by Alien1Being 1 day ago

Meanwhile America is having its glorious McCarthy moment....

Comment by SilverElfin 2 days ago

What’s interesting is no one is talking about political censorship in models and how DeepSeek, Kimi, and the rest have to abide by CCP rules. It’s a big opportunity for China to control information.

Comment by Danox 2 days ago

And that censorship will fail too, like sanctions, tariffs, and keeping down open source software or exchanging ideas across borders, ultimately it will fail, but that won’t stop governments across the world from trying.

Comment by Covenant0028 1 day ago

The primary customers of these models are enterprises, and the most common use cases are office work and coding. How often do the questions of Tiananmen Square or the Uighurs become relevant in those contexts?

Comment by dboreham 1 day ago

It's going to be so complicated that nobody quite understands how it works?

Comment by nunez 1 day ago

I respect Knaup's opinion but disagree with the claim he makes here.

Kubernetes took off because everyone could run it on pretty much anything. Bigger hardware meant bigger clusters, but developers could spin up a cluster on their laptops and deploy their apps into it. More importantly, companies could repurpose their decommissioned servers as k8s clusters, a massive unlock seeing how huge companies had heaps of these in their racks.

This isn't possible with open weights models. Not in the same way.

First, you're out of the game if you don't have a data center class GPU (or it's sort of prosumer equivalent). Model servers support CPU inference, but you might as well watch paint dry as you wait for results...and you'll still have to run super quantized low-parameter models that aren't as good.

Realistically, companies will need to purchase millions of dollars of nVIDIA gear (through suppliers) to serve agentic-capable models at scale, an activity that is being made more expensive and complicated by the day as the hyperscaleds slurp up the demand.

It also really is a huge problem that all of the open weights models are coming out of one country that also happens to be a superpower. I don't think this can be handwaved away, and it's concerning to see so many folks here minimize this.

American companies running on Chinese intelligence. As a country that prides itself on being the knowledge capital of the world, the optics manifested by this are horrible, not to mention the absolutely massive supply chain risks (the counterfeit Cisco devices comes to mind, except worse because the rangers are in the weights and might not be possible to distill out).

This is a threat even if you focus solely on the individual developer. Recall how AWS became...AWS. They "fanatically" focused on the developer experience. They designated this as key to their growth strategy, and rightfully so. How people talk about using Qwen or Kimi and the like on here feels like that (ignoring how these models are ALSO from huge for-profits).

The only solution here is for the big labs to make some of their most capable models open-weights. There is a lot of secret sauce around routing, inference, hosting and other stuff that makes them work as well as they do, but the community can figure that out. Of course, this is basically a death sentence to those companies, but that's what they get for playing with monkey paws, I guess.

Comment by memedesimo 9 hours ago

It is funny to see how you feverishly seeking "legitimate ways" to suppress successful concurents.

I would call you "sore losers", but as I know how insane you actually are, you might just start a war.

(Yes yes, they "stole from you", and "Tiananmen square", and "Uigur genocide", and.. and I just can't wait for the monent someone locks your crazyness behind the bars of an asylum.)

Comment by chrisjbg 2 days ago

Dario is a FUD-spreading douche

Comment by claw-el 2 days ago

Essentially doing this, https://health.clevelandclinic.org/catastrophizing while making other’s mental health worse..

Comment by cr125rider 2 days ago

Way over complicated for what most users need? Huh?

Comment by bmiekre 1 day ago

This seems like a pretty balanced and thoughtful idea. I’m HN will hate it

Comment by pich 2 days ago

[flagged]

Comment by hsienchuc 2 days ago

[flagged]

Comment by nttylock 2 days ago

[flagged]

Comment by debarshri 2 days ago

Shameless plugin. Funny enough we just made agents kubernetes native at adaptive [1]

[1] https://adaptive.live

Comment by petilon 2 days ago

Enormous amounts of money is being invested in the development of AI models. Investors expect returns on their investment or they will not continue investing. Open weights make it harder for investors to get their money back, so it harms the industry.

Once the weights are out, it makes no sense to ban them in the US while the rest of the world takes advantage of it. But that doesn't mean developers of frontier models shouldn't take steps to prevent their weights from being stolen.

Comment by danny_codes 2 days ago

Ridiculous. If investors wish to set their money on fire investing in over-valued companies, they are welcome to do so. Protectionism to preserve ROI is a dumb policy. Just look at the US car industry. We're building dinosaurs. On this trajectory we'll have 0% market share abroad in 10 years. Chinese EVs will probably get market share at a 100% tariff because US automakers fell so far behind.

Same situation in protectionism in AI. The rest of the world will simply lap us.

Another shill account I gues

Comment by petilon 1 day ago

FYI, developers of frontier models taking steps to prevent their weights from being stolen is not "protectionism". If you guard your wallet from being stolen is that protectionism?

Comment by jumpkick 2 days ago

How are weights stolen from the frontier model developers? What is that actual mechanism?

Comment by eigenspace 2 days ago

The weights themselves aren't stolen. The claim is that Chinese companies are using VPNs and proxies to buy massive amounts of Claude Pro and Codex subscription accounts, and then selling usage on those subscriptions as cheap white-label LLM API usage.

While selling that LLM API usage, they then capture all the prompts, outputs, and intermediate thinking the LLM does, and then sell those logs to the companies making open-weight models. The open-weight model developers then train on those logs to 'distill' a model.

Comment by ForHackernews 2 days ago

...good for them?

We used to call that competition.

Imagine making this argument with a straight face in any other industry:

"The claim is that Japanese car companies are buying Ford vehicles, and then leasing them to American consumers at cut-rate prices. In return for the cheap cars, the customers are letting the Japanese observe their driving behavior, studying how they use their F-150 and then the Japanese car companies are applying that data to design new vehicles that will directly replace Ford!"

Comment by eigenspace 2 days ago

I wasn't arguing against the practice, I was clarifying that the weights of these models are not being stolen, and explaining what these companies do to create their models.

Comment by petilon 2 days ago

> observe their driving behavior

That's not what they are observing, it is the behavior of the engines.

Comment by ForHackernews 1 day ago

:sigh:

fine, s/observing the way the Ford cars work

Comment by jumpkick 2 days ago

The parent I replied to said frontier weights were “being stolen” which is not the literal case. I think precision is important here.

Comment by eigenspace 2 days ago

I agree. That's why I said that the weights were not being stolen, and then explained what is actually being done.

Comment by singingtoday 2 days ago

Fantastic, thank you for the explanation!

Comment by kalu 2 days ago

The sentiment in this article is nice. But open source software is a weak analogy for frontier models. Principally because software requires zero capital investment (actually zero) while frontier models demand billions. Open models can only survive in the long run if they can (eventually) generate significant cash flows or if they are paid for by governments. Now China essentially has a monopoly on open weight models. And so supporting open source models means either supporting long term economic capture by China or supporting Chinese government control of your intelligence. Both of these outcomes are unequivocally bad from an American perspective. If you live in the valley and benefit from the US venture ecosystem you should be highly skeptical of open weight models. Banning them may very well be the best course of action.

Comment by Danox 2 days ago

They don’t have a monopoly and if they do, it would only be because the EU and the United States let them or I should say they let all their decisions be made by private companies who scraped the public Internet and now want to fence it in for their up-and-coming IPO.

Comment by danny_codes 2 days ago

Hilariously bad take. Open weight models can be retrained of fine-tuned, that's the entire point. The idea that the "Chinese government controls your intelligence" is laughable in the case of open weight models. Once the weights are released you can do whatever you want with them. The idea that there's economic capture by the Chinese for products they're literally giving away is stupid to the point of inanity.

I can only assume this account is pure shilling for the closed-source AI labs.

Comment by noncoml 2 days ago

Open Source is free as in speech

Open Weights is free as in beer

Comment by singingtoday 2 days ago

Ah, a true scholar!

Comment by 2 days ago

Comment by 2 days ago

Comment by pianopatrick 2 days ago

"Open source" open models can also survive if they are seen as a necessary cost of doing business. Same logic as tech companies working together on Linux. "We need an operating system. But an operating system is expensive to make on our own and does not really provide an edge. So let's just work on and use the open source one."