Hetzner is working on LLM Inference
Posted by jonas_scholz 16 hours ago
Comments
Comment by swiftcoder 14 hours ago
Comment by eliaskg 12 hours ago
It works great.
Comment by hoppp 13 hours ago
Comment by Aldipower 13 hours ago
Comment by beernet 7 hours ago
All of these are lightyears behind US/China LLM offerings. None of them offer any model close to open-source SOTA. Runtime and reliability is a disaster and sure enough you have to pay a "sovereignty" mark up.
Comment by expedition32 4 hours ago
Comment by frizlab 3 hours ago
(Also no, modern science is (still) not incompatible with (the Christian) religion. I want my children to have both, as do I, or at least be taught both, so they can decide for themselves what they want when they are grown enough to do so.)
Who is sacrificing what, really?
Comment by beernet 4 hours ago
Comment by someone4958923 12 hours ago
Comment by rf15 9 hours ago
Comment by ghaering 9 hours ago
Comment by npodbielski 11 hours ago
Comment by rf15 9 hours ago
Comment by jonas_scholz 14 hours ago
Comment by sofixa 14 hours ago
Comment by jonas_scholz 14 hours ago
Comment by PaoloBarbolini 10 hours ago
Also, they have a feature request that has been getting many votes recently: https://feature-request.scaleway.com/posts/1251/prompt-cachi...
Comment by t857295612 11 hours ago
Comment by atherton94027 12 hours ago
Comment by swiftcoder 10 hours ago
That's a pretty big part of why, honestly. How many of us are running server workloads that actually need the price/performance tradeoff of cutting edge server hardware?
Comment by atherton94027 7 hours ago
Comment by yjftsjthsd-h 4 hours ago
Comment by idiotsecant 5 hours ago
Comment by atherton94027 4 hours ago
They're great for hobbyists but people on here keep bringing them up as the flagbearer of euro hosting
Comment by MathiasPius 12 hours ago
Comment by atherton94027 12 hours ago
Comment by ffsm8 11 hours ago
Understandably too, because hetzner is running things on consumer hardware vs the enterprise hardware the others use
Comment by ano-ther 16 hours ago
Comment by drcongo 13 hours ago
"highly unmanaged" didn't fill me with confidence, and "uses our API" is a very weird reason to give for not offering support. I wrote off ever using them for anything beyond email after that reply.
Comment by jonas_scholz 15 hours ago
Comment by archerx 15 hours ago
Comment by sureglymop 13 hours ago
Worse even, in my account panel I still had control over the e-mail service of the domain, so there was most definitely a pretty bad security critical bug there.
Comment by MASNeo 14 hours ago
Comment by sam_lowry_ 14 hours ago
You will at least not need months of training and tough exams to be an expert in Hetzner Cloud. It is simple, but that's the point.
Comment by jonas_scholz 13 hours ago
Comment by drcongo 13 hours ago
Comment by Saris 13 hours ago
Comment by mark_l_watson 14 hours ago
Comment by jonas_scholz 14 hours ago
Comment by mips_avatar 12 hours ago
Comment by pmg1991 9 hours ago
I'm waiting for that day so that inference will be affordable just like web hosting. 200$ per month is in no way affordable by everyone.
Comment by embedding-shape 16 hours ago
Straight up the opposite, which the name makes abundantly clear, with the option it does reasoning, without it it doesn't...
Comment by jonas_scholz 15 hours ago
Comment by cyanydeez 14 hours ago
This allows both the client and server to customize it. I typically use a message that tells it to compact the conversation and use subagents. I find the reasoning gets bloated when it's failed to do whatever task it's doing and often times it either has too little context (subagents) or its context is bloated (compact).
This works fairly well to get it to extend workable life up to ~1M on a local model.
Comment by perelin 12 hours ago
Comment by danlitt 11 hours ago
Comment by jonas_scholz 10 hours ago
Comment by danlitt 10 hours ago
Comment by jonas_scholz 8 hours ago
Comment by toomuchtodo 10 hours ago
Comment by danlitt 10 hours ago
Comment by toomuchtodo 9 hours ago
Crucial Memory, historically a retail RAM provider, closed to focus on AI compute sales, for example.
Micron Announces Exit from Crucial Consumer Business - https://news.ycombinator.com/item?id=46137783 - December 2025 (392 comments)
> “The AI-driven growth in the data center has led to a surge in demand for memory and storage. Micron has made the difficult decision to exit the Crucial consumer business in order to improve supply and support for our larger, strategic customers in faster-growing segments,” said Sumit Sadana, EVP and Chief Business Officer at Micron Technology. “Thanks to a passionate community of consumers, the Crucial brand has become synonymous with technical leadership, quality and reliability of leading-edge memory and storage products. We would like to thank our millions of customers, hundreds of partners and all of the Micron team members who have supported the Crucial journey for the last 29 years.”
Micron is killing Crucial SSDs and memory in AI pivot to serve on AI companies - https://news.ycombinator.com/item?id=46152915 - December 2025 (0 comments)
> Secondly, the supply environment has changed permanently. AI infrastructure requires every single wafer with memory it can consume, something that has never happened with any industry megatrend previously. This means that every wafer Micron assigns to consumer parts is a wafer not going to a hyperscaler or enterprise contract. As a result, keeping a consumer line would directly limit Micron's ability to fulfill orders from its largest customers, which is a risk for profits and strategic relationships.
So! Until this AI capital investment exuberance ends, compute hardware manufacturing pipelines are committed for years into the future, we're all bidding up the limited supply of compute for rent/lease, which in turn is constrained by the limited supply of compute hardware available for purchase.
(at least in the scope of RAM, price fixing is also a potential component, but it'll take some time to determine how much of that is contributing to price inflation versus bona fide supply constraints: https://finance.yahoo.com/technology/articles/ram-crisis-one...)
Comment by NetOpWibby 11 hours ago
Comment by satvikpendem 11 hours ago
Comment by cousinbryce 11 hours ago
Comment by Havoc 13 hours ago
Comment by jonas_scholz 13 hours ago
Comment by _pdp_ 11 hours ago
Comment by nubg 14 hours ago
> For now, the API is fast, free, and fun to try. The next hardware announcement will tell us much more than another small model would.
Comment by jonas_scholz 14 hours ago
Comment by nubg 13 hours ago
Comment by jonas_scholz 13 hours ago
Comment by nubg 12 hours ago
Comment by petesergeant 11 hours ago
I would much rather read someone's carefully written and LLM-assissted article than shallow and lazy dismissals like this one. What, precisely, do you feel you've added to the conversation here?
Comment by nubg 7 hours ago
Comment by petesergeant 7 hours ago
Comment by jonas_scholz 6 hours ago
Comment by petesergeant 6 hours ago
Comment by scoriiu 14 hours ago
Comment by dk970 11 hours ago
Comment by rebelde 13 hours ago
Will this be the new division of labor?
Americans - best proprietary models
Chinese - best open weight models
Europeans - best / most efficient inference service
Comment by andsoitis 13 hours ago
Maybe Nvidia's Nemotron can get there.
Comment by 9dev 12 hours ago
What we need is frontier open weight models in the EU, Canada, and Australia.
Comment by slashdev 12 hours ago
Comment by 9dev 11 hours ago
Big, slow wheels have turned everywhere around the world. They won't turn back so easily.
Comment by ForHackernews 13 hours ago
Comment by andsoitis 13 hours ago
a) broad competition is good
b) jurisdiction diversification (you don't want to be dependent on the regulatory winds of a single jurisdiction)
Comment by ForHackernews 13 hours ago
If I'm running a Kimi model on my local cluster, the CCP can't shut it off.
It's not like the Finnish government has veto power over Linux.
Comment by andsoitis 13 hours ago
It is not beyond imagination that a government can limit availability to future frontier versions of the open weight models produced within their jurisdiction.
It pays to be a little paranoid and have a diversified portfolio of options to choose from. Anything that can have major geopolitical consequence is especially susceptible to such structural risk.
This has nothing to do with any particular government/country/jurisdiction except insofar as said country is a leading power and hence has immense leverage across a very wide range of dimensions. Said more plainly, it is safe to assume that a country with immense power is unlikely to be willing to cede or distribute such power to others.
If you're hung up on my choice of "the West", consider that I mean it geopolitically and economically, so broadly includes North America, Europe, Australia, New Zealand, Japan, and South Korea. I could have also said the Global South, but I think it is a fair statement that none of those countries have the means, but I'd be totally cool and happy if I'm proven wrong!
Comment by frangonf 13 hours ago