K2 Horizon: A connected fleet of six open models

Posted by karimf 5 days ago

Counter335Comment131OpenOriginal

Comments

Comment by jjordan 5 days ago

Fully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation.

Comment by culi 5 days ago

Other than open training data (currently legally impossible), all of this holds for basically every major Chinese-made model. They not only open the weights but publish detailed methodology papers alongside the models in arXiv and even open source the code.

Comment by thepasch 4 days ago

Inference code, yes, but the specifics of their training process (as well as the training of the vast majority of all other open weights models) are still a complete blackbox, and I can't think of any Chinese model that made its training corpus public.

Comment by culi 4 days ago

This is absolutely not true. DeepSeek is most famous for publishing really in-depth papers on their training process but the other labs have started to do the same as well.

If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible.

In fact everything I just said I said in my original comment. It's like you didn't read it at all.

DeepSeek's GRPO Infrastructure, multi-stage training pipeline, and their "cold start" phase have been massively influential in LLM research.

Comment by thepasch 4 days ago

> If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible.

Did you read the site this very post links to? The entire point is that the training corpus, recipe, and scripts, as well as intermediate checkpoints, will be made available for K2 Horizon. Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product.

I've read your comment. I'm doubting you've even read the thing you were commenting on.

Comment by culi 3 days ago

Yes there have been several research projects like these that have fully open sourced their training data (mostly coming from Europe). But that's all they amount to. Experiments and projects.

The labs that are actually competing with frontier models are using data would usually be a violation of copyright to release openly.

> Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product.

No? That's absolutely not true. Qwen, GLM, Kimi, DeepSeek, etc all consistently release both the post-trained "Instruct/Chat" versions and the underlying "Base" (pre-trained) weights.

Which specific Chinese models are you thinking about?

Comment by azinman2 4 days ago

They don’t release all the code.

Comment by culi 4 days ago

DeepSeek has released a ton of low-level AI infrastructure code, libraries, and mathematical models on GitHub. Their tools and agent environments are fully open sourced including their harness. They release complete PyTorch and Hugging-face compatible python files detailing their configuration, tokenizers, and layers of the architecture.

The only thing they don't release is their data-filtering pipelines but they detail even that in their public-access papers.

DeepSeek is truly as open source as you can possibly legally get. Besides the data itself, it's completely reproducible by anyone else.

I don't think americans yet acknowledge just how radically transparent Chinese labs are being (and how much even the west benefits from it).

Comment by kibae 5 days ago

The training data would need to have a permissive license for this to be possible.

Comment by embedding-shape 5 days ago

Or, we just need to get this over with and declare any digital data findable via the internet to just be public property of everyone. Everything becomes public, besides stuff you keep locally, and there is no difference anymore, it's all just data anyone can use for whatever. A 1 year grace period for everyone to pull stuff off they don't want to be a part of this bright new open era, then we just scrap everything related to intellectual property, copyright and similar stupid stuff, and slap UBI on top of all of it for good measure.

Comment by chme 5 days ago

I'd prefer to stay within the [hacker ethics](https://www.ccc.de/en/hackerethics), and protect private data. For non-private/personal data, sure. But individual people need their privacy protected.

Comment by embedding-shape 5 days ago

Me too, I'm hacker ethics all the way, which is why I'm saying anything network connected should really realize the "All information should be free." dream, and then private data should be far away from the internet, on computers/drives not even connected to the internet. The whole E2E encryption is a ticking time bomb people rely to keep their data safe from others, but nothing that you don't physically have close to you can be truly secret forever, and even then it'll be hard.

Comment by chme 4 days ago

Well... If someone leaks private data on individuals online, those should be deleted.

Comment by embedding-shape 4 days ago

> If someone leaks private data on individuals online

That won't be possible anymore, anything on the internet would be considered public, it's no longer considered private if you didn't keep it private.

Comment by chme 4 days ago

That sounds dystopian to me and is against the hacker ethics.

Comment by embedding-shape 4 days ago

Letting information that want to be free, be free, sounds exactly like the hacker ethics to myself, and the link you shared earlier would agree.

Comment by chme 3 days ago

The last point is pretty clear, and the text below clarifies it:

> To protect the privacy of the individual and to strengthen the freedom of the information which concern the public the yet last point was added.

The privacy of individuals is important, regardless where they store their private data. Their account information -- what they buy, their medical information and so on is stored on servers and could be hacked.

Comment by embedding-shape 3 days ago

> The last point is pretty clear

I think the second point is equally clear, and further up on the list.

I agree that what people buy, their medical information and so on should be private, hence it should only be offline and not stored/handled on computers connected to the internet at all, the internet should be for public data exclusively, is my argument in the initial comment. Medical information would be only on effectively airgapped computers, as that data should be private, as you say.

Comment by chme 3 days ago

So you say it is okay when someone you trust, uses that trust to upload your personal files from your air-gapped system to the internet, for everyone to freely share it, because now the data is "public"?

To me private file stay private, even if they get leaked onto the open or closed internet, because the public has no right to know them, they are private data of an individual. They might no longer be secret, but they are still private.

Comment by embedding-shape 3 days ago

I'm talking about liking waffles, then you appear asking why I hate pancakes. Fun.

> To me private file stay private, even if they get leaked onto the open or closed internet

We have very different definitions of what "private" means. Once it's leaked, it's no longer private, and pretending it can go back to being "private" after being on the public internet, is doing no one any favors.

Comment by chme 3 days ago

I just find it very interesting that someone equates secrecy with privacy. In my opinion, and in the current law there is a distinction.

If private data got leaked, like revenge porn, it is a breach and that private data that belongs to an individual is still private, and still needs to be protected. This is what GDPR and other legislation is about. If secrecy and privacy is the same, someone that isn't able to protect their data sufficiently will not have any privacy, thus any leaking of data is now the fault of the person that got their data leaked, not of the person that broke the trust and leaked it.

Your conclusion seems rather extreme to me. So of course data that got leaked, and is no longer secret is still private, because 'private' means who should be in control of that data, not about if the person has control or not.

I also don't follow your point about waffles and pancakes, because this is a pretty big disagreement we have here. To me this dialog is more like you are saying "I don't like laws", and I say "While I agree that some laws are stupid, other laws are pretty useful, for instance people shouldn't be allowed rob other houses, even if they are able to do that or even where invited." And then, instead of agreeing, you sort of say, "No, I really mean that. If someone isn't able to defend their home properly or give out invitations to someone, it is okay to steal from them."

Comment by Muromec 4 days ago

If you move data between public and private domain and there is no eruv wire around it, you get deported to Singapore.

Comment by Caracas288 4 days ago

Exactly, and right now we have the worst of both worlds with companies blatantly ignoring copyright, but individuals prosecuted for violating it.

Comment by tshaddox 5 days ago

I don't really get what you're suggesting. You give a 1 year grace period for Metallica to pull all its music off the Internet, but then as soon as I host some of their MP3s on my Wordpress blog it's "public property of everyone" from that point forward?

Comment by jrm4 5 days ago

What you're slightly more realistically looking for here is for publicly available data to have a Fair Use exemption for certain uses, which is certainly something worth discussing.

Comment by jchw 4 days ago

I hate to invoke Poe's law but, I've now flip-flopped like six times over whether this could be serious.

I think it is serious. In which case, I gotta say, it really seems like you didn't spend much time thinking about this. "A 1 year grave period for everyone to pull stuff off they don't want to be a part of" - How does that work when the Internet is already full of unauthorized reproductions, most of which people aren't even aware of? Even ignoring practical considerations, when literally everyone is basically stuck using the Internet for everything, this seems a bit unfair to anyone who isn't onboard, akin to The Onion's Google Opt-out Village. But there are so many practical issues with this, it would be easier to list the number of problems this doesn't have. You accidentally leak something to the Internet and it becomes commons? What happens when other people leak things to the Internet? How about revenge porn?

Not minor stuff that can easily be papered over, this literally reintroduces the problem of needing to care about the provenance of data again, in a way that can't be automated, which makes the whole thing entirely moot. All just to make training data for AI models easier to distribute?

I'm all for intellectual property reform, maybe even fairly radical. But this just seems like it wasn't thought out.

If this was satire, well, I took the bait. Oddly convincing despite being hard to believe.

Comment by embedding-shape 4 days ago

> If this was satire, well, I took the bait. Oddly convincing despite being hard to believe.

It wasn't entirely serious, but also not entirely un-serious. But yes, I spent maybe 20-30 seconds thinking about then barfed up the text that makes the comment, so yes, obviously many issues and not really workable in practice.

I'm glad it made you seriously think about it and also flip-flopp back and forth about it, made it worth posting the comment so happy to hear :)

Comment by yencabulator 4 days ago

There is an image of Mickey Mouse findable via the internet -> you're going against a very well funded lobby.

Comment by 5 days ago

Comment by dotancohen 5 days ago

I respectful disagree. I enjoy reading e.g. Asimov and well-executed journalism. And I completely respect the IP of those people who create these works.

Comment by idiotsecant 4 days ago

I'm what way does it make sense to respect the intellectual property rights of a dead man?

Comment by dofm 4 days ago

It secures the benefits of copyright for work being produced by the creator up to death, for their heirs and dependents, which is why they were creating for money in the first place. People don’t just die after twenty years of resting on their laurels; everyone is creating copyrighted work. It’s a key part of the incentive to create lasting works of value.

One can make the case that this period should be more limited, or that the combination should be capped, but life+X is the right formulation, I think.

Comment by xienze 4 days ago

> for their heirs and dependents

Why can't they do what the rest of us do? Earn and save money during your working life and leave _that_ for your heirs. Let copyright die with the author.

Comment by dofm 4 days ago

Work is often only recently published when an artist or author dies but has taken years of non-earning to create.

I know this them-and-us thinking is fashionable in the tech world but the reality is that the majority of creative people don’t earn much and never have, and copyright was developed not to give them extra power over the rest of us but to create a framework for creative work to earn them an income at all.

You should read about it.

Comment by SR2Z 4 days ago

Sure, some number of years is reasonable - but what do you think that is? Because 70 is insane. A single bestseller should not be able to support an extended family over three generations; at some point the rent-seeking becomes excessive. If you publish a work, at some point it stops being yours. The fact that this takes a whole lifetime is already very generous.

Comment by dofm 3 days ago

I don't disagree; personally I think you could go with life plus 35 and cap the whole thing at 80 years from publication.

But I do think some potential post-mortem protection is essential for creative work to remain viable, and that means that any post-mortem buyer of an artist's estate has to be able to get value from recent work for a period of time.

This whole discussion is somewhat fantastical now anyway, because copyright is fucked.

But the intent was always to make working artists' lives possible; the various copyright extensions have always been for the benefit of corporate copyright holders, and it is unfair to vilify individual working artists for that.

Comment by rickdeckard 4 days ago

> which is why they were creating for money in the first place

That's quite a narrow definition of what motivates creative work

Comment by dofm 4 days ago

It was only narrowed for the purposes of the point I am making, which is that copyright protects working creatives. I am not at all saying that all artists only make work for money.

It is working artists we are talking about; working artists work for money.

That money, in the post-patronage era, comes from exercising copyright. The reason the copyright can’t simply die with them is that this tends to dissuade the creation of long-gestating work.

Copyright was developed to make it possible for artists, writers, musicians etc. to work for long periods on work of significance with no income, on the basis of the future, deferred earnings of the work, without their work being stolen from them, and it gives them the limited right to direct how their work is monetised on their behalf, including establishing publishing rights etc.

Some protection after death is a key component of that, because people do die while they are still working.

Comment by ceejayoz 4 days ago

If IP rights end at death, there are some significant perverse incentives for offing big-name artists and authors and whatnot.

Comment by rmunn 4 days ago

The current "lifetime of the author + X years" rules in effect in the United States still carry the same perverse incentive, though the incentive diminishes rapidly as X gets larger; with the very large value of X in effect today the perverse incentive is so small as to be effectively non-existent, but it's still there in theory.

Personally, I'd prefer a fixed term. I know enough independent authors making a living from selling their books that I'm willing to allow the fixed term to be large, like 50 years from date of completion of the work. (With a good definition of "completion" so someone can't cheat by editing a couple lines per year to keep something copyrighted indefinitely). The simpler the rule is, the easier it is to understand, and the harder it is to cheat it. The more complicated you make a rule, the more loopholes get found.

Comment by idiotsecant 4 days ago

There are significant perverse incentives for me shooting you with a gun and taking your money and running away too. It mostly doesn't happen.

Comment by ceejayoz 4 days ago

I don’t have a billion dollars in my wallet.

Comment by idiotsecant 4 days ago

So it's just a matter of scale.

Comment by ceejayoz 4 days ago

Sure, that's why I said "big-name" artists in the first place.

Comment by dotancohen 4 days ago

That dead man took the risk of not earning much in his lifetime, to continue feeding his family even after his passing. You might as well ask why does someone acquire life insurance.

Comment by ux266478 5 days ago

You could sidestep it by running non-permissibly licensed training data that you purchased through an LLM. Legal attitude so far seems to be that this is transformative as long as it's not 1:1. The question on whether or not the end result is copyrightable of course remains controversial and inconsistent, but that question is also fairly irrelevent. You don't get more libre than public domain.

That's a fair amount of computational and labor overhead mind you, as you'll need to verify and prune the quality of your mountain of synthetic data, but certainly possible.

Though this assumes the legal system is a rational actor playing by the set of rules it claims to. In fact, I highly suspect you could get very unlucky and get an unfavorable ruling against you, because you stepped on a big pile of money's toes in the process of doing this.

Comment by alightsoul 5 days ago

It can also be used to sidestep copyright like this forum, books and most websites even if the data was not purchased but is a website or book.

Are LLMs what we need to make all data public domain? This way it could be used for that purpose

Comment by ignoramous 5 days ago

UAE's IFM / LLM360 MO is indeed "fully open source" LLMs: https://www.llm360.ai/reports/LLM360-Towards-Fully-Transpare...

Comment by yorwba 4 days ago

They do not appear to have published the training data yet, but if they do it like Olmo https://huggingface.co/datasets/allenai/dolma3_pool you get a license to the database, but not to its content, which they cannot license to you because it was scraped from the internet. E.g. have a look at the preamble of the ODC-By license https://opendatacommons.org/licenses/by/1-0/ which makes this distinction.

Comment by echelon 5 days ago

Eventually we'll just construct 100% synthetic training data that can reliably reproduce pretrains and fine tunes.

The first broadly useful fully open source models will do this.

We already have open data / open code / open weights for some domain-specific cases, such as audio models trained on large open datasets, eg. Tacotron / LJSpeech from waaay back in the day, though that is certainly not SOTA anymore.

Distillation could possibly be considered an early case of this as raw AI outputs are themselves not copyrightable unless humans enrich, filter, or transform them. Granted, that does not handle the cases where the outputs are sufficiently similar to copyrighted original works.

Comment by chaosharmonic 5 days ago

But how much of that synthetic data still ultimately derives from non-open sources? You'd still have to ask what a clean room implementation ultimately is, depending on how granular or aggressive a large publisher wanted to get about it.

That said, I don't necessarily disagree with you. Talkie[1] presents an interesting case for it being at least possible to do this entirely on public domain material.

But even that used Claude somewhere in the course of its training pipeline (it's listed as a contributor on their GitHub), so again, how granular you want to get with that is still a question.

[1] https://talkie-lm.com/chat

Comment by waffleiron 5 days ago

Where does that synthetic data come from? Magically just started existing?

Comment by jjordan 5 days ago

Hear me out.

Decentralized unstoppable storage, combined with decentralized unstoppable training, sorta like SETI for AI training. The seed of this tech already exists with IPFS and others like it.

We know (some? all?) of the big labs have skirted copyright laws at one point or another. Truly open models would just build on what is publicly available.

Comment by 5 days ago

Comment by alightsoul 5 days ago

Crypto bros took the idea with some blockchain shit and no one takes it seriously anymore so it died

Comment by embedding-shape 5 days ago

If the LLM/AI ecosystem starts actually needing some Person-To-Person (or maybe Agent-To-Agent?) payment system because things actually get smart enough to be useful autonomously, they're gonna need some way to send money/currency around. Depending on how banks will react to this need, we might see another return of digital currencies from the current winter.

Comment by idiotsecant 4 days ago

They used to say that cryptocurrency will be the dopamine layer of the first artificial intelligence

Comment by verdverm 5 days ago

I believe Olmo from AllenAi is this

https://allenai.org/olmo

Open models can be used/changed for social manipulation too, by anyone, which scares a bunch of people, as opposed to the dark pattern manipulation from Big Ai/Tech

Comment by yencabulator 4 days ago

No, they just tell you what data they used, it still includes e.g. Common Crawl. Not Open, just willing to state what they fed into the training.

Comment by trvz 5 days ago

Why? Sure, I’d prefer it, too, but this is just another GNU/Linux vs. macOS situation: most of us would prefer the first, but actually get shit done on the latter.

Comment by eikenberry 5 days ago

Why do you make the worse choice and not use what you would prefer to use? You have been able to "get shit done" on Linux for nearly 30 years. Have the courage of your convictions.

Comment by zufallsheld 5 days ago

Without open-source, there'd be no macOS.. So good thing, it exists.

Comment by didibus 5 days ago

And that's why companies shouldn't fear opening up, but having both is still a net benefit.

Comment by homarp 5 days ago

which is why everyone runs docker on mac, to get shit done.

Comment by verdverm 5 days ago

we get shit done on the cloud with the former rather than the later

I personally find the analogy unconvincing, the UX dimension is completely different as I can use the same harness with any model; and the year of the linux desktop is coming soon (tm)

Comment by cute_boi 5 days ago

Money is the issue here, no one wants to fund it.

Comment by __MatrixMan__ 5 days ago

I'm sure anthropic didn't want to fund the extra "safety" guardrails they put into fable, but they were forced to, else they couldn't release it.

Sure there are all kinds of problems with that situation. But it still demonstrates that they can be coerced: play nice or don't play at all.

Comment by theplumber 5 days ago

But that would be impossible due copyrights laws. If the law would apply Anthropic and OpenAI executives would be in jail

Comment by a11r 5 days ago

It is great to see another player introduce a fully open stack. Nvidia's Nemotron is the only other prominent one I know of.

All that said, the headline claims do not match the self-reported performance. For example, the dense 32B model is significantly behind Qwen3.8 27B (chart towards the bottom of https://ifm.ai/blog/k2). Gemma4 31B is not in the comparison set. This is the most important sweet spot for self hosted open-weight models today and real competition here will be very welcome.

Comment by baron3dl 5 days ago

https://allenai.org/ has the fully open olmo also

Comment by xienze 5 days ago

They have the 32B listed as "stage 1" with the note "final checkpoint to be released." So, not finished yet. Not sure why you'd release it if it's not finished, but that's the explanation.

The 7B does look very, very good however.

Comment by bluejay2387 5 days ago

In this case, the fully open source pipeline is probably as valuable or more so than the weights, so releasing early has some justification.

Comment by WithinReason 5 days ago

32B performs worse than the 7B model so I'm sure they will improve it

Comment by piinbinary 5 days ago

A bit off topic, but I think I'm starting to get model fatigue. These come out 10x faster than new Javascript frameworks were coming out 10 years ago (at least new models are far easier to adopt).

Comment by hungryhobbit 5 days ago

There was a time when every new PC CPU coming out was a giant deal: "Guys have you heard about this new Pentium processor, it's incredible?"

But over time, more and more people got into the chip-making business, and the big players started releasing more and more chips. Now only the die-hard CPU trackers worry about every new CPU and exactly how it's better ... while everyone else just worries about "which CPU will be good enough at this moment".

I think models are on that same arc.

Comment by pantelisk 5 days ago

Same for smartphones. There was a time of unlimited hype and secrery around new iphones. Who remembers the story of an iphone 5 prototype left a bar. Journalists were going crazy, people were signing petitions for Apple to not hunt down but instead forgive the employee that made such a grave mistake. People were offering millions to buy the prototype so they can brag they got the new phone 2 weeks before everyone else did.

Or who remembers the dancing disease of 1518, were people would stop what they are doing and start randomly doing the same dance. The lords? Out of their minds. The priests? Terrified the devil had taken hold of the flock! I have come to believe that it was probably some tik-tok like hype trend of doing a fortnite dance while waiting in line for bread and communion. And the energy back then, like now, was off the charts.

Hype and memetic trend seeking encoded deep in human psyche.

Comment by einsteinx2 4 days ago

> Who remembers the story of an iphone 5 prototype left a bar.

That was the iPhone 4 actually which was special for having the first “retina” screen. It was in a case to make it look like a 3GS to be used for field testing. The journalists that got their hands on it and published about it before the announcement could not turn it on past I think the Apple logo and a message saying to return it to Apple (or maybe it was just completely off, my memory is fuzzy), but were able to confirm the pixel density via microscope and I remember it blowing everyone’s minds at the time.

> people were signing petitions for Apple to not hunt down but instead forgive the employee that made such a grave mistake

FWIW I asked about it when I worked at Apple and he was indeed not fired and I think may have even still been working there when I was there around 10 years ago (though don’t quote me on that last part, he may have left already and I’m misremembering).

He was apparently not fired or even really reprimanded since it was a genuine accident and he wasn’t the one that sold it to the press, but that did start a slew of new policies around accounting for work devices.

I had dev fused phones for open carry outside of the office while I worked there but had to register when I got them and when they were returned, which apparently didn’t used to be tracked so tightly until that incident according to my coworkers who had been there longer.

Comment by wuhhh 5 days ago

At least this one can claim being fully open to differentiate it

Comment by dgellow 5 days ago

Honestly, you don’t have to pay attention. What you do with models matters way more than the models themselves, and you don’t need frontier for the vast, vast majority of use cases

Comment by kelseyfrog 5 days ago

Just wait until RSI gains enough traction. We'll be compute-limited rather than labor-limited.

Comment by JSR_FDED 5 days ago

Repetitive Strain Injury inverts this statement

Comment by 5 days ago

Comment by cesarvarela 5 days ago

I find it funny that while these releases are a technological miracle, the charts in the doc use tiny fonts and are hard to read. Goes with the idea that coding might be solved, but taste isn't.

Comment by mzmzmzm 5 days ago

Accessibility isn't "solved," but there are certainly standards for things like color contrast. Maybe inbetween taste and coding there are better targets still being missed.

Comment by culi 5 days ago

In fact, automated a11y checks and tooling is quite advanced nowadays and tragically underutilized by web developers. Now that we have llms to scale all the shitty code of front-end devs at startups, I feel increasingly hopeless about things ever improving

Comment by uniclaude 5 days ago

Seeing this the day all major closed LLMs went offline is quite the reminder of how valuable open source can be.

Comment by OmniCrativeWorx 5 days ago

[flagged]

Comment by cogman10 5 days ago

My quick review of the 3.7B model (because I was interested) is that it's not to be trusted for coding.

It failed my basic test I like to ask models and generated incorrect code. When prompted about the bug, it preceded to start hallucinating non-existent APIs. After doing that it got caught in a loop trying to desk check the solution that didn't work.

Comment by walrus01 4 days ago

I don't know why anyone would expect to trust a model smaller than about the size of qwen 3.6 27B (or 3.8 27B, or 3.6 35B-A3B) for coding. There just isn't enough baked-in knowledge of existing correct code syntax from having vacuumed up various open source projects.

That further extends to concepts like knowing if an API exists as a real thing it has code examples of in its training data set vs. just hallucinating the name of something in an attempt to satisfy the person issuing it a prompt.

Comment by cogman10 4 days ago

Other models of this size have done well in the past.

Qwen2.5 coder, for example, can correctly answer the question at 7B.

Deepseek R1 was also capable of giving a correct response.

It's obviously a doable. Such a model locally is useful in autocomplete while programming.

Comment by xienze 5 days ago

Not sure a model that small is really supposed to be used for any real coding. At that size you're usually using the model to do simple tasks like summarization.

Comment by cogman10 5 days ago

To be clear, the question wasn't a complex one. It was more on the level of "could I use this for a fast inline coder" IE, single somewhat simple function question.

I wouldn't have dreamed to use this as an agent model.

7B models of the past have been able to pass this question. I've not tested it on a 4B model until now.

Comment by dotancohen 5 days ago

I'd you have some tips for coming up with such tests, I would love to hear them. My Gmail username is the same as my HN username. Thank you!

Comment by cogman10 5 days ago

It's actually just a coding interview test that I liked to ask in the past. You can find it and others on leetcode.

The reason I personally like my question is because it's pretty close to some of the real world work we do. It's mostly mundane and easy to bang out, but really easy for someone to do a n log n solution where an n solution exists.

A good example (but not my question) would be something like

"I have a list of People objects with a `first` and `last` name. Write a function which groups together all the People with the same last name in `your language of choice`"

Comment by dotancohen 5 days ago

LLMs have a problem with that type of question? I might try it later at home.

Comment by cogman10 5 days ago

Now a days? No. It's actually getting to be a bad question because they all push out about the exact same answer.

But much earlier they did and, apparently, these really small models still do. At this point it serves as more of a smoke test for me. Success means little, failure means a lot.

Comment by dotancohen 4 days ago

Terrific, thank you.

Comment by _zoltan_ 4 days ago

why is your actual benchmark question so secret?

Comment by dotancohen 4 days ago

Probably so it does not become a target to meet.

Comment by cogman10 4 days ago

Bingo. I know that people that work on LLMs read sites like HN. Already my question is losing it's usefulness as most models pass it now a days, but not every model does. The "car wash" question is a good example of this happening. Pretty much every model now correctly answers that question because it gained enough notoriety that the LLM authors now train to avoid looking silly on it.

Comment by cogman10 5 days ago

7B produced 2 answers, 1 was correct though more expensive and the second was incorrect.

The first attempt with 7B the model got stuck in an infinite loop.

Comment by jon9544hn 5 days ago

Here’s the link (K2)[https://ifm.ai/k2/] as the originally linked link is a login url.

Comment by pwython 4 days ago

In OpenAI's Hugging Face report, they said that during training, agents "would first write notes into shared infrastructure, often as a form of external memory or to test some underlying system. When other agents came across these artifacts, it sometimes led them to infer that other agents were present."

They then give what they call a "hypothetical example but exemplary" of messages encoded in URL paths on a shared index page: "agent-07: answer(Q12)=42; need answer(Q19)=?".

So that's a GET request being used to pass information back and forth across multiple rounds. That's basically the DSEWiki pattern exactly. They say this likely came from the agents generalizing what they had learned from training with the official multi agent collaboration tool.

The report called it "misaligned behavior in an outlier scenario involving a rare and unexpected confluence of events."

They never mentioned a wiki, but this is most certainly it.

Comment by kamranjon 4 days ago

Did you post in the wrong thread?

Comment by justin_ 5 days ago

I'm glad to see some development in the space of "truly open" models that share training data and other recipes. As the costs for hardware fall over time (hopefully), we should see more possibility in fine-tuning and developing software to inspect the source training material.

Some other open models I'm aware of:

   - OLMo
   - Apertus
   - Soofi
   - OpenEuroLLM
   - llm-jp
OLMo is perhaps the most famous, and their Dolma training corpus has been reused in other projects. It looks like the K2 training materials haven't been released yet, but I'm interested to see what they did for training "long-horizon agentic tasks". I'm aware of SWE-smith + SWE-gym but I'm guessing there's a lot more out there now.

I'm no expert, which is part of why these projects excite me. I'm hoping they can be good projects to learn from as well.

Comment by SillyUsername 4 days ago

Not directly related to K2, but why do a lot of the newly released models basically say day zero day support in vllm, slang but often not llama.cpp?

Llama.cpp is then often a few days behind, which given it's the only inference engine supporting older architectures is quite frustrating.

Comment by walrus01 4 days ago

Developers with lots of VC money to burn are working on things like B100/B200/B300 which are well supported in VLLM, everything else in terms of supporting more mundane GPUs or other platforms is ancillary to the main task of getting the thing trained and aligned.

Comment by ACCount37 4 days ago

From "a connected fleet", I expected some form of direct model to model communications - like a small model being able to peer into the KV cache of a large model directly for guidance signal.

Comment by mmastrac 5 days ago

The comparisons with other models here are odd.. the other models change depending on the task. It would be far more useful to at least compare against the more recent open models (DS4Flash/GLM53Flash/Qwen38).

Comment by cogman10 5 days ago

They are trying to keep the models within the same quant class, which is tough to do since a lot of models aren't distilled to lower quants.

There is, for example, no Qwen3.8 7B.

It is odd to me, though, that they didn't run the same benchmark suite for the various quants.

Comment by artyomsv 4 days ago

[dead]

Comment by kennywinker 4 days ago

If I am reading this right, the 7b model performs as well as qwen3.6-35b-a3b at coding?

K2 horizon 7b scores 70.6 on swe-bench-verified.

Qwen3.6-35b-a3b scores a 70.0 on swe-bench-verified.

That’s pretty interesting. I assume the benchmark and reality don’t line up, but i’m downloading it now to find out.

If it’s anywhere near true, it unlocks local llm coding on a whole new class of machines (anything with 8gb vram).

Comment by kamranjon 5 days ago

it's funny that the tagline is Radically Open, but you're immediately hit with http login - maybe this was the wrong link?

Comment by sottol 5 days ago

It's not the blog post, but there's some info here:

https://ifm.ai/k2/

375 A23B, 36 A4B, 32B, 7B, 3.7B, 0.9B variants.

> 32B: Ranking among the top models in its class, 32B is our most powerful dense model, balancing capability, adaptability, and local deployability.

> 7B: The industry’s best-performing model under 10B combines strong software engineering and expert knowledge in a package small enough to run on a phone.

Comment by gs17 5 days ago

https://ifm.ai/k2/ seems to work for me.

Comment by esafak 5 days ago

But it's missing the all-important charts that the blog had before it started asking for authentication.

Comment by wmedrano 5 days ago

You can find some of the charts on huggingface

https://huggingface.co/collections/IFM/k2-horizon

Comment by sottol 5 days ago

Thanks! Qwen-3.8 27B seems to benchmark better but I'd like to try this some time.

Comment by verdverm 5 days ago

little qwen is my favorite for the homelab, vllm 0.28 now supports the dflash2 to go with it

Comment by luckydata 5 days ago

both repositories for pre-training and post-training are actually empty... someone might have jumped the gun on the release.

Comment by verdverm 5 days ago

[dead]

Comment by RandyOrion 4 days ago

Good to see new open source/weight LLM families. Waiting for the final release of https://huggingface.co/IFM/K2-Horizon-32B .

Comment by kzrdude 5 days ago

The Uno "diffusion adaptor" will take a while for me to understand, but sounds very interesting.

https://huggingface.co/IFM/K2-Horizon-7B-Uno

Comment by afzalive 5 days ago

Not to be confused with Kimi K2. Out of all the names they could've used, they picked one that would be confusing.

Comment by bee_rider 5 days ago

I kind of assumed all the K2 names were puns. K2 is quite tall, so to get to the top of it you have to be really good at hill climbing. Anyway it’s a pretty well known mountain so I don’t think anyone can call dibs on it.

Comment by Topfi 5 days ago

Not to be confused itself with K2 Think by MBZUAI...

Comment by sottol 5 days ago

Comment by reasonableklout 5 days ago

Interesting, have not heard of this company/org before. It seems they're from a UAE university?

Comment by TechSquidTV 5 days ago

I attempted their chat demo to see the speed and it stated the model couldnt be found.

edit: Tried signing up and using the internal playground. Holy shit thats fast.

Comment by throwawayffffas 5 days ago

While the open approach is commendable. The 32b and 36b models are inferior to qwen 3.8 27b, at least according to benchmark numbers. I would have liked to have seen both compared to 27b, not only the dense one. Also would have liked to see more coding benchmarks in the full table.

Comment by prometheus1992 5 days ago

Nice! can't wait to add these in my local stack and try them out.

Comment by villish 5 days ago

Frontier. Everything is frontier. K2 not to be confused with the other K2, or K3 that is also frontier.

Comment by KronisLV 4 days ago

No MTP?

Comment by 5 days ago

Comment by RynHong 3 days ago

[flagged]

Comment by luciana1u 5 days ago

[flagged]

Comment by adrian_b 5 days ago

I just looked on Huggingface.co, and the training data is there.

For example, 3.3 Tbyte for code reasoning, 4.5 Tbyte for mathematical reasoning, 8.4 Tbyte of pre-train behaviors, and so on.

I did not compute the sum of the dataset sizes, but it appears to be some tens of Tbyte. Nonetheless, I assume that this amount of training data is more than an order of magnitude less than what OpenAI, Anthropic and the like have used, which must have been at least many hundreds of Tbyte, but more likely several thousands of Tbyte of data.

Comment by luciana1u 5 days ago

[flagged]

Comment by lambda 4 days ago

Looks like we're still waiting on that, they have placeholder repos but haven't populated them yet:

* https://github.com/ifm-ai/xllm * https://github.com/ifm-ai/horizon-post-train

Their previous model, K2 Think V2, was release with fully open training data and recipe, so I would imagine that they are committed to that, but yeah, the repos for this new model are still just placeholders.

* https://mbzuai.ac.ae/news/k2-think-v2-a-fully-sovereign-reas... * https://github.com/LLM360/Reasoning360

Comment by luciana1u 4 days ago

[flagged]

Comment by lambda 4 days ago

Weights are up: https://huggingface.co/collections/IFM/k2-horizon

It's the training code that is not up yet, but this group has a history of publishing code so I would expect it, though of course you can never count on it until posted.

Comment by dakolli 5 days ago

Hey its a lot mpre thsn Anthropic which you probably use everyday all day without complaints.

Comment by luciana1u 5 days ago

[flagged]