Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

Posted by maxloh 1 day ago

Counter29Comment63OpenOriginal

Comments

Comment by throwa356262 1 day ago

This data sort of disqualifies itself: unless Moonshot has a time machine, K3 should be more similar to Opus 4.5-4.8 than Fable 5.

Keep in mind, Anthropic started limiting access and introduced anti-distillation measures around 4.5-4.6 (?). So the majority of distillation should have happened on earlier models.

Maybe a better explanation is that they have access to the same training datasets? Which if private can again raise questions about theft, but on a very different level.

Comment by 1 day ago

Comment by causal 1 day ago

So this shows distance relative to other models, but I don't have a good sense for what these numbers say in absolute terms.

K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant?

Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?

Comment by causal 1 day ago

I also suspect the questions asked matter a lot, and the system prompts matter a lot, because "the map is built from nothing but the words they choose" - so this is more a measure of linguistic style than anything.

If you use the Claude Code harness on two models you will probably get very similarly styled output. I would not be surprised if K3 "stole" a lot of the harness that was leaked.

Notice how the only rows that even come close to being gray is GPT-5.4+Mini, diverging even from other GPT models. Is this because it has a wholly different training set, or (more likely IMO), did it just have a system prompt that leads to a different style?

Comment by c0n5pir4cy 1 day ago

I think it needs some more explanation - Fable is 0.42 from itself apparently so Kimi K3 is basically indistinguishable?

Comment by kingstnap 1 day ago

Keep in mind if you ask Claude what model it is in Chinese it says its Deepseek or Qwen or Kimi.

So who's training on who's outputs?

https://news.ycombinator.com/item?id=48990086

https://x.com/stevibe/status/2026227392076018101

Comment by 1 day ago

Comment by orbital-decay 1 day ago

>character trigrams

Worthless. Any LLM output similarity metric that uses n-grams as a source might as well measure the average temperature on Mars surface, no matter how much lipstick you put on it. It just can't have enough certainty. There used to be an n-gram benchmark popular on Twitter that showed extreme similarity of grok-3-beta to gpt-4.5-preview, while these models were trained on new base ones, came out 2 weeks apart, and were unmistakably different. Results were wildly inconsistent run to run. It didn't stop the crowd believing its creator in that DeepSeek R1 was trained on o1-preview (which was obvious bullshit as well, they were as different as two models can be). It's amazing how you can put anything on the web and everybody will believe you without checking or even understanding of what they're looking at.

K3 was trained on Claude's outputs, though - it repeats Anthropic's prompt injections 1:1 in its reasoning, which you should know if you ever tinkered with both models long enough. Good for them.

Comment by 1 day ago

Comment by josh-paul 1 day ago

This looks to be more behavior based, not logit based? What is the actual claim?

Comment by weego 1 day ago

It seems to boil down to China bad, US good, but with tech people speculating so it's definitely valid

Comment by zobzu 1 day ago

honestly 99% of hn, reddit, etc. is the exact opposite of that statement.

Comment by codelion 1 day ago

How do you do cross entropy analysis for Claude without logits?

Comment by shiandow 1 day ago

I assume Kimi does provide logits. You can do cross entropy both ways. With slightly different results.

Edit: No that doesn't seem to be what's happening here. I think it's some kind of word frequency analysis.

Comment by 1 day ago

Comment by great_psy 1 day ago

Shouldn’t we expect a convergence as models become more powerful ?

There is an optimal answer to any question. Something that maximizes utility and minimizes tokens.

Comment by sebastianconcpt 1 day ago

Honest take: different surface, same category.

Comment by pandoro 1 day ago

All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"

Comment by skybrian 1 day ago

Does distilling actually violate any IP protection laws? Sure, it’s against their terms of service.

Comment by tills13 1 day ago

I get what you're saying and two wrongs do not make a right but the irony, and why people are even talking about this, is that the thing being distilled clearly, and knowingly, violated copyright & terms across the entire internet.

Comment by skybrian 1 day ago

Yeah, everyone is saying that but I don't find the irony very interesting.

Comment by Laurel1234 1 day ago

If it does then their original training did as well.

Comment by applicative 1 day ago

One mentions distillation to intimate that one is in fact still the leader. It has zero to do with fairness.

Comment by parineum 1 day ago

For me, I don't really care about the theft aspect but when people are claiming that these open models are better value or going to overtake anthropic/openai models, the implication that the open models are training of distilled data means all the "progress" they are making is just mimiced from the closed models.

It's a bit interesting how the open models are able to keep pace with the closed models except whole maintaining a steady following time.

Comment by abdullahkhalids 1 day ago

It's important to note that even if these open models are distilled, they are showing genuine improvements in their architecture, which enables inference costs to be several factors below what equivalent closed models have.

The interesting question is: will Anthropic release a Fable like model with an architecture similar to Kimi, and get the inference cost gains? They should surely beat Kimi because they can internally distill as much as they want.

Comment by orbital-decay 1 day ago

Distillation is absolutely not the reason they're good. It's not necessarily even done on a more capable model. It can even be done on itself and still bring improvement, or on a weaker model as well (see GLM and Gemini, which is definitely true because it repeats Deepmind's injections).

Comment by parineum 1 day ago

If distillation doesn't make them better, why do it?

If using Anthropic models for distillation doesn't make them better, why do it?

Comment by orbital-decay 1 day ago

It does, it's just not the reason. Most of the work is done before that point, and as I said z.ai used a weaker model. Besides, this is all strictly one-sided, as nobody knows how much Anthropic and OpenAI borrowed from Chinese labs' open research and weights (and they innovated a lot, to put it mildly, starting with first reasoning models worth talking about long before OAI did the same). Chinese labs are also severely restricted on hardware.

This entire story makes certain American AI shops look cartoonishly evil and Chinese ones relatively sane. Not only they want to grab without giving anything back, they also want to sabotage everyone else's AI research and do plenty of terrible things like media manipulation on the global scale and getting in bed with the government. This can't possibly end well, for the Americans in the first place.

Comment by nonethewiser 1 day ago

Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.

Comment by mosura 1 day ago

What of the costs for creating the data that was used to train the model being distilled?

Comment by yandie 1 day ago

At least 1.5B by stealing books, per a recent ruling.

I don’t feel sorry for the model companies

Comment by mosura 1 day ago

They only had to pay for storing the books on a server for later possible use. They did not have to pay anything for the training which was declared fair use.

Comment by 1 day ago

Comment by nonethewiser 1 day ago

This actually undermines the argument that distilling is harmless because its founded on the idea that Anthropic did the same thing and didn’t have any repercussions.

Comment by 1 day ago

Comment by applicative 1 day ago

Do you people seriously believe that Moonshot and the Chinese personal-cult-state dont possess and train on the same torrents?

Comment by mosura 1 day ago

They aren’t hypocritically crying foul about it, so no we don’t care if they do.

Comment by 1 day ago

Comment by nonethewiser 1 day ago

Well of course. But they dont hate Chinese model companies.

Comment by shlewis 1 day ago

Yes. I bet Moonshot paid for API access as opposed to pirating like Anthropic did.

Comment by nonethewiser 1 day ago

I understand you think Anthropic should have paid for the information it trained the models on. But im talking about all the costs to build a model. Do you think Anthropic didnt spend money to build these models? Did you not know that its actually very expensive?

Comment by shlewis 1 day ago

Moonshot pay for access at the price point Anthropic sees fit, potentially more, since they likely had to jump through multiple hoops?

It's hard to feel sorry about the breach of their ToS, which ultimately is all Anthropic can argue, when they are constantly being sued by countless IP owners

Comment by 1 day ago

Comment by pandoro 1 day ago

Could you give an example of the value that only training a model can create but none of the rights-holder could? I feel like if you got a direct, instant communication channel to any of the rights-holder that created the content in the training set of those models, you'd get more value than what the LLM could ever give you on any specific subject.

Comment by nonethewiser 1 day ago

>Could you give an example of the value that only training a model can create but none of the rights-holder could?

Yeah.

LLMs

Comment by 8note 1 day ago

the outputs of the model have no property protections, and training a model on the outputs of another model does the exact same thing - its expensive and creates new value over what was in the input - a set of documents.

Comment by jayd16 1 day ago

I'm going to hope this was sarcasm and if it is, it's great.

Comment by titzer 1 day ago

An incredibly ironic comment.

Comment by verdverm 1 day ago

There is a lot more one has to be good at to make a fable level model. Distillation won't get you there, it will help refine some at the end

Comment by rvba 1 day ago

A lot of those arguments could be said about writing a book, or a decent forum guide.

Comment by 8note 1 day ago

given that github is full of vibe coded repos, reddit is full of bot comments, and most blog posts are ai slop, isnt it guaranteed that anyone training on big public data sets will be closely tracking each other in text style?

Comment by Fokamul 1 day ago

I would gladly make a donation of Claude accounts to Kimi, to support more distillation.

Comment by throwa356262 1 day ago

On a more serious note:

If I am working on something not sensitive and afterwards upload my chat log to a public dataset, can open models use that data for training?

Keep in mind, I paid for access to model and half of the work is mine (the prompts). But can the model provider forbid me from publishing my result as open data?

Comment by josefritzishere 1 day ago

AI stealing from AI? That's shocking that an industry built on stealing would engage in stealing.

Comment by refulgentis 1 day ago

I didn’t participate in the discussion yesterday because I find it implausible Fable was available long enough (3-4 weeks cumulative?) to get data and train and have it fundamentally affect.

But I don’t grok training enough to know that’s silly.

If my new prior is you can…that’s a pretty thin moat that’s essentially indefensible.

Comment by credit_guy 1 day ago

My guess is they used the other Anthropic models extensively for synthetic data generation. The top most similar models for K3 are (lower means more similar)

Fable 5 -> 0.42

Opus 4.8 -> 0.45

Sonnet 5 -> 0.45

Opus 4.7 -> 0.46

Grok 4.3 -> 0.52

There's an obvious jump at Grok 4.3, and it would not surprise me that the similarity there is because Grok used Anthropic models for training too (you can get the similarity list for Grok and it does look like the top most similar models for Grok are either Anthropic models or some Chinese models).

The damning evidence that K3 used Anthropic models for training is that K3 is more similar to those models than it is to K2.6. If you look at the Anthropic, OpenAI or Google models, they are most similar with their own other models. Not so with K3, where K2.6 is less similar than 15 other models.

Now, why is Fable 5 the most similar to K3 and not Opus 4.8. I think it's quite likely that K3 did some fine tuning at the end, when Fable 5 became available. They probably had all the infrastructure in place, and Mythos had been announced for months, so they were probably waiting for the second the newest Anthropic model was released to start using it for synthetic data generation.

Comment by throwa356262 1 day ago

Note that until last generation all Chinese models were relatively small. This was mainly due to lack of training hardware, but as soon as the new Huawei NPUs started shipping some Chinese labs switched to larger models:

Deepseek: 670B to 1.6T

Moonshot: 1.1T to 2.8T

Xiaomi: 310B to 1T

The new models also use a different architecture so I would assume in tests they will look different from previous generations.

Z.ai (GLM) and MiniMax on the other hand have continued training existing models (with some surgical changes to improve long context memory) so they should score similarly in these tests.

Comment by isoprophlex 1 day ago

Stolen data was stolen. Oh no! Anyway.

Comment by nonethewiser 1 day ago

Can you elaborate?

This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet.

Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizing new technology including hardware, software, transformer architecture, etc.

Comment by wongarsu 1 day ago

I don't believe I understand your argument? Are you claiming a moral, legal or practical difference? Or are you saying that Anthropic spending resources on training an LLM is somehow different from an author spending resources on writing a book?

In any case, I doubt Kimi was trained without "stealing" the same data. Assembling all of your training data from Claude responses seems infeasible. It's much more likely that Kimi's base model was trained similarly to any other base model, with terabytes of data from all imaginable sources. Then the model was fine-tuned with "high-quality" data, followed by reinforcement learning. Throwing in lots of chat transcripts from other chatbots into the "high-quality" dataset would be expected, and is done to some degree by everyone, but maybe a lot more for Kimi. And likely they did a lot of reinforcement learning against the Claude API

The model would exist without Claude, it just wouldn't be nearly as coherent or smart

Comment by maxloh 1 day ago

Think of it in terms of distilled knowledge, not distilled LLMs.

I find both claims unsound, though. Knowledge or model behavior itself is not copyrightable, so all these claims just boil down to the "I am not happy with that" argument. You cannot claim someone is stealing something you don't own in the first place.

Comment by nonethewiser 1 day ago

You misunderstand the point. You can think both or either are morally right/wrong or good/bad for society. But distilling a model and building one using information online are fundamentally different. Even of you think the information the frontier models used was not fairly accessed.

Comment by maxloh 1 day ago

I don't see any US labs suing their Chinese counterparts. It is practically impossible. That makes it a verbal battle, not a legal one.

From a strictly moral standpoint, it is illogical to state, "You stole things from my archive of stolen goods." LLM vendors need to morally own the knowledge before their accusations hold. It is unsound to claim ownership of something resulting from stolen property, regardless of the work you put into building it.

Either those accusations don't hold at all, all training is just "fair use", or the two processes you described are just different forms of "stealing."

Bottom line, even if we consider knowledge copyrightable, it is just stolen goods changing hands. You cannot claim ownership of a derivative work if you deny ownership to the sources your work was derived from. No one holds a higher moral standing than another.

Comment by jayd16 1 day ago

You might as well say a photocopy of a book was "built using information and utilizing new technology including hardware, software, transformer architecture, etc."

Comment by nonethewiser 1 day ago

Yeah but thats to my point. You can’t legally resell a photocopy of a book.

Comment by jayd16 1 day ago

But you can use it to make an AI so then wouldn't distillation be fine?

Comment by ghm2199 1 day ago

All ML/AI models are comparable to some form of compression(al beit lossy) of information and in this case copyrighted information. The OP is pointing to this as stolen data(by all the companies that started with pre trained models)

Comment by nonethewiser 1 day ago

Models being compression algorithms doesn’t really make your point. Compression algorithms are copyrightable.

Comment by yreg 1 day ago

So what?

If training on copyrighted data without authors consent is ok then distilling is ok as well.

Comment by nonethewiser 1 day ago

Maybe both are ok. Maybe neither. Maybe one. The point is this does not follow

>If training on copyrighted data without authors consent is ok then distilling is ok as well.

Because building a model from distillation and building a model from raw data are not the same. You have to evaluate them independently. And legally its different as well. IP vs ToS (civil).

Comment by yreg 13 hours ago

No, it follows because the AI labs can at best argue that some copyright over the model output was broken.

They can hardly successfully sue a rival for breaking ToS.