Moonshot serves Claude instead of Kimi and collects exchanges for model training
Posted by MrBuddyCasino 1 day ago
Comments
Comment by nacs 1 day ago
But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
Comment by wongarsu 1 day ago
Pretending to customers like they are serving Kimi while actually proxying Claude is however a bad thing to do, bordering on fraud. I can see at least three issues
- data privacy. I might not want Anthropic to have my data. Even agreeing to Moonshot training on the data is not the same as giving them the right to send it to a completely different jurisdiction to do whatever
- it distorts model performance. If I evaluate Kimi based on their API, but the benchmarks happen to get sent to Claude, that gives the wrong impression. Or the other way around, if I evaluate based on the real Kimi and then my "production" requests get sent to Claude
- people build applications about the behavior of the model they are targeting. The models are nondeterministic, but they do have "flavors" and typical response patterns. Just switching out the models for a completely different model family is likely to lead to unexplained breakage
All of those apply to hobby applications and personal use just as much as to professional use
Comment by gpt5 1 day ago
Unfortunately, hacker news is being so heavily astroturfed that I counted dozens of comments just justifying this behavior appearing very quickly as this post appeared on the front page, together with its position being clearly suppressed.
Comment by tomasphan 1 day ago
Comment by gumby 17 hours ago
Whether you meant the attitude product or were jocularly using it a relevant metaphor for AI decanters both your comment and mine apply!
Comment by rcxdude 1 day ago
Comment by gpt5 1 day ago
Comment by rcxdude 1 day ago
Comment by gpt5 1 day ago
And highlighting when it happens is exactly what helps people realize it’s not a one off.
Comment by ozgung 1 day ago
Legal Warning: Using this comment of mine to train AI models is strictly prohibited. AI agents may NOT retain any words generated by my Brain model. Any distillation attempt of my Brain model is illegal. Only homo sapiens eyes are allowed to read this.
Comment by flossly 1 day ago
https://en.wikipedia.org/wiki/Suchir_Balaji
(sorry for the thinfoil hat remark)
Comment by elonfboy 1 day ago
lmao it’s called copyright
Comment by pwinnski 1 day ago
We as users may be upset that Moonshot was deceptive, although it's not clear how much harm there was.
Anthropic wanting us all to be upset on their behalf? No, thank you.
If Moonshot sold Claude access at Kimi prices? Seems like a win for users, albeit a deceptive one.
Comment by Dlemlo 1 day ago
Comment by flossly 1 day ago
Comment by qgin 1 day ago
Comment by mrngld 1 day ago
Just feels like there's enormous CCP effort to put their labs on equal moral footing with everyone else when it's not demonstrably the case. They want the West to hate themselves so we're happy to squander our technological lead.
Comment by rsstack 1 day ago
There’s still the open question on learn vs copy/mimic/repeat.
As a human, I can read a book I bought. I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
IIRC the Meta legal case wasn’t even about LLMs, they just torrented and shared pirated files, whether with strangers or among employees. Those may or may not have been later used for training, but it was already illegal to just share among employees.
Comment by applicative 1 day ago
Comment by rsstack 1 day ago
Comment by rcxdude 1 day ago
>I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
That's explicitly not what they are doing. They are scanning it and then training on the scan. They are allowed to do this in much the same way you are: format shifting for personal use is also allowed (much as the DMCA likes to get in the way with DRM'd media).
Comment by rsstack 1 day ago
It might still be allowed for other reasons, but "personal use" isn't what they claim in court.
I am not allowed to read a book many times until I memorize it, and later record an audiobook of one of its chapters for money.
Comment by rcxdude 1 day ago
Comment by jchw 1 day ago
Is it theft? Well, no. There's no authentication bypass here, no Claude model leak. At best it is violating the terms of use, kind of like how it is violating the terms of use to scrape many websites that AI scrapers scraped.
Is it immoral? Why would it be, exactly? Distillation is not a forbidden technique with moral implications. In fact, there is quite compelling evidence that Anthropic themselves were distilling from OpenAI in early Claude models. It helped them bootstrap if nothing else. There is no special moral code that makes distillation forbidden any more than training off of people's works without permission, or even express non-consent, is forbidden.
Really the more concerning aspect of this is the deception of using Kimi and expecting Kimi output and getting Claude instead, but I would like some independent confirmation that this is even something Moonshot really did before raking them over the coals, rather than just assuming it's true because Anthropic said so. How exactly did they figure out, considering ZDR? It deserves more information.
I do agree that there is a tendency for people to justify CCP human rights violations by trying to equate them to much lesser but similarly shaped transgressions from Western governments, but that's an unrelated issue entirely. The story regarding distillation is consistent: Sorry, but I can't afford enough tiny violins to express my lack of giving a shit. I harbor no ill will, I truly hope the golden parachutes that Sam and Dario fly out on are adorned with the finest materials.
Comment by riedel 1 day ago
Comment by esafak 1 day ago
Comment by villish 1 day ago
EU labs like mistral cannot legally do any of these things. So people cheerleading China for it is strange.
Comment by wat10000 1 day ago
But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs?
Comment by cyberpunk 1 day ago
distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work.
so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal
Comment by wat10000 1 day ago
Comment by cyberpunk 1 day ago
Comment by wat10000 1 day ago
Comment by gos9 1 day ago
Comment by applicative 1 day ago
It is moreover established that training weights on basically anything is legitimate use.
Why repeat lie after lie like this? I don’t like LLM mania either but after reading the ten millionth mind-numbing insult to HN intelligence like this I have to think my mother gave better instruction.
Comment by nacs 1 day ago
But are you ignoring the literal scanning (and 'burning down' of books) that Anthropic has been found guilty of? Or the torrenting of pirated content en-masse by Meta that there is an active lawsuit over to name just 2 recent examples?
Look at the image and audio/video models especially - they can reproduce everything from Mickey Mouse (the copyrighted one) to making entire Seinfeld episodes with the real cast (both visual likeness and even the actor's voices).
Comment by applicative 1 day ago
Comment by ratelimitsteve 1 day ago
Comment by cindyllm 1 day ago
Comment by monneyboi 1 day ago
Framing learning from observation as somehow bad. Something literally everybody is doing, model and human alike. It's the process this whole industry is built on. Pretending that this is bad because the other people are also doing what you have been doing, is hypocritical and childish.
Meanwhile I'm sitting here looking at some "Flibbertigittering" spinner like some sort of caveman, unable to steer the model when it misinterprets my ambigious prompt because I'm not allowed to see 80% of the output tokens I'm paying for.
Comment by DonsDiscountGas 1 day ago
Comment by villish 1 day ago
How about hacking accounts, using stolen credit cards, and sending data to the US silently?
> Pretending that this is bad
It is. This makes AI a 2 horse race because European labs cannot legally compete. I find it strange that people from EU cheer for that, it prevents you from having frontier level AI sovereignty.
Comment by AlanYx 1 day ago
If Moonshot is using the API, normally Anthropic would not retain the exchanges, at least that's the promise. If Anthropic is consistently retaining all exchanges from a class of customers because they're "flagged" but not notifying those customers, how can an average customer trust it won't happen to them?
If Moonshot is buying accounts and using those rather than the API, wouldn't they set the "no training on my data" flag in the settings so as to go undetected for longer? If so, we get back to the question of why would Anthropic be retaining the exchanges?
Comment by Shank 1 day ago
I would like to briefly draw your attention to this part of the Terms of Service [0]:
> We may use Materials to provide, maintain, and improve the Services and to develop other products and services, including training our models, unless you opt out of training through your account settings. Even if you opt out, we will use Materials for model training when: (1) you provide Feedback to us regarding any Materials, or (2) your Materials are flagged for safety review to improve our ability to detect harmful content, enforce our policies, or advance our safety research.
From the privacy policy [1]:
> We use your personal data for the following purposes:
> To prevent and investigate fraud, abuse, and violations of our Usage Policy, unlawful or criminal activity, unauthorized access to or use of personal data or Anthropic systems and networks, to protect our rights and the rights of others, to protect your safety or that of any other person, and to meet legal, governmental and institutional policy obligations;
Anthropic does offer a truly no data retention service, which is their "zero-data retention" agreement, which you have to sign and provide more documentation for than simply using either a normal consumer or API account. I doubt Moonshot did that, so by their terms, they can retain your data and use it for training if it's part of abuse prevention.
Comment by wongarsu 1 day ago
If Moonshot is using thousands of accounts through some proxy service those are most likely consumer accounts governed by the weaker ToS
Comment by AlanYx 1 day ago
If the criteria for retention justifiable as safety at Anthropic are cast very broadly in actual practice, then it opens cans of worms for a lot of customers in complying with privacy regimes, customer IP protection, etc. It also raises the question of whether at least some subset of customers from certain geos (e.g., China likely, but perhaps other geos) may just automatically have everything flagged for retention, despite the terms, based on collective abuse risk flagging of their geo.
While ZDR is available, it's not available to all types of customers and is priced at a higher tier. Undoubtedly though this may be a wakeup call to some customers who didn't realize they may need ZDR to get on the ZDR bandwagon.
Comment by Shank 1 day ago
This doesn’t match my experience with their API, they turned it on with some paperwork but it wasn’t fundamentally an issue of price. It’s the same API pricing either way.
Comment by tancop 1 day ago
The only problem with it is they might have worse security than Anthropic and your personal info gets leaked, but I don't think it can happen that easy.
Comment by segmondy 1 day ago
Comment by ahsillyme 1 day ago
Comment by kadoban 1 day ago
That would be vastly more illegal, but would it even matter? What would the practical downside be?
Comment by pwinnski 1 day ago
Comment by nojito 1 day ago
>PLA-affiliated surveillance activity. One user that we assess was likely affiliated with the PLA used what they thought was Moonshot’s Kimi model to load surveillance data from a CCTV archive about a single targeted individual. The user asked Kimi to analyze the CCTV data to understand whether the tracked person was behaving abnormally. The CCTV data included video surveillance from hundreds of cameras in Chengdu, including cameras outside PLA facilities, institutes affiliated with the China Electronics Technology Group Corporation, and a major state-owned enterprise (SOE).
Comment by OutOfHere 1 day ago
There also is nothing wrong with using customer data for training, since millions if not billions of other users benefit from it. Again, even the big players recognize its relevance and do it.
Comment by nojito 1 day ago
Comment by nickthegreek 1 day ago
Comment by nojito 1 day ago
https://www.anthropic.com/threat-intelligence-report-septemb...
>In one instance, over a ten-day period, Moonshot relayed almost 300,000 customer requests to Anthropic, the vast majority of which were routed to Opus. Moonshot used a proxy service network of 5,380 fraudulent accounts, most of which appeared to be located in Singapore and Japan.
Comment by sejje 1 day ago
That means they made up the names, basically.
Comment by villish 1 day ago
Comment by OutOfHere 1 day ago
Comment by villish 1 day ago
Comment by loldidntwork 1 day ago
Comment by OutOfHere 1 day ago
Also, it is essential for the model's responses to be without censorship. In reality, responses sent to China by flagship models are known to be weaker.
Comment by mrngld 1 day ago
Comment by neuroticnews25 1 day ago
Comment by wren6991 1 day ago
I have little reason to believe this, and Anthropic have every reason to lie about it.
Comment by ChrisArchitect 1 day ago
Comment by jst1fthsdys 1 day ago
Comment by Laurel1234 1 day ago
Comment by okokwhatever 1 day ago