Kimi K3 Now Available via Telnyx Inference API

Posted by fionaattelnyx 13 hours ago

Counter100Comment53OpenOriginal

Moonshot AI released open weights for Kimi K3 today and it's live on Telnyx Inference. The architecture and Moonshot's own benchmarks are in their technical blog. This post is about running it on Telnyx.

What we are adding: K3 is now available via the Telnyx Inference API, hosted on GPUs that we own and operate.

This matters due to the size of this model. A 2.8T model needs dedicated infra to serve well. Because we own and operate the GPUs, we can contorl throughput. with no inter-provider hops and no cloud tenant, latency is minimized. We have GPUs in each of the US, EU, APAC, and MENA. Inference runs in the region you pick, with zero data retention. We do not store prompts or completions after the response returns. Because we own the infra, the per-token price reflects the cost of running the model, not the cost of renting someone else's plus their margin.

We have not benchmarked K3 ourselves yet. Moonshot's own numbers and early third-party evaluations put it at frontier level for coding and agentic work, trailing only Claude Fable 5 and GPT 5.6 Sol on aggregate. Full breakdown in their blog.

Pricing on Telnyx: $2.70/1M input tokens, $13.50/1M output tokens, $0.27/1M cached input tokens. Prompt caching enabled by default. Served via an OpenAI-compatible endpoint so you can test easily.

Model ID: [MODEL_ID] API: https://api.telnyx.com/v2/ai/chat/completions Docs: https://developers.telnyx.com/docs/inference Technical blog (Moonshot): https://www.kimi.com/blog/kimi-k3

Comments

Comment by apexalpha 12 minutes ago

Is this an ad?

Why is this provider specifically on the front page?

Also why is this not available over openrouter?

Comment by dan_gee 5 minutes ago

[dead]

Comment by zius 4 hours ago

In my first interaction ("hi there kimi k3!"), Kimi K3 identified twice out of three times as Claude:

> Hi there! Quick note — I'm actually Claude, made by Anthropic, not Kimi. But no worries!

https://imgur.com/a/jqpc2Jc

and

> Just a quick heads-up — I'm Claude, made by Anthropic, not Kimi! Kimi is a different AI assistant (made by Moonshot AI), so it looks like there might be a little mix-up.

https://imgur.com/a/AKxeysH

Comment by tomashubelbauer 3 hours ago

FWIW Claude sometimes identifies as DeepSeek when asked in Chinese: https://x.com/stevibe/status/2026227392076018101

Comment by Razengan 3 minutes ago

The meme/trope of China copying everything really keeps playing into itself

Comment by victorbjorklund 2 hours ago

I remember when Claude identified as ChatGPT a long time ago. It proves nothing else than that there is a lot of training material on the internet with Claude as the AI.

Comment by nojs 3 hours ago

This happens with other models too - Gemini often identifies as ChatGPT for me, confusing many a debugging attempt

Comment by Razengan 22 minutes ago

It's distillations all the way down!

Comment by cleaning 3 hours ago

How many times do people need to point out that every model has this behavior until this stops being posted?

Comment by dan_gee 4 minutes ago

[dead]

Comment by croes 1 hour ago

Could be on purpose to disguise as a US made model

Comment by 23 minutes ago

Comment by FooBarWidget 3 hours ago

Models don't have an inherent identity. It should be obvious by now that every models trains on public AI chat session transcripts. I've seen Claude say it's Qwen.

Comment by TZubiri 3 hours ago

Bootleg AI

Comment by lukewarm707 55 minutes ago

when open source models are banned maybe i will become a cartel kingpin smuggling open source weights into the usa. find a sufficiently shifty street corner, 'what do you need', "kimi". are you familiar with my product? pure mxfp4, $500 per TB.

(joke and walter mitty, i know the government reads my messages)

Comment by victorbjorklund 2 hours ago

Comment by 2 hours ago

Comment by croes 1 hour ago

That’s how a smoking gone turns into a water pistol

Comment by Grimblewald 3 hours ago

if this qualifies bootleg, point me at non-bootleg frontier ai.

Comment by mesmertech 1 hour ago

Why are you guys not on Openrouter? I assume you'd get way more volume that way no?

Or does openrouter have like a specific contract you have to sign with them and requirements or smth? https://openrouter.ai/moonshotai/kimi-k3#providers

Comment by hmokiguess 7 minutes ago

Who is "you guys"? Telnyx?

Comment by hmokiguess 29 minutes ago

Is it also offered via a ZDR + BAA / HIPAA eligible? Would love to use it in prod

Comment by Mossy9 5 hours ago

Also available from Nebius, via Cortecs: https://cortecs.ai/detailedServerlessView/kimi-k3

€2.693/M input €13.464/M output Surprisingly, cache is not mentioned

Upd: Tensorix joined the fray, with the same prices, with cache at €0.673/M read

Comment by morpheuskafka 55 minutes ago

I have used Telnyx for, on the opposite end of cool-ness, their fax API. Curious if their AI pricing and quality are actually competitive or if this is just a rapid pivot to try to ride the AI wave.

Comment by theredsix 9 hours ago

10% cheaper than official! Let the inference pricing wars begin!

Comment by 4 hours ago

Comment by norbert515 2 hours ago

I'm really curious how far and how fast prices will drop (if at all)!

Comment by tokai 32 minutes ago

That is not going to be cheap for long if K3 prices does the same as GLM5.2 prices. If nothing else open weight models are great to get providers to compeete on price.

Comment by bedros 4 hours ago

any of these providers are HIPAA compliant?

Comment by vdfs 1 hour ago

Amazon Bedrock, but they don't have this model yet

Comment by maelito 4 hours ago

Where is it hosted ?

Comment by madhu_ghalame 4 hours ago

Consider publishing latency, throughput, and cost metrics under different workloads to help teams make informed decisions.

Comment by smallerize 13 hours ago

Very cool. What are your throughput and latency like?

Comment by rvz 4 hours ago

Jevon's paradox depends on how cheap the tokens get as the price of tokens get driven to zero as intelligence gets better and cheaper.

Comment by inigyou 2 hours ago

Well this is a very large and expensive model. Try something like Qwen-7B if you want ultra-cheap.

Comment by neuroticnews25 1 hour ago

Larger than Mythos? How is that?

Comment by jakswa 10 hours ago

what quantization? FP4?

Comment by NitpickLawyer 7 hours ago

The model is native mxfp4 w/ mxfp8 activations via QAT.

Comment by TZubiri 4 hours ago

I think telnyx is a good product, with the only stain to its name being the supply chain attack on their python library.

But I don't feel like providing inference is a professional move, it feels like out of scope for a telephony IaaS company. Feels like a FOMO moment where a reputation of years is crashed in a couple of weekends of being drawn into a fad.

And the fact that it's a chinese model doesn't quite help? I guess it's on brand with the 'cheap' pay as you go brand telnyx might already be associated to.

But more so it reads like Telnyx is trying to 'jump' into the trend of the 'open weights' discussion to compete with closed source incumbents. But we are at the tail end of the boom, anti ai sentiment is ever growing, customers now despise AI, especially in support channels, which is presumably the hook that Telnyx would have into 'AI'(LLMs). At this stage any company or individual that tries to join into the buzzword fueled cycle will pay the full fixed cost reputational price, but only reap the leftover hay from when the sun shone.

AI(LLM) on support channels is essentially a decapitalization of a company/brand, the company has a reputation that customers value, and might be worth good money in the market, and by implementing AI (LLMs) on support, a lot of costs can be cut, while the brand loses value, not sustainable. And by Telnyx (or any B2B company)asking their clients to participate in this decapitalization move, they essentially gamble their reputation as well, albeit with better odds as shovel sellers.

Comment by jbstack 3 hours ago

Large profitable corporations with very poor customer support have existed long before AI, so why should it be different? Using AI to provide poor customer service is just an implementation detail.

Comment by inigyou 2 hours ago

Companies changing their scope or opening side products is not at all unusual. Microsoft is Windows, what's this Xbox thing? Google is search, why do they have email? Y Combinator is a startup accelerator, why'd they make their own Reddit?

Comment by LoganDark 10 hours ago

Telnyx was cool until they started demanding KYC. I would use them for burner phone numbers until they started saying they needed my government ID. Fuck that.

OVH same thing. Tried to buy a VPS from them some years back and they said no VPS unless I provided ID. Would not refund me. Tried to dispute but my bank just gave me a credit instead.

Comment by illliillll 5 hours ago

You can go on telegram and pay someone $10 to do the KYC for you.

Comment by tomr75 4 hours ago

hi nsa

Comment by nujabe 7 hours ago

There were new regulations passed to combat robocalls that forced these companies to tighten up.

Comment by oxidant 3 hours ago

10DLC in North America

Comment by TZubiri 4 hours ago

Seems like a good riddance, telnyx is for business usecases, not for personal use.

Such use would be incompatible because it would lure in fraud, which would ruin the reputation of shared comms resources like ip blocks, phone number blocks, etc..

If you are doing real business, you are providing KYC as a daily matter in procurement, thereby protecting consumers from Sybil scum.

Comment by inigyou 2 hours ago

> thereby protecting customers from Sybil scum

No? Firstly you can't do a Sybil attack when each identity costs actual resources, and secondly you can still have lots of phone numbers, they just know who each one belongs to.

Comment by 13 hours ago

Comment by nttylock 5 hours ago

[flagged]

Comment by marsven_422 7 hours ago

[dead]

Comment by hiherer 9 hours ago

[dead]

Comment by buffer_overlord 13 hours ago

That’s huge

Comment by teravor 9 hours ago

    > Because we own the infra, the per-token price reflects the cost of running the model, not the cost of renting someone else's plus their margin.

is not compatible with

    > Pricing on Telnyx: $2.70/1M input tokens, $13.50/1M output tokens, $0.27/1M cached input tokens.
since you asserted something false and bizarre, how about telling us what is your markup?

Comment by alexeldeib 9 hours ago

Why are those incompatible? Pricing is 10% under moonshot, and pure infra providers also want money

Comment by gruez 9 hours ago

I don't get it, what's the contradiction supposed to be?

Comment by rubslopes 9 hours ago

They are selling it even cheaper than Moonshot AI. Why are you so sure they are lying?

Comment by teravor 8 hours ago

    We estimate that the true blended price per million tokens for running Opus 4.7 on agentic tasks at $0.99 despite the sticker price being $5/$25 per MTok.
https://newsletter.semianalysis.com/p/ai-value-capture-the-s...

according to Semianalysis those prices would be far above actual costs.

Comment by petu 4 hours ago

They're not saying anything about Anthropic serving costs in that quote, just calculating what MTok price is for running agents. Next sentence after your quote:

> Agentic workloads have extremely high input-to-output ratios (our Claude Code usage has a ratio of about 300:1) and high cache hit rates (90%+). Because cached input tokens only cost $0.50/MTok, most of the tokens end up in the cheapest tier.

90% cache hit input blend: 0.9 * $0.5 + 0.1 * $5 = $0.95 per MTok.

300:1 input/output blend: (300 * $0.95 + $25) / 301 = $1.03 per MTok.

They don't say exact cache hit rate they calculated for ("90%+"), so close enough.

Comment by Grimblewald 2 hours ago

Please correct me if I'm incorrect, but it seems to me those numbers are describing a fairly different situation to this one. Anthropic serving their own model to their own users at that scale gets cache hit rates and machine utilisation that someone standing up another company's 2.8T model in four regions isn't going to get, and this thing needs 64 accelerators minimum before it will run at all, so a rack sitting mostly idle through a quiet hour costs the same as a busy one. The margin figure quoted is also just the price against the cost of producing the tokens, it doesn't have buying the hardware in it, or depreciation, or maintenance, or the money they'd have made renting those machines out instead, which with rental prices up 40% since October isn't nothing.

I largely agree things are overpriced, I just don't think that article is the right basis for saying it about this one.

Comment by FergusArgyll 8 hours ago

**reflects** the cost...

Not the **literal** cost

Comment by tpm 4 hours ago

If they are indeed far above actual costs, then surely price discovery will be done by the overall market in due course.