> Just a quick heads-up — I'm Claude, made by Anthropic, not Kimi! Kimi is a different AI assistant (made by Moonshot AI), so it looks like there might be a little mix-up.
Models don't have an inherent identity. It should be obvious by now that every models trains on public AI chat session transcripts. I've seen Claude say it's Qwen.
I think telnyx is a good product, with the only stain to its name being the supply chain attack on their python library.
But I don't feel like providing inference is a professional move, it feels like out of scope for a telephony IaaS company. Feels like a FOMO moment where a reputation of years is crashed in a couple of weekends of being drawn into a fad.
And the fact that it's a chinese model doesn't quite help? I guess it's on brand with the 'cheap' pay as you go brand telnyx might already be associated to.
But more so it reads like Telnyx is trying to 'jump' into the trend of the 'open weights' discussion to compete with closed source incumbents. But we are at the tail end of the boom, anti ai sentiment is ever growing, customers now despise AI, especially in support channels, which is presumably the hook that Telnyx would have into 'AI'(LLMs). At this stage any company or individual that tries to join into the buzzword fueled cycle will pay the full fixed cost reputational price, but only reap the leftover hay from when the sun shone.
AI(LLM) on support channels is essentially a decapitalization of a company/brand, the company has a reputation that customers value, and might be worth good money in the market, and by implementing AI (LLMs) on support, a lot of costs can be cut, while the brand loses value, not sustainable. And by Telnyx (or any B2B company)asking their clients to participate in this decapitalization move, they essentially gamble their reputation as well, albeit with better odds as shovel sellers.
Large profitable corporations with very poor customer support have existed long before AI, so why should it be different? Using AI to provide poor customer service is just an implementation detail.
Telnyx was cool until they started demanding KYC. I would use them for burner phone numbers until they started saying they needed my government ID. Fuck that.
OVH same thing. Tried to buy a VPS from them some years back and they said no VPS unless I provided ID. Would not refund me. Tried to dispute but my bank just gave me a credit instead.
Seems like a good riddance, telnyx is for business usecases, not for personal use.
Such use would be incompatible because it would lure in fraud, which would ruin the reputation of shared comms resources like ip blocks, phone number blocks, etc..
If you are doing real business, you are providing KYC as a daily matter in procurement, thereby protecting consumers from Sybil scum.
We estimate that the true blended price per million tokens for running Opus 4.7 on agentic tasks at $0.99 despite the sticker price being $5/$25 per MTok.
They're not saying anything about Anthropic serving costs in that quote, just calculating what MTok price is for running agents. Next sentence after your quote:
> Agentic workloads have extremely high input-to-output ratios (our Claude Code usage has a ratio of about 300:1) and high cache hit rates (90%+). Because cached input tokens only cost $0.50/MTok, most of the tokens end up in the cheapest tier.
90% cache hit input blend: 0.9 * $0.5 + 0.1 * $5 = $0.95 per MTok.
In my first interaction ("hi there kimi k3!"), Kimi K3 identified twice out of three times as Claude:
> Hi there! Quick note — I'm actually Claude, made by Anthropic, not Kimi. But no worries!
https://imgur.com/a/jqpc2Jc
and
> Just a quick heads-up — I'm Claude, made by Anthropic, not Kimi! Kimi is a different AI assistant (made by Moonshot AI), so it looks like there might be a little mix-up.
https://imgur.com/a/AKxeysH
FWIW Claude sometimes identifies as DeepSeek when asked in Chinese: https://x.com/stevibe/status/2026227392076018101
This happens with other models too - Gemini often identifies as ChatGPT for me, confusing many a debugging attempt
How many times do people need to point out that every model has this behavior until this stops being posted?
Models don't have an inherent identity. It should be obvious by now that every models trains on public AI chat session transcripts. I've seen Claude say it's Qwen.
Bootleg AI
if this qualifies bootleg, point me at non-bootleg frontier ai.
Also available from Nebius, via Cortecs: https://cortecs.ai/detailedServerlessView/kimi-k3
€2.693/M input €13.464/M output Surprisingly, cache is not mentioned
Upd: Tensorix joined the fray, with the same prices, with cache at €0.673/M read
10% cheaper than official! Let the inference pricing wars begin!
Consider publishing latency, throughput, and cost metrics under different workloads to help teams make informed decisions.
any of these providers are HIPAA compliant?
Where is it hosted ?
Very cool. What are your throughput and latency like?
Jevon's paradox depends on how cheap the tokens get as the price of tokens get driven to zero as intelligence gets better and cheaper.
what quantization? FP4?
The model is native mxfp4 w/ mxfp8 activations via QAT.
I think telnyx is a good product, with the only stain to its name being the supply chain attack on their python library.
But I don't feel like providing inference is a professional move, it feels like out of scope for a telephony IaaS company. Feels like a FOMO moment where a reputation of years is crashed in a couple of weekends of being drawn into a fad.
And the fact that it's a chinese model doesn't quite help? I guess it's on brand with the 'cheap' pay as you go brand telnyx might already be associated to.
But more so it reads like Telnyx is trying to 'jump' into the trend of the 'open weights' discussion to compete with closed source incumbents. But we are at the tail end of the boom, anti ai sentiment is ever growing, customers now despise AI, especially in support channels, which is presumably the hook that Telnyx would have into 'AI'(LLMs). At this stage any company or individual that tries to join into the buzzword fueled cycle will pay the full fixed cost reputational price, but only reap the leftover hay from when the sun shone.
AI(LLM) on support channels is essentially a decapitalization of a company/brand, the company has a reputation that customers value, and might be worth good money in the market, and by implementing AI (LLMs) on support, a lot of costs can be cut, while the brand loses value, not sustainable. And by Telnyx (or any B2B company)asking their clients to participate in this decapitalization move, they essentially gamble their reputation as well, albeit with better odds as shovel sellers.
Large profitable corporations with very poor customer support have existed long before AI, so why should it be different? Using AI to provide poor customer service is just an implementation detail.
That’s huge
Telnyx was cool until they started demanding KYC. I would use them for burner phone numbers until they started saying they needed my government ID. Fuck that.
OVH same thing. Tried to buy a VPS from them some years back and they said no VPS unless I provided ID. Would not refund me. Tried to dispute but my bank just gave me a credit instead.
You can go on telegram and pay someone $10 to do the KYC for you.
hi nsa
There were new regulations passed to combat robocalls that forced these companies to tighten up.
10DLC in North America
Seems like a good riddance, telnyx is for business usecases, not for personal use.
Such use would be incompatible because it would lure in fraud, which would ruin the reputation of shared comms resources like ip blocks, phone number blocks, etc..
If you are doing real business, you are providing KYC as a daily matter in procurement, thereby protecting consumers from Sybil scum.
Why are those incompatible? Pricing is 10% under moonshot, and pure infra providers also want money
I don't get it, what's the contradiction supposed to be?
They are selling it even cheaper than Moonshot AI. Why are you so sure they are lying?
according to Semianalysis those prices would be far above actual costs.
They're not saying anything about Anthropic serving costs in that quote, just calculating what MTok price is for running agents. Next sentence after your quote:
> Agentic workloads have extremely high input-to-output ratios (our Claude Code usage has a ratio of about 300:1) and high cache hit rates (90%+). Because cached input tokens only cost $0.50/MTok, most of the tokens end up in the cheapest tier.
90% cache hit input blend: 0.9 * $0.5 + 0.1 * $5 = $0.95 per MTok.
300:1 input/output blend: (300 * $0.95 + $25) / 301 = $1.03 per MTok.
They don't say exact cache hit rate they calculated for ("90%+"), so close enough.
If they are indeed far above actual costs, then surely price discovery will be done by the overall market in due course.
**reflects** the cost...
Not the **literal** cost