← All episodes

[Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model

semianalysis · Jul 18, 2026 · 31:43

AI transcriptdiarizedcorrected5,459 words48 nuggetssource ↗audio ↗
Jordan NanosSpeaker B
Jordan Nanos0:00

All right, Max, we're going to do a podcast. We're going to talk everything about KIMI K3 and maybe some other models that just came out. How you doing?

Speaker B0:07

Doing great. Uh, looking forward to it. And thanks for having me, Jordan.

Jordan Nanos0:12

I'm not having you. It's us. Thanks for having me. Okay. On the docket, uh, let's say KIMI K3 hot takes. Is this the third best model in the world? Impact on OpenAI Anthropic architecture changes, personal usage that we've had so far. What we think about their open source strategy. And maybe more. All right, Max, quick hot take. Is this the third best model in the world right now?

Speaker B0:37

I think the answer is a clear yes. People, people love shitting on benchmarks. I think benchmarks definitely have their problems, but I think sort of if you take a composite of all the main benchmarks and just look at the model rankings, they have been directionally correct over time. And I think if you look at that composite today, there's like a pretty clear top 3 with Fable, uh, SOL 5.6, and now Kimmy K3. And there's sort of just like always above everyone else, uh, which includes of course other open source guys like, you know, Deepseek and Jilin, whoever. But also like very notably it includes Google and Meta and SpaceX. Um, which I think is honestly, it's an extremely impressive and a very remarkable feat from the Moonshot guys. Uh, Google in particular, I think, should feel incredibly embarrassed right now that at one point, guys, remember, as, as recent as like, you know, November, December 2025, everyone thought that like the clear AI Big Three was Google, Anthropic, and OpenAI. And even when I talk to like, you know, boomers today, they still seem to think their top three is Google, Anthropic, and OpenAI. And it's just like, clearly not the case anymore. Um, so yeah, I'd say definitely third best model in the world. I do think it's overall still worse than Fable and Soll 5.6. Uh, kind of funny that they explicitly said that in their like model release blog post. Um, maybe it's some like good old-fashioned Chinese humility. Maybe it's like they don't want to, you know, incur scrutiny from the US government or anything, but, uh, because obviously there were some, uh, like delays with the Fable 156 release. Um, but overall very impressed with the model.

Jordan Nanos2:26

Yeah, they, in the limitation sections of the blog post, they said despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Fable-5 and GPT-5.6.

Speaker B2:38

Cool.

Jordan Nanos2:39

So my experience using this personally is that it is good. It's really slow, which is really annoying. It's motivated me to try, um, open source harnesses for the first time. And so I feel like I'm learning more about the harnesses than I am about the models because Frankly, all these models are like good enough to do the basic work that I've been doing so far. I can't really find a lot of complicated stuff that it can't do, which in and of itself is a bit of a feat. Here's my hot take. For me, this might be the second best model in the world right now because every time I try and do something meaningful with Fable, I get rejected and I get sent down to Opus. Even though I don't know if this is better than Opus, it is less annoying to not get rejected whenever I'm trying to do something. However, I am getting, I'm not getting rejected when I use my API key and paper token, but I am hitting limits whenever I try and, you know, just use the web console or deep research or like the coding plan. I haven't used the coding plan, but some other guys at Semianalysis have. And so it leads me to be like, what, what is the strategy here? Because these guys clearly just do not have enough GPUs to serve the demand that they've got for this model. And previously that was solved by an open source strategy where they just drop the weights and then other people serve it and serve that demand, but they haven't dropped the weights yet. So I think they said weights in 10 days or something. What do you think the strategy is for the delay between announcement of the model, the API being available, and no weights yet?

Speaker B4:25

Yeah, I mean, I think to be clear, this is all just pure speculation on my part. But I think one big reason is like they need to give the vLLM and SuLAN guys enough time to make sure they can serve like this model performantly. Because if they just like dropped it today, you have all this hype, but then like everyone else sort of the model is only giving you like 20 tokens per second or something, that's probably really bad for their brand. Doesn't really sort of like capture— so they have this like incredible opportunity to get a bunch of huge PR and a bunch of adoption, and that was sort of like kneecapping a little bit. I think another possibility is that they are actively talking to the Togethers, Fireworks, Nebulas, like CoreWeaves of the world to figure out how they can sign some sort of licensing deal and have them serve like, you know, the incremental capacity on, you know, GV300s or whatever. Um, in my mind, those are sort of like the two main reasons why you would wait 10 days to actually draw the model weights.

Jordan Nanos5:31

Yeah, makes sense. And I think, um, functionally this is really interesting because just to talk about the model architecture for a second, it's 2.8 trillion parameters. Um, this does not fit in a B200, so you need to have B300, GV300, or I guess MI355X in order to be able to serve this model on a single system, like a single 8-way HGX server. Of course you can do unique things where you have pipeline parallelism across multiple nodes and stuff, but that's going to really impact performance. So, I think there's a lot of recipes being cooked up and only people who have the latest and greatest chips are going to be able to serve this model. Um, just to, let's, let's go back to your comment on, on Google for a second. The idea that this model is truly competitive in at the frontier at 2.8 trillion parameters, um, kind of gives us some insight into how big the closed source frontier models are, right? Like it would be even more embarrassing if they're hitting these levels of performance and they're being compared to 10 trillion parameter models. With a lot more active, right? It, we have to be, we have to assume that this is in the same range as what SOL and Fable are, right?

Speaker B6:48

Yeah, I think that's a great point. And, uh, you just like have to be correct. I think I still believe in sort of, uh, the competence and the correctness of all the OpenAI and Anthropic researchers. And if there are some people on Twitter who like to claim that like, oh, you know, current, uh, closed source is like 10 trillion total parameters or something. If that's actually true, like Guys, it's time to pack up the bags. Like, you know, SPY probably should crash like 50% like tomorrow. Like, you know, it's over. I'm pretty confident that like KIMI K3 cannot be much bigger— I'm sorry, much smaller. If anything, it might even be slightly bigger than the leading open source models today. Sorry, the leading closed source models today. And if that's true, this actually further highlights a point we've been harping on for a while on some analysis, which is that the margins for these like closed source labs have to be absolutely mind-boggling. Because if you're telling me, you know, Kemi is probably not at negative margins when they're serving K3 at $3.15, $3 per million input tokens, $15 million output tokens, that's sort of the same price as Sonnet. And so if you're telling me that like Fable is probably similarly sized and Anthropic can charge $10 per million input tokens and $50 per million output tokens, then this should just like immediately dispel any of the remaining concerns people have about the AI labs being these unprofitable businesses, like selling tokens at API prices is just, it might be even better than like SaaS, honestly, it's just, it's an incredible business today.

Jordan Nanos8:28

Yeah, makes sense. No, no cost of employees, just the GPUs. So can you compare this pricing strategy to the previous stuff? Because you said it's at $315, the previous version from Moonshot directly was at 95 cents and $4. So we're talking about more than 3x pricing increase from 2.7 Code to Kimi K3. Do they have more, even more room to increase pricing? Like what's the curve to get the frontier open source intelligence? Frontier soon-to-be open weight, yet, yet to be determined what the license will be, intelligence?

Speaker B9:14

Honestly, I don't think they have that much more room to push pricing up. Uh, like I would guess that even at like this $315,000, there will be a lot of people who are like, this is a little too expensive for me. Uh, my task is easy enough for like a GLM 5.2 or a Minimax M3, and I might just like use one of those models instead. And I actually think sort of, so on one end sort of you have like the semi-analysis of the world, right? Where we don't really care how much money we're costing Dylan when we like burn tokens all day. Like we're very happy, you know, using Fable for even like a relatively easy task that we're like pretty confident one of these open source models can do pretty well. And then on the other end, you have people who are like extremely cost conscious, like, you know, you maybe only get like $200 worth of tokens per week. Like as we heard some large companies like, you know, Tesla and Uber are implementing. And I think like pretty much all the people in that second bucket are going to be wanting to use like the GLM kind of pricing tier models because they are already good enough. For, uh, like most everyday tasks. And then everyone in the semi-analysis bucket is still using like FiveSixTool and Fable. Um, so I think there actually is a pretty interesting question of who is the user that is actually going to be switching towards ChemiK3. Um, I think it might just be a lot of people who like philosophically love open source and are excited to like try this new hype model and support it. Um, But it wouldn't surprise me at all if there's like not actual serious adoption amongst, say, like large enterprises of this model.

Jordan Nanos11:10

Okay. What do you think about where we go from here? Like, this is obviously a new base model. It's a fully new architecture for these guys. 2.8 trillion parameters. They've got— give me Delta Attention, attention residuals, the Stable Latent MoE that they keep using, the um, it's a, it's like a scaled up bigger version of the previous models, clearly about, about 2 times bigger. Um, but previously with K2.5, we saw Cursor Train Composer based on just continued pre-training as well as some RL. And then we saw KIMI give us 2.5, 2.6, 2.7 checkpoints as they just continued the RL., this is a new base model. It seems pretty complete. Like in my usage, it's working pretty well. It's not screwing up anything basic when it comes to writing a PR description or like totally going off the rails the way that we've actually seen some other models that are kind of raw without a bunch of RL have some rough edges at the beginning. I'm not seeing those yet. So where do we go from here? When does 3.1 come out? How does pricing change over time? Like does Composer, do we get a Composer based on KIMI K3?

Speaker B12:22

Uh, I mean, Composer based on coming K3, definitely not, because I think the Cursor guys are pretty set on training their own model from scratch now. As for when, like, you know, K3.1, K3.2, whatever come out, I imagine you probably see like 2 or 3 updates within like the next, each like a month or 2 apart or something, as they just continue post-training this thing. I would guess sort of like pricing stays about the same. Just because I don't, like, they're not going to be able to run it on new hardware in the next, like, 2 to 3 months. So they're not going to get, like, a huge, you know, throughput increase there to reduce pricing. Maybe it's possible some, like, really crafty kernel engineers figure out how to reduce the cost to serve this thing. So it's closer to, like, like DeepSeq v4 pricing or something. That would be really impressive. But given that it's a 3 trillion parameter model, I think I'm a little skeptical. I would guess that like the current pricing we see for like the Minimaxes and the GLMs is kind of already pushing the limits of like what you can serve like a 1T to 1.5T model at and not have like just embarrassingly bad margins. So I feel like this pricing is probably here to stay at least for the next few months. I think the most interesting question is whether or not the open versus closed gap is going to continue shrinking and if open will ever fully match closed source, like true frontier-level parity. Curious what your thoughts are there, Jordan. I think it has some serious implications for our whole industry if it actually happened, obviously.

Jordan Nanos14:09

Yeah. I mean, my view is that I believe the reason that this gap has closed right now is squarely put on the US government imposing restrictions on Anthropic and resulting in us not getting the actual best models that these guys have. And so they've artificially caught up basically.

Speaker B14:27

Interesting.

Jordan Nanos14:29

Clearly we see this with, with Mythos versus Fable. I can't use Mythos. I can only use Fable sometimes if I ask it nicely. 5.6 SOL. I think our house view is that it's not the biggest model OpenAI has ever trained. It's not the size of 4.5. To me, they have a bigger model somewhere. And I think the result is that we're going to be able to only access frontier intelligence if government entities allow us to. And that's a very interesting change to the setup going forward. Because I think it represents an opportunity for many of the players that are in 4th, 5th, 6th, 7th place to catch up to a limit. At which point it's okay to release everything and start to battle for user share without really being able to find the frontiers and have the frontier dominate. I think it's possible that, you know, we see the frontier take another big step towards the end of the summer. It's possible the politics change a little bit. Yeah. It's possible that we start to find other, you know, modalities beyond coding, uh, at which these guys can really improve and they start exploring those like areas. Um, we didn't, you know, intend to talk about this right away, but I just loved the, uh, release of Inkling by Thinking Machines. Um, thought that the native audio input would be super interesting, super useful in the future and kind of a sign of what's to come. But anyway, Yeah, I—

Speaker B16:09

And there's, on the topic of Inkling, there's definitely like huge, huge, huge demand for a Western open source model that does not suck. Like it is, like I'm shocked how markets are still so inefficient and like we haven't had a single American company that's at least on par with like the 5th best Chinese company. But if you just like, I mean, one, it's probably just a matter of time until the US government like bans Chinese open source models like entirely, uh, and maybe that's a can of worms we don't have to go down on this conversation. Uh, but two, even if that like doesn't happen, uh, I feel like the average large American enterprise is simply unwilling to like put all of their proprietary data through a Chinese open source model, even though you can make sort of like all the logical arguments of like, Dude, you know, you're like, you're just like loading their weights in like your air-gapped data center or whatever. There's like no way the CCP is actually going to like see any of your data. But like, I don't think the executives will actually buy that. And I don't think they really care. And I think there are a lot of people who like, A, care about token budgeting and B, are only interested in running a Western model or like a non-Chinese model. And so it is like shocking to me. That I guess Inkling is the best one that we have now. Um, but it's shocking to me that we're not like actually closer to the open source frontier in America.

Jordan Nanos17:36

Yeah. I mean, it was NemoTron and then it was Inkling and I'm, I'm, yeah, it's really inspiring to see Inkling go for it. I think they have two business opportunities there. They've got to be better than the bulk of Chinese open source. Like they have to be in the game there to be considered, but then they also need to be better than Sonnet. Or better than Terra, Luna, the tier 2, tier 3 models from the frontier labs, because, you know, you can build a bunch of cheap applications using close to frontier intelligence using Bedrock or Foundry or whatever, and just get access to the Anthropic or OpenAI models and save your money there by going with their second best model. So I never really understand the, western open source angle of saving people money. I think it is real and getting those models into the ecosystem of companies like Fireworks and Together and Base10 and anybody who's serving open source is a good thing because it is a market. But to me, the bulk of the market is government. Like one of the views on the Chinese models actually that's interesting is that Xi has been encouraging the Chinese companies to keep the models open source. That is the view from their party. And, uh, I think the, the big reason for that is that, uh, a bunch of the Chinese government wants to download the weights and run it on servers that they own and they want the support of the local ecosystem. And I think that the American government should work the exact same way. I think that's a pretty pragmatic strategy. It's like you need to give the people in your country, um, access and support to run this stuff. Um, maybe the, the, the other thing worth commenting on is that in the K3 blog, they mentioned the post-training, uh, sorry, the, the quantization during the SFT stage, right? And in that they were commenting on natively using MXFP4 and MXFP8 weights and activations respectively for quote, broad hardware compatibility.

Speaker B19:55

What other hardware do you think Moonshot cares about, Jordan?

Jordan Nanos20:00

I got a list of 11 Chinese accelerators. You should subscribe to the SemiAnalysis Accelerator model and learn more. But yeah, Huawei Ascend, you've got Baidu, you've got Kunlunxin, uh, you've got the Moore Threads guys. There's all sorts of different chips that are being, uh, they're showing up in papers. We're seeing code. Like it is a national priority for China to get these frontier models. These are frontier models now.

Speaker B20:29

Yeah.

Jordan Nanos20:30

Um, running on their domestic accelerators.

Speaker B20:34

Yeah, I mean, if, uh, I guess if we were calling Google a frontier lab, you know, end of 2025, we gotta call, they call Moonshot a frontier lab now. Kind of crazy, dude.

Jordan Nanos20:45

It's like Vantage. Yeah, no, my, you know, you're gonna, I'm still, I'm still a 34 waste. Yeah, yeah. And so are the 7 other Chinese labs.

Speaker B20:54

Yeah. No, no, funny enough, my, my dad is actually visiting China right now and he's telling me how like the hotel he's currently staying at. Is totally booked because Xi Jinping is going to be in the area soon and he's going to give a speech about how AI is a top priority for China. So I think a lot of what you said is right. Circling back to what you said earlier about the US government and how if they keep knee-tapping the frontier models OpenAI and Anthropic have, forcing them to delay them, forcing them to only have their second-best model actually publicly available and therefore giving all the other players, the Googles, the SpaceXes, Metas, whoever, time to catch up. Do you think that just completely destroys like the Frontier Lab business model? Like if you're open-eyed Anthropic, you just lose all pricing power at that, at that point, right? Like I don't see how Anthropic can still accelerate net new ARR if their model is on par or comparable with like the Meta model, the SpaceX model, the Google model, the Moonshot model, the DeepSeek model, like what happens to our industry at that point, Jordan?

Jordan Nanos22:09

Yeah. I mean, first of all, no, I don't think that's going to happen. And I think I can explain why, but second of all, I don't, I don't know for sure. So we'll have to see it play out. Interesting to think about. So I think the biggest thing that I've realized in my personal usage of this stuff is one, how hard it's getting to differentiate between using the Absolute frontier model and the max thinking mode versus high versus medium effort on those models. It's really, really hard for me to find day-to-day tasks that these models can't figure out. And I'm just, my behavior is default to the biggest and hardest thinking because I don't care about Dylan's budget, but Yeah. When it comes to actually, um, using this, there is an aspect of the harness being part of the product. So testing Kimi K3 requires me to take a serious look at OpenCode and Hermes and Py. And so the harness is totally part of the product still. Um, simple, simple things can cause me to want to use one model over the other. Like, can I install it on my remote SSH server? How easy are the keystrokes to get stuff in? Can I Edit previous commands. I mean, you know, it's, it's silly, but these, these little tiny features in the harness actually impact where I'm going to send my tokens, which results in where I'm going to send my budget. Right. Yeah.

Speaker B23:48

That's an interesting point because I think a lot of people, they talk about token budgeting, right? But I think, and correct me if I'm wrong, but what I'm hearing from your description of your own workflow, is that like, even for tasks where I'm pretty confident that like a GLM could successfully do it, like I'm happy routing it to Claude, doing that on max intelligence because like the ROI of that task is still worth like the Claude price to me. And so, and then like there's obviously there's like always going to be some risk in the back of your mind where it's like, oh, if I, you know, use GLM instead, even on like on medium thinking mode, it's way cheaper. Like maybe it's not actually like as high quality as Claude actually would have been. Right. And so, uh, even if the benchmarks claim that like, oh, a lot of your tasks can move to GLM, uh, you're still fine, like keeping them on Anthropic models or OpenAI models for the foreseeable future.

Jordan Nanos24:45

Um, mostly yes, but I use a lot of Slack bots right now and I actually don't know what model's running behind the scenes on that Slack bot. So specifically at Computer with Perplexity. And I think if it starts routing it to Kimi K3, if it starts routing it to GLM, starts routing it to Sonnet, and I know it's doing it today because I looked at my usage a few weeks ago and found how much of the OpenAI models I was using because it was making that decision. First cut at a PR before I go in and actually fix some stuff up. I don't really care which model they're using, right? And that is about the, that is about the quality of the harness there for what I'm, what I'm using. The justification—

Speaker B25:30

this might actually be a this might be a pitch for outcome-based pricing, if anything. Like one of these labs could potentially just like get 95%+ margins if they do outcome-based pricing for you because, yeah, as you said, all these tasks you're like happy to pay even stable pricing for when they could probably like get it done at a fraction of the price even today.

Jordan Nanos25:50

Yeah. Yeah, 100%. Certainly with the dialing in the thinking mode, which is where a ton of the expense ends up going. Yeah, I can totally, totally imagine them building a router. The second thing though, just on the competition thing you said earlier, is like, I don't think that we're out of use cases or ideas for these guys to work on. I don't think, I think they can continue to train incredible models to try and hit RSI on the coding side without ever releasing it to us in the proletariat and keep their—

Speaker B26:22

Is there a permanent underclass?

Jordan Nanos26:24

Yeah, keep their bourgeoisie models training each other, um, and keep distilling them and, and, you know, whatever, giving us the, the little tastes of it, uh, while still pursuing a research objective that includes all sorts of other uses of AI. Like we're, we're really exploring coding right now, but we do some video generation stuff. We do a lot of audio to audio stuff. We do lots of like deep research that really doesn't look like coding. In some ways. I think there's lots of, lots of use cases that they can continue to explore without really encountering like the cybersecurity issues. Robotics and world models is a simple one, right? What if Anthropic sets their sights on automating away a whole bunch of manual labor jobs instead of knowledge work jobs? I mean, the idea that there's no way for them to build a sustainable business with great ROI for their greatest technology the world's ever seen. I don't believe that at all.

Speaker B27:24

Yeah, it doesn't pass the smell test.

Jordan Nanos27:27

No, I, I, not at all. But even, even beyond that, the, I use these models so much every day and I see, first of all, how much my friends who work in technology, who are software engineers, spend 10 times less than me and use it 10 times less. Right now, right? One person using Claude is like, you know, a person using Sonnet and a person using Claude, they use it both the same amount, the same day. The person using Claude spends 10 times more at 90% margins. They make up the bulk of the business, right? As soon as those people, of which I would say there's maybe, maybe I'm in the top 10%, maybe even the 1% of the industry start using the bigger models, they use them more. That's just more demand for all of this business. And the models don't even need to get any better for them to discover they can use them for the really important tasks or the bigger ideas that they have. And then I need to go and talk to my neighbors who don't work in technology. And there I'm certainly in the 1%, probably in the 0.1%, maybe 0.01%. And maybe we've got 1,000 times more to go from here. And so I get back to, you know, Masa-san golden goose exponential chart to the right, you know, sort of slides at this point of just like, it's an exponential, hop on. Why are you— okay. Just to take this totally off the rails.

Speaker B29:03

But actually, before we go there, I want to say that I think what she said about being early is totally right. And this is like exactly why I don't think Kimi K3 is going to cause like net new ARR at Anthropic and OpenAI to decelerate. But I think even if, even if you want to like assume that some non-negligible portion of people who are using Claude and Claude Sonnet today are going to switch to like Kimi K3 because it's cheaper and can do their workloads. I think that is like completely overwhelmed by the people who still haven't like seriously tried this technology. You know, the people who've like kind of tried it a little bit, but are every day like discovering new use cases, like new cool things, high ROI things they can do with the models. I think all those people are going to be using like 5.6 SOL or Claude 5 as a default to like unlock these new use cases. And that is like, you're just not gonna see like a, a, a ARR slowdown, ARR growth slowdown, uh, because like that is gonna continue ramping so fast.

Jordan Nanos30:13

Yeah, I think we're in agreement on that. Think about how many people there are left to subscribe to this podcast and follow SemiAnalysis, man.

Speaker B30:20

Dude, it's crazy to me. Like, so I went to ICML last week, I went to AI Engineer the week before that, These are like, you know, nominally AI conferences, right? I thought people would be pretty plugged in there. Like, I would say 80%+ of people had never heard of SemiAnalysis before. I was like, guys, what are we doing, man? Like, you claim to work in AI, but you haven't read it? You've never even heard of SemiAnalysis? Like, we're still so early.

Jordan Nanos30:43

It's insane. That's an ego check, Max. Come on, man.

Speaker B30:46

Calm it down. Maybe tune our own horn a little bit.

Jordan Nanos30:56

Okay, man. I, I, I think we could keep talking about this all day, but probably good to wrap here. Anything you think was left upset, left unsaid? Any burning questions?

Speaker B31:07

I, I just, if the stock market crashes because all the investors, you know, have DeepSeek R1 movement part 2, buy the stocks, guys. Not investment advice though. Do your, do your own diligence. Not investment advice.

Jordan Nanos31:22

Love it. Let's end it on that. Clip it. Clip out Max saying anything about stocks and finish with me saying, do your own due diligence. Good job, man.

Speaker B31:34

Okay. Yeah. Cool.