← All episodes

Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos

semianalysis · Jul 1, 2026 · 34:56

AI transcriptdiarizedcorrected5,830 words35 nuggetssource ↗audio ↗

Synthesized from 35 insights · Jul 12, 2026

Huawei Ascend emerges as a credible NVIDIA alternative

Chinese hardware plus open-source software velocity is closing the gap NVIDIA long assumed was unbridgeable.

  • •DeepSeek reportedly delayed its V3/V4 release specifically to optimize the model for Huawei Ascend hardware first, and the Huawei team shared real benchmarks and kernel profiles at launch (Speaker C).↗↗
    quote
    “'Cause DeepSeek was like planned to release on Chinese New Year and then it got pushed back and back and back. And every time it was like, oh, it's gonna drop this weekend. Oh, nevermind, it's this weekend. Yeah, I said there were conflicting, like, uh, reasons. Some, some people on X were like talking about that's actually trying to get the model working on Huawei. That's why they were delaying it and like optimizing it first before releasing it. And of course there's the other explanation of, of like they just want that better evaluation results. But regardless, the performance on Huawei during release was real. There were benchmarks and like profiles shared by the Huawei team. And looking at profiles, they were very elegant. The Kernels optimization, as of course, as I previously mentioned, were quite sophisticated at that point. So they did get, they probably did get much longer access than vLLM or SGLang to optimize the architecture. And yeah, in the future we would release another follow-up article on this. Regarding Huawei's results and comparing it with other chips.”
  • •Both Zhipu AI's GLM models and DeepSeek V3 now run on Huawei Ascend chips, signaling Ascend is becoming a strong GPU alternative (Speaker D).↗
    quote
    “So first of all, it's pronounced Huawei. There's an H in there. I'm a bit too annoyed by this. Sorry. No, no, no, it's my bad. I have this, I have an issue. So, so yeah. Okay. Aside from that, yeah, I think ever since we see Zhipu AI with GLM are able to serve their models on Huawei Ascend chips. We also see like this time DeepSeek V3 can do that too. It shows that Huawei chips are getting there, being like a very strong option in addition to what we know from the GPUs and the TPUs. And Chinese have a very different ecosystem. They built a lot of things on their own, like Brian mentioned, the HiCoCo, MC², and everything. And they also, the engineers have been very aggressive in optimizing everything. I think these, the accumulation of all these things end up with a DeepSeek being able to offer a very low price. And I imagine that one of the reasons that they could probably keep the price is that they've also pushed a lot more optimizations since release. While I would also suspect that they are trying to capture the market. Based on my understanding is that the chat or the AI model market for China, ByteDance is still like the taking the majority. So everyone else is fighting for market share. So they are trying very hard at the cost of probably negative margin or at least, uh, just zero margin, trying to get people to use it. So that's why they could serve at a very low margin. And, um, uh, so that's, I think that's like the two main things, uh, two main implications.”
  • •Huawei's fully open-source CANN stack is hosted on Gitee with good documentation, frequent China meetups, and weekly public engineering calls; combined with '10x' Chinese developer velocity it positions Huawei as a long-term rival to NVIDIA's closed ecosystem (Speaker E).↗↗
    quote
    “Yep. And I saw on Gitee and it is very interesting to see how, how they do things like compared differently to NVIDIA or AMD, their documentation. In my opinion, has been very good. They do frequent meetups in China where you can talk to the— they call it CANN, C-A-N-N, and you can talk to the CANN developers during meetups. And they have like weekly calls, I think. They have a schedule of weekly calls where you can hop on and talk to engineers. So, like, I would take from this is just that Huawei is very enthusiastic about getting this working and open source contributions on CANN. So yeah, the direction has been very good. And some minor notes is that like Huawei has implemented some optimizations from like papers, I think slightly before NVIDIA. Like the one I talked in the article was about fusing communications and computation kernels. If I'm not wrong, NCCL released it in 2024, something like that. But Huawei released it much earlier, right after the paper on it was released actually. Which is quite actually quite interesting to me. I mean, the Chinese developers are indeed 10x developers and their velocity is actually great and will be compounded with, uh, uh, open source. Yeah. Thanks Jordan. The MC squared. Yeah. And their velocity will definitely be propelled by the Chinese open source. Yeah.”
  • •Huawei implemented fused communication-computation kernels shortly after the academic paper—possibly before NVIDIA's NCCL (released ~2024)—suggesting its software velocity may exceed NVIDIA's in certain areas (Speaker E).↗
    quote
    “Yep. And I saw on Gitee and it is very interesting to see how, how they do things like compared differently to NVIDIA or AMD, their documentation. In my opinion, has been very good. They do frequent meetups in China where you can talk to the— they call it CANN, C-A-N-N, and you can talk to the CANN developers during meetups. And they have like weekly calls, I think. They have a schedule of weekly calls where you can hop on and talk to engineers. So, like, I would take from this is just that Huawei is very enthusiastic about getting this working and open source contributions on CANN. So yeah, the direction has been very good. And some minor notes is that like Huawei has implemented some optimizations from like papers, I think slightly before NVIDIA. Like the one I talked in the article was about fusing communications and computation kernels. If I'm not wrong, NCCL released it in 2024, something like that. But Huawei released it much earlier, right after the paper on it was released actually. Which is quite actually quite interesting to me. I mean, the Chinese developers are indeed 10x developers and their velocity is actually great and will be compounded with, uh, uh, open source. Yeah. Thanks Jordan. The MC squared. Yeah. And their velocity will definitely be propelled by the Chinese open source. Yeah.”
  • •Chinese engineers have built an independent software ecosystem (HiCoCo, MC², MindSpore) and aggressively optimized at every layer, cumulatively enabling low-cost inference on Huawei hardware (Speaker D).↗
    quote
    “So first of all, it's pronounced Huawei. There's an H in there. I'm a bit too annoyed by this. Sorry. No, no, no, it's my bad. I have this, I have an issue. So, so yeah. Okay. Aside from that, yeah, I think ever since we see Zhipu AI with GLM are able to serve their models on Huawei Ascend chips. We also see like this time DeepSeek V3 can do that too. It shows that Huawei chips are getting there, being like a very strong option in addition to what we know from the GPUs and the TPUs. And Chinese have a very different ecosystem. They built a lot of things on their own, like Brian mentioned, the HiCoCo, MC², and everything. And they also, the engineers have been very aggressive in optimizing everything. I think these, the accumulation of all these things end up with a DeepSeek being able to offer a very low price. And I imagine that one of the reasons that they could probably keep the price is that they've also pushed a lot more optimizations since release. While I would also suspect that they are trying to capture the market. Based on my understanding is that the chat or the AI model market for China, ByteDance is still like the taking the majority. So everyone else is fighting for market share. So they are trying very hard at the cost of probably negative margin or at least, uh, just zero margin, trying to get people to use it. So that's why they could serve at a very low margin. And, um, uh, so that's, I think that's like the two main things, uh, two main implications.”
  • •Huawei publicly shared a detailed day-zero optimization video (in Chinese) on DeepSeek V4 kernel fusions and index optimizations, serving as a practical inference-tuning guide (Speaker C).↗
    quote
    “Yeah. So I guess they had a head start with implementation. That's why there wasn't as many hiccups. But I don't remember, I don't know if you guys remember seeing it, but, uh, there was a Twitter On Twitter, there's this video circulating of like Huawei talking about the optimizations, although it was in Chinese, but someone gave a translation and there were very good like optimizations talk about during Day Zero. So like if you just follow what Huawei did, I think you'll get a quite good performance if you were optimizing from Day Zero. Yeah. But they talked about kernel fusions, they talked about index optimizations. Yeah, it's a really good guide, I guess, for D0 optimization.”

DeepSeek V4's architectural leaps: MegaKernel and radical KV cache cuts

V4 is a genuinely new architecture whose engineering tricks redefine the efficiency frontier for long-context inference.

  • •V4 is a fundamentally different architecture—more total parameters but fewer active parameters than V3—not just a weight update (Speaker A).↗
    quote
    “And it's also a different architecture in terms of like more total parameters, less active parameters when compared to V3. So literally like a different model, not just an update to the weights. Brian, maybe kicking to you, what does this mean when it comes to actually getting support from the various inference runtimes and like what, what work is required to actually get performance results out of a system on day zero when the model's released?”
  • •The open-source MegaKernel (MegaMoE) delivers a 1.5x–1.73x performance improvement, validated on both NVIDIA GPUs and Huawei Ascend NPUs (Speaker A).↗
    quote
    “Okay. I wanna jump into throw, throw something on screen here. Just little screenshot from the v4 paper where it explains the open source MegaKernel. And it specifically says two things. One is that it's been validated on NVIDIA GPUs, nothing else, and Huawei Ascend NPUs. Obviously we've seen the open source CUDA-based MegaKernel called MegaMoE, which then these inference serving runtimes such as VLLM and SGLang and, uh, TensorRT can implement for themselves. And the performance improvement claim is somewhere between 1.5 and 1.73 times faster. So clearly they've, they've realized the benefits of this engineering work by, by doing this. Um, but I want to maybe share a second thing on screen, which is that like the open source community, just because the code is out there, doesn't necessarily have the opportunity to benefit just from day one. And so, um, what we've been tracking really, and Brian, maybe you can take us, take us through this like experience in, in detail, is how, how quickly these configs can come online. In other words, how, how quickly can a given inference serving runtime actually support things? And I thought this little GIF video from the article was quite nice in terms of how it shows the progress. Over time as we count up the days and we start to see more hardware being supported as the B200, B300, GB200, GB300 come online, the MI355 from AMD. And then you can start to see the performance improve, right? For people watching, like these lines going from the left to the right means either more throughput or lower latency per user or, you know, higher throughput per user in terms of interactivity. And, yeah, maybe you can take me through like, okay, day zero, there's 2 or 3 different hardware platforms supported, and then a month later everything's supported, but in that whole time, a lot of performance improvements happen. What, what are like some examples of some performance improvements that were just kind of like dropped? Overnight and then we're a big win where people can suddenly have, I don't know, 20% more throughput, which means 20% more users or 20% more profitability for the existing users or whatever, you know, for a given inference endpoint provider.”
  • •MegaKernel breaks typical kernel boundaries to eliminate register-to-HBM round trips and aggressively overlap compute and communication in a single kernel, cutting latency (Speaker D).↗
    quote
    “So, uh, fused kernel, I think fused kernel is one thing. The whole concept of MegaKernel is that you are breaking the kernel, the typical kernel boundaries. So compared to kernel fusion, which is either manually written by kernel engineers where they just go through the math and they see what can be, what can be merged together, uh, compared that and also like compiler optimization where they remove the kernel launches. MegaKernel is more about breaking the kernel, the kernel's boundaries that are not typical. For example, sorry, I suddenly can't remember which ones, but the, some concrete examples of what can be broken is that the register usage. So you can imagine that in, in a typical kernel, what happens is that you launch a kernel, you do computation. In a Von Neumann architecture, you would load data into the register, you do computations in a register, and then at some point you'll, you'll store it back to HBM. So what MegaKernel does is that they fuse different kernels so that they could remove the round trip between data movement between registers and the HBM, while also when a register is free, they could automatically do the next operations of the future kernels. So instead of waiting for a whole kernel to complete, and for MegaMoE, the case for MegaMoE is that in addition to that, they also merged the compute and communication, computation and communication overlapping. So it's like all merged into one kernel and then they could do like very aggressive overlapping.”
  • •V4 achieves roughly a 100x reduction in KV cache usage versus a standard MoE model, via sparse attention, embedding compressors, and sliding windows, enabling a 1 million token context length (Speaker D).↗↗
    quote
    “So, um, I think the headline changes or the features of V4 compared to V3 and R1 would be the 1 million context length. This ability in order to achieve or to get to a million context length, DeepSeek did like a very aggressive innovations on the attention mechanism. That's like the first thing. And the second thing is the MegaMoE to like speed up things, to speed up the expert FFN computations. First, the attention mechanism, DeepSeek V4 comes with two variants of sparse attention. One is the compressed sparse attention. And heavily compressed attention. So both of them are built on top of the previous version, DeepSeek sparse attention that came out with 3.2. The idea is that by sparsifying your attention between query and the key values, you could reduce the memory reads. So that allows you to do a 1 million context length without exploding your memory requirements. Both of them, heavily compressed attention and compressed sparse attention, in addition to that, have like one additional embedding compressor. So that further decreases the cache, the size of the KV cache entry. And then after that, there is, they both incorporate some sort of sliding window, which then further decreases the KV cache usage. Combined with all this, I think DeepSeek is quoting like a very, I think around 100, around that order, 100x reduction of the KV cache usage compared to a standard MoE model. And, and yeah, that's like the headline feature of DeepSeek V4.”
  • •MegaKernel's benefits are concentrated in low-latency inference; in large-batch training with hundreds of millions of tokens, compute-communication overlap is already sufficient and returns diminish (Speaker D).↗
    quote
    “Yeah. In the extreme case where we see like the Hazy Research, they literally fuse a whole model, like a Llama B, a Llama 8B. And I would say that the downside, or I guess the reason that is preventing people from doing that is that it's a lot of engineering work. And in in principle, in addition to it's a lot of engineering work, you have to justify it with something, right? So the thing is that MegaMoE is about aggressively reducing the latency, and in a large batch scenario or in a training scenario, it kind of doesn't make sense, or there is a limit to where it stops making sense. So let's say if you're doing a training with a super large batch, like a couple hundred million tokens per batch, your computation and communication is sufficiently overlapped. Your kernel launch time is not the bottleneck. In that case, it doesn't make sense to do MegaMoEs. And the other downside would be, I guess another difficulty for mega-kernels is that it's, because of doing like the aggressive resource allocation and on the fly, it creates a lot of memory pressure. So I imagine that would be, that'll require a lot of work on managing like the GPUs, I guess, so that they don't, I guess some physics related, maybe they'll overheat or some related issues, which would in turn affect performance.”
  • •MegaKernel creates significant on-the-fly memory pressure that risks thermal/performance issues and demands heavy engineering, limiting broad adoption despite theoretical gains (Speaker D).↗
    quote
    “Yeah. In the extreme case where we see like the Hazy Research, they literally fuse a whole model, like a Llama B, a Llama 8B. And I would say that the downside, or I guess the reason that is preventing people from doing that is that it's a lot of engineering work. And in in principle, in addition to it's a lot of engineering work, you have to justify it with something, right? So the thing is that MegaMoE is about aggressively reducing the latency, and in a large batch scenario or in a training scenario, it kind of doesn't make sense, or there is a limit to where it stops making sense. So let's say if you're doing a training with a super large batch, like a couple hundred million tokens per batch, your computation and communication is sufficiently overlapped. Your kernel launch time is not the bottleneck. In that case, it doesn't make sense to do MegaMoEs. And the other downside would be, I guess another difficulty for mega-kernels is that it's, because of doing like the aggressive resource allocation and on the fly, it creates a lot of memory pressure. So I imagine that would be, that'll require a lot of work on managing like the GPUs, I guess, so that they don't, I guess some physics related, maybe they'll overheat or some related issues, which would in turn affect performance.”

NDA access decided who won day zero

Early model access under NDA—not raw engineering talent—determined which stacks were ready at launch.

  • •vLLM and SGLang received early access to DeepSeek V4 under NDA, while NVIDIA did not, giving the open-source runtimes a head start on day-zero implementation (Speaker B).↗
    quote
    “I think it's worth mentioning, right, that some companies like VLLM and SGLang had like early access under NDA, but companies like NVIDIA like did not. So that made it more difficult to implement on day zero.”
  • •NVIDIA hit a day-zero hiccup because V4 used a new MHA dimension differing from all prior DeepSeek versions—something NVIDIA hadn't anticipated without early access (Speaker C).↗
    quote
    “Yeah, great question. So there were a lot of changes, not just like, as you say, not just the weights, but the sizes of certain things as well. For example, some hidden sizes and dimensions. So it's not just adding support for like newer attention mechanisms, attention blocks. It might also be like changing or tuning certain things that the community thought was constant. Like in this case, it was the— for NVIDIA, I guess it was the MHA dimension. DeepSeek for the previous times have always released MHA with What was it? I think half of what Pro used. So Flash, V4 Flash, and the previous iterations of DeepSeek V3 all use the same hidden dimension. Sorry, not hidden dimension, MHA dimension. Uh, but Pro used a new one, which caused a hiccup for NVIDIA on day zero. Yeah, but for the rest of the guys, I guess everyone adopted quite well, uh, on day zero.”
  • •AMD hardware supported only FP8 on day zero with no native FP4, a notable disadvantage since large gains came from getting V4 working in FP4 (Speaker C).↗
    quote
    “Wait, before I move on to that, huge shout out to our frontend engineer, Alec, for making that video feature. It is very cool and like shows you very nicely how improvements were made. Yeah. And so moving on to the question, uh, I guess one thing that we can like look at to look at these improvements is AMD. Like as the paper said earlier, a lot of the stuff optimized already for NVIDIA GPUs and like, because of how big the CUDA community is and how many people use CUDA, there's huge incentives for SGLang and VLLM in the limited time they had with the model to prioritize NVIDIA. So AMD catching up is actually an interesting look at like real optimization from getting to work on day zero when— so maybe a first step would be like getting the quantized versions to work. AMD on day zero had FP8 only. There were no native FP4 support. Which was quite unfortunate. And actually a very large gains came from getting it to work in FP4. Of course, due to the calculations being faster and lesser memory needs to be transferred. Actually, we have a very nice plot on InferenceX, if you may bring it up, Jordan. Thanks. On the AMD improvements. Yeah, but the biggest jump was converting a lot of the kernels, uh, into Either instead of using torch.fallbacks. Yeah. So you can see the SGLang graph actually evolving the most. There's a breakdown of optimizations as well, I think, below in the image. It's a pretty messy image with like orange text in, yeah, in Excalidraw. Sorry. Yeah. So there were very big improvements just from changing kernels from Torch fallback into either Triton or etc. So it's very cool to see like the impacts these individual improvements have. And of course, not just like a humongous step, it's always step by step. You optimize one part and you optimize another. Yeah, but very cool stuff.”
  • •Huawei's day-zero kernel optimizations were 'sophisticated and elegant,' implying far longer architecture access than vLLM or SGLang teams received (Speaker C).↗
    quote
    “'Cause DeepSeek was like planned to release on Chinese New Year and then it got pushed back and back and back. And every time it was like, oh, it's gonna drop this weekend. Oh, nevermind, it's this weekend. Yeah, I said there were conflicting, like, uh, reasons. Some, some people on X were like talking about that's actually trying to get the model working on Huawei. That's why they were delaying it and like optimizing it first before releasing it. And of course there's the other explanation of, of like they just want that better evaluation results. But regardless, the performance on Huawei during release was real. There were benchmarks and like profiles shared by the Huawei team. And looking at profiles, they were very elegant. The Kernels optimization, as of course, as I previously mentioned, were quite sophisticated at that point. So they did get, they probably did get much longer access than vLLM or SGLang to optimize the architecture. And yeah, in the future we would release another follow-up article on this. Regarding Huawei's results and comparing it with other chips.”

Open-source runtimes are eating the inference stack

SGLang and vLLM are becoming the default infrastructure even for the largest players, while vendor-locked libraries lose ground.

  • •Microsoft, Poolside, and GLM (whose entire inference stack is SGLang) all rely on open-source runtimes rather than building in-house (Speaker A).↗
    quote
    “They have like, I just want to say one other thing though, which is that I don't think it's completely a historical artifact the way Kimbo describes it. I think there's obviously stuff that's downstream of these runtimes now, whether it be customers or other libraries that depend on them. And so having two different providers allows you to make a decision about who's going to prioritize your features, who's going to merge your stuff, where you're going to get support from, how you're going to work with vendors, because they're just like one point and there's stuff that's upstream of them, like the vendor-specific libraries in some cases, or like the hardware, literally CI testing, InferenceX stuff., and then there's stuff downstream, which is like, you know, RL libraries, like you look at SLIME or you look at VRL or you look at PrimeRL and stuff. Like if there's features that they want to see from the runtime, they need to make a choice and go with the ones that are going to support them. And there's, there's many customers, like you mentioned, the inference endpoint serving providers, as well as the labs. You know, we know companies like Microsoft are using these technologies. They don't have their own stuff in-house. Uh, Poolside talked about that in their paper, GLM, the whole thing's actually SGLang, right? So a lot of people depend on this stuff right now, and I think that's like, it's at least good to have two options to, to buy from in open source. If you're somebody that's making a library and you need support from somebody and you can have multi, you know, runtime backend support if you really need to, if you're really not getting what you want from the community and you need to go somewhere else. The, the other, okay, not to, uh, not to take this in too far a direction, but I, I, I really, really want to hear about Huawei and this is making me think a lot about proprietary runtimes and software experience and like what this NPU chip is going to look like when there's a lot more users and it goes from what we thought was one or two lines in a paper saying, yeah, we have Huawei support and then no real proof. But Brian, man, there's real proof now. That DeepSeek is running on Huawei. Can you take us through what you've, uh, what you found in the analysis that you did on the numbers that they've shared?”
  • •Vendor-specific closed libraries like TensorRT-LLM and AMD's AITER are hyper-specific, less portable, and not always the best performers versus vLLM/SGLang (Speaker B).↗
    quote
    “The way I think about it is TRT-LLM and, you know, like AITER, the AMD optimized engine for ROCm is they're like hyper, hyper specific to certain AMD and NVIDIA architectures that will run really fast, but they're not necessarily like completely portable. And not always 100% open source, like, whereas vLLM and our SGLang, you know, you can fork it and go start like your own Baseten together, whatever, Fireworks, you know, it's very like user-friendly and it has, you know, API specifications and such. So that's kind of how I look at it, but I don't know, maybe Bryan and Kimbo have different opinions.”
  • •The coexistence of two competing runtimes isn't a historical accident—it gives downstream developers and enterprises optionality over who prioritizes their features and support (Speaker A).↗
    quote
    “They have like, I just want to say one other thing though, which is that I don't think it's completely a historical artifact the way Kimbo describes it. I think there's obviously stuff that's downstream of these runtimes now, whether it be customers or other libraries that depend on them. And so having two different providers allows you to make a decision about who's going to prioritize your features, who's going to merge your stuff, where you're going to get support from, how you're going to work with vendors, because they're just like one point and there's stuff that's upstream of them, like the vendor-specific libraries in some cases, or like the hardware, literally CI testing, InferenceX stuff., and then there's stuff downstream, which is like, you know, RL libraries, like you look at SLIME or you look at VRL or you look at PrimeRL and stuff. Like if there's features that they want to see from the runtime, they need to make a choice and go with the ones that are going to support them. And there's, there's many customers, like you mentioned, the inference endpoint serving providers, as well as the labs. You know, we know companies like Microsoft are using these technologies. They don't have their own stuff in-house. Uh, Poolside talked about that in their paper, GLM, the whole thing's actually SGLang, right? So a lot of people depend on this stuff right now, and I think that's like, it's at least good to have two options to, to buy from in open source. If you're somebody that's making a library and you need support from somebody and you can have multi, you know, runtime backend support if you really need to, if you're really not getting what you want from the community and you need to go somewhere else. The, the other, okay, not to, uh, not to take this in too far a direction, but I, I, I really, really want to hear about Huawei and this is making me think a lot about proprietary runtimes and software experience and like what this NPU chip is going to look like when there's a lot more users and it goes from what we thought was one or two lines in a paper saying, yeah, we have Huawei support and then no real proof. But Brian, man, there's real proof now. That DeepSeek is running on Huawei. Can you take us through what you've, uh, what you found in the analysis that you did on the numbers that they've shared?”
  • •Competition between vLLM and SGLang creates a competitive spirit that forces both to iterate faster than in isolation (Speaker B).↗
    quote
    “I mean, we tend to stray away from comparing vLLM to SGLang on InferenceX specifically, just because mainly because we didn't find it beneficial necessarily to have, like, make it a competition. We, we more so just wanted to showcase each one in isolation. But, you know, on the other hand, it does kind of create some competitive spirit, which makes both sides move faster, which can also be a problem because we only have limited compute to run things. So when you double the amount of, of submissions on, on a, on a certain model, it's kind of like, you know, a lot of CI time. But yeah, we've, we've seen, we've seen great improvements by both vLLM and SGLang, and I think a little bit of competition is good.”
  • •Performance gains compound incrementally—from PyTorch native fallbacks to progressively better custom kernels, each adding ~5% throughput, yielding large cumulative gains over weeks to months (Speaker B).↗
    quote
    “I want to like take a second to just talk about how cool this is and also shout out all the AMD engineers and NVIDIA engineers that we work with because This is all of them. Like we contribute a few things to upstream, like kernel libraries and stuff, but I mean, we're just a team of 3 people and this is all of the engineers. But I think this is something so cool about InferenceX and like this was the whole thesis of it, the project, like when we started it is like, you know, a lot of things online, a lot of benchmarks, you just see like the end product in terms of performance, but really, it's like a lot of hard work, tiny iterations, increasing throughput by 5% at a time. And that eventually compounds to sort of, you know, get performant. And like, you can see that here, right? So whenever new models release, like Minimax M3 just released, we're doing the same thing. And over the course of a month or two, you're going to see like, you know, you start out with PyTorch, you know, just Torch native fallback, and then you kind of use custom kernels and then you make the custom kernels even better and then you, you just keep making them better and better and pushing the frontier forward. So yeah, this is really cool to see.”

AMD's AITER is falling behind

Despite AMD's commitment, AITER's thin engineering and weak community support leave it structurally disadvantaged.

  • •SemiAnalysis benchmarks note AITER has no customers and roughly 20x fewer engineers than competing projects (Speaker C).↗
    quote
    “I mean, in my opinion, it's just more competition. I mean, on our benchmark, we also have a disclaimer saying that AITER has no customers. Yeah, but that's the other problem is your market just isn't, or rather isn't as big as NVIDIA's. Like unless your product is really, really much better. Like AITER, AITER is a lot better though than the previous fork. Yeah. But AITER, the development is okay. Development is fast. AMD is putting support into it, which you'd like to see. Uh, but yeah, SGLang is just. There's much more help from the open source community, uh, kernel development, which is really good. And AITER just needs that gap, I guess. And who knows, maybe it'll take some time for that gap to close and AITER gets more accepted by the community, I guess.”
  • •SGLang enjoys far greater open-source community help on kernel development, a gap AITER must close to become competitive (Speaker C).↗↗
    quote
    “I mean, in my opinion, it's just more competition. I mean, on our benchmark, we also have a disclaimer saying that AITER has no customers. Yeah, but that's the other problem is your market just isn't, or rather isn't as big as NVIDIA's. Like unless your product is really, really much better. Like AITER, AITER is a lot better though than the previous fork. Yeah. But AITER, the development is okay. Development is fast. AMD is putting support into it, which you'd like to see. Uh, but yeah, SGLang is just. There's much more help from the open source community, uh, kernel development, which is really good. And AITER just needs that gap, I guess. And who knows, maybe it'll take some time for that gap to close and AITER gets more accepted by the community, I guess.”
  • •Because the CUDA community is so large, SGLang and vLLM prioritize NVIDIA optimization first, leaving AMD to catch up over time after each model release (Speaker C).↗
    quote
    “Wait, before I move on to that, huge shout out to our frontend engineer, Alec, for making that video feature. It is very cool and like shows you very nicely how improvements were made. Yeah. And so moving on to the question, uh, I guess one thing that we can like look at to look at these improvements is AMD. Like as the paper said earlier, a lot of the stuff optimized already for NVIDIA GPUs and like, because of how big the CUDA community is and how many people use CUDA, there's huge incentives for SGLang and VLLM in the limited time they had with the model to prioritize NVIDIA. So AMD catching up is actually an interesting look at like real optimization from getting to work on day zero when— so maybe a first step would be like getting the quantized versions to work. AMD on day zero had FP8 only. There were no native FP4 support. Which was quite unfortunate. And actually a very large gains came from getting it to work in FP4. Of course, due to the calculations being faster and lesser memory needs to be transferred. Actually, we have a very nice plot on InferenceX, if you may bring it up, Jordan. Thanks. On the AMD improvements. Yeah, but the biggest jump was converting a lot of the kernels, uh, into Either instead of using torch.fallbacks. Yeah. So you can see the SGLang graph actually evolving the most. There's a breakdown of optimizations as well, I think, below in the image. It's a pretty messy image with like orange text in, yeah, in Excalidraw. Sorry. Yeah. So there were very big improvements just from changing kernels from Torch fallback into either Triton or etc. So it's very cool to see like the impacts these individual improvements have. And of course, not just like a humongous step, it's always step by step. You optimize one part and you optimize another. Yeah, but very cool stuff.”
  • •AMD is actively putting engineering support into AITER, a positive sign of commitment (Speaker C).↗
    quote
    “I mean, in my opinion, it's just more competition. I mean, on our benchmark, we also have a disclaimer saying that AITER has no customers. Yeah, but that's the other problem is your market just isn't, or rather isn't as big as NVIDIA's. Like unless your product is really, really much better. Like AITER, AITER is a lot better though than the previous fork. Yeah. But AITER, the development is okay. Development is fast. AMD is putting support into it, which you'd like to see. Uh, but yeah, SGLang is just. There's much more help from the open source community, uh, kernel development, which is really good. And AITER just needs that gap, I guess. And who knows, maybe it'll take some time for that gap to close and AITER gets more accepted by the community, I guess.”

The race to zero threatens closed-source moats

DeepSeek's willingness to run frontier open models on cheap hardware at near-zero margins may erode the pricing power of proprietary model providers.

  • •DeepSeek's strategy of running frontier open-source models on Huawei hardware and driving prices toward zero threatens the competitive moat of closed-source providers (Speaker A).↗
    quote
    “Fascinating stuff. Kimbo, let's bring you in here. When you, when you hear this stuff about support on Huawei and then you see something like DeepSeek V3, 75% off discount at the start, which then becomes permanent, what does that make you think for the future of like open source model competitiveness with closed source models if they're just going to try and drive the price down to zero and run it on Huawei hardware or whatever GPUs they can get their hands on? Yeah.”
  • •DeepSeek V3 launched with a 75% discount that became permanent, signaling aggressive pricing for open frontier models (Speaker A).↗
    quote
    “Fascinating stuff. Kimbo, let's bring you in here. When you, when you hear this stuff about support on Huawei and then you see something like DeepSeek V3, 75% off discount at the start, which then becomes permanent, what does that make you think for the future of like open source model competitiveness with closed source models if they're just going to try and drive the price down to zero and run it on Huawei hardware or whatever GPUs they can get their hands on? Yeah.”
  • •DeepSeek's ultra-low pricing is likely driven by zero or negative margins as it fights for market share in China, where ByteDance currently dominates (Speaker D).↗
    quote
    “So first of all, it's pronounced Huawei. There's an H in there. I'm a bit too annoyed by this. Sorry. No, no, no, it's my bad. I have this, I have an issue. So, so yeah. Okay. Aside from that, yeah, I think ever since we see Zhipu AI with GLM are able to serve their models on Huawei Ascend chips. We also see like this time DeepSeek V3 can do that too. It shows that Huawei chips are getting there, being like a very strong option in addition to what we know from the GPUs and the TPUs. And Chinese have a very different ecosystem. They built a lot of things on their own, like Brian mentioned, the HiCoCo, MC², and everything. And they also, the engineers have been very aggressive in optimizing everything. I think these, the accumulation of all these things end up with a DeepSeek being able to offer a very low price. And I imagine that one of the reasons that they could probably keep the price is that they've also pushed a lot more optimizations since release. While I would also suspect that they are trying to capture the market. Based on my understanding is that the chat or the AI model market for China, ByteDance is still like the taking the majority. So everyone else is fighting for market share. So they are trying very hard at the cost of probably negative margin or at least, uh, just zero margin, trying to get people to use it. So that's why they could serve at a very low margin. And, um, uh, so that's, I think that's like the two main things, uh, two main implications.”

What InferenceX is building next

The benchmark is expanding toward agentic, long-context, and multi-chip testing—where the next competitive surprises will surface.

  • •InferenceX is developing an agentic benchmark using real internally collected Claude code traces, and will showcase KV Block Manager, DynamoDB, and PD disaggregation optimizations in the coming weeks (Speaker B).↗
    quote
    “So one of the things with Minimax M3 and DeepSeek V3 are, is the 1 million context length. Right now, obviously on InferenceX, we are not testing that. We're testing fixed sequence length, 8K, 1K, 1K, 1K, which nobody uses. Nobody wants to. But I said this in the last podcast and it's still good to have the 8K, 1K, and the 1K, 1K, because it's testing, it's showcasing just in my opinion, basically pure chip performance, right? There's no prefix caching or anything like that. So we're moving it up and we're going to showcase, I mean, inference is a systems problem. So we're moving up the stack. We're going to showcase, we're going to run an Agentic benchmark with real Claude code traces that we've collected internally. We're going to showcase things like DynamoDB and KV Block Manager, other sort of KVBMs and, uh, yeah, showcase like different PD disagg optimizations. And I think it's gonna be really cool and that, that's gonna come in the next couple weeks, so be on the lookout for that. Yeah, that's, that's all I have to say about it right now, but we'll, it'll be really cool when it comes out.”
  • •Current benchmarks use fixed sequence lengths (8K, 1K), so 1M-token context testing for models like Minimax M3 and DeepSeek V3 is not yet covered but is planned (Speaker B).↗↗
    quote
    “So one of the things with Minimax M3 and DeepSeek V3 are, is the 1 million context length. Right now, obviously on InferenceX, we are not testing that. We're testing fixed sequence length, 8K, 1K, 1K, 1K, which nobody uses. Nobody wants to. But I said this in the last podcast and it's still good to have the 8K, 1K, and the 1K, 1K, because it's testing, it's showcasing just in my opinion, basically pure chip performance, right? There's no prefix caching or anything like that. So we're moving it up and we're going to showcase, I mean, inference is a systems problem. So we're moving up the stack. We're going to showcase, we're going to run an Agentic benchmark with real Claude code traces that we've collected internally. We're going to showcase things like DynamoDB and KV Block Manager, other sort of KVBMs and, uh, yeah, showcase like different PD disagg optimizations. And I think it's gonna be really cool and that, that's gonna come in the next couple weeks, so be on the lookout for that. Yeah, that's, that's all I have to say about it right now, but we'll, it'll be really cool when it comes out.”
  • •The team is running the same iterative optimization process on the newly released Minimax M3 that it applied to DeepSeek V4 (Speaker B).↗
    quote
    “I want to like take a second to just talk about how cool this is and also shout out all the AMD engineers and NVIDIA engineers that we work with because This is all of them. Like we contribute a few things to upstream, like kernel libraries and stuff, but I mean, we're just a team of 3 people and this is all of the engineers. But I think this is something so cool about InferenceX and like this was the whole thesis of it, the project, like when we started it is like, you know, a lot of things online, a lot of benchmarks, you just see like the end product in terms of performance, but really, it's like a lot of hard work, tiny iterations, increasing throughput by 5% at a time. And that eventually compounds to sort of, you know, get performant. And like, you can see that here, right? So whenever new models release, like Minimax M3 just released, we're doing the same thing. And over the course of a month or two, you're going to see like, you know, you start out with PyTorch, you know, just Torch native fallback, and then you kind of use custom kernels and then you make the custom kernels even better and then you, you just keep making them better and better and pushing the frontier forward. So yeah, this is really cool to see.”
  • •Several new chips are in the benchmark pipeline that could reveal new competitive dynamics in AI inference hardware (Speaker B).↗
    quote
    “That's what I, that's what we were just talking about. So that's coming in the next couple weeks, but I'd say be on the lookout. New chips. Yeah, be on the lookout for new chips. We have quite a few in the pipeline, so that should be really, really interesting.”

Stock read-through

DeepSeekBullishV4's open-source MegaKernel (MegaMoE) delivers 1.5x-1.73x speedups validated on NVIDIA and Huawei; release delayed to optimize for Huawei Ascend; aggressive low pricing to capture Chinese market share.↗
HuaweiBullishAscend chips increasingly viable as GPU alternative, running GLM and DeepSeek V3/V4; software velocity via open-source CANN stack may rival NVIDIA in some areas.↗
NVIDIAMentionedStill dominant with CUDA community driving runtime optimization priority and FP4 advantage, but did not get early V4 access and trails Huawei in some kernel optimizations.↗
AMDBearishLagged on day-zero V4 support (FP8 only, no native FP4); AITER runtime has no customers and ~20x fewer engineers, trailing SGLang in community kernel support.↗
SGLangBullishGot early V4 access under NDA, strong open-source community kernel support, used by GLM's full stack; competition with vLLM drives rapid innovation.↗
vLLMBullishReceived early V4 access under NDA; prioritizes NVIDIA optimization; competition with SGLang accelerates innovation.↗
MinimaxMentionedM3 model supports 1M token context; InferenceX beginning iterative optimization process for it.↗
Zhipu AIMentionedGLM models run on Huawei Ascend chips, signaling Huawei's growing viability.↗
MicrosoftMentionedUses open-source inference runtimes (SGLang/vLLM) rather than in-house solutions.↗
PoolsideMentionedUses open-source inference runtimes per its paper rather than building in-house.↗