← All episodes

Ep. 014 - Finding Miscompiles For Fun, Not Profit (AI Infrastructure) | Justin Lebar & Jordan Nanos

semianalysis · Jun 4, 2026 · 23:19

AI transcriptdiarizedcorrected4,863 words24 nuggetssource ↗audio ↗
Jordan NanosJustin Lebar
Jordan Nanos0:05

Hello everyone, welcome back to Semi Analysis Weekly. This week I'm with Justin LeBar. We're gonna talk about an article that he put up last week called Finding Miscompiles for Fun, Not Profit, also known as You Don't Need Access to Claude Mythos to Spend $10,000 in an Afternoon. Justin, welcome to the show. Hot off a $10,000 afternoon. How are you feeling? How you doing?

Justin Lebar0:26

I'm happy to be here.

Jordan Nanos0:29

Okay. At a high level, can you, uh, can you talk me through what you covered in the article? It wasn't one of our longest articles recently, but it's actually gotten the most, uh, most likes on Substack or something that we've put out in a while. So I think people really enjoyed this one.

Justin Lebar0:42

Happy to be winning some internet points here. What I went through in the article was the process of finding bugs in two compiler stacks. The first one being Nvidia's closed-source compiler for PTX. So PTX is this low-level assembly, or maybe I should call it high-level assembly language, that's emitted by various compilers, and that goes into, that's the input into PTX-AS, and the output from PTX-AS is the machine code that actually runs on the GPU. So the process of finding bugs in that, and also the process of finding bugs in LLVM, and I looked at two backends in LLVM, the AMD GPU backend, so that would be if you write like HIP code and it compiles through Clang, and also the x86 backend, and the shared code between them that's sort of used by all or most of the backends in LLVM. And I looked at kind of two different ways of finding bugs in the compilers. The first way is kind of a more traditional way called fuzzing, where essentially you generate random programs and run them through the compiler and then somehow check that the output after you compile the code is the same as before. And there's some subtleties to that that we can go into if we want, but essentially that's all you're doing. So I did some fuzzing, and then the other thing I did was LLM-assisted bug finding, where I just literally unleashed the agents on the code and said, please read the code and find bugs. And we talked about the difference between these two approaches and kind of Also speculated a little bit on what might be possible in the future.

Jordan Nanos2:16

Awesome. Could you talk a little bit about the motivation to do a project like this, like where it came from in your mind? And then we can dig into, you know, what exactly you did in more detail.

Justin Lebar2:25

I've been working on compilers for a decade, compilers really for GPUs and machine learning and find— and bugs are, bugs are terrible in compilers, but I think actually it's kind of worse in the ML world than it is in the CPU world. Um, and that's because if you think about like a regular C++ compiler, it's had hundreds of millions, maybe billions of lines of code run through it. And if there's a miscompile, like there's some chance that someone is going to notice, that chance goes up the more lines of code you run through the compiler. In contrast, like when we're compiling for GPUs, um, there's just so much less GPU code in the world. You think about like you know, an LLM, there's relatively few programs, relatively few GPU kernels that go into an LLM as compared to like all of the code that has to go into Google Search. So that's one big motivation for doing it is just that like my suspicion, which is somewhat borne out, somewhat not borne out, is that it would be easier to find bugs in the ML compilers just because they're like less mature or less well-tested in terms of the amount of code that gets run through them.

Jordan Nanos3:35

Can you talk about why that might be true, might not be true? Like maybe a little more details on what you actually went through with the fuzzing and then the, uh, the other approach with the model itself directly reading stuff?

Justin Lebar3:45

Yeah, I mean, it's really hard to say with these approaches when you get the results, like it's really hard to interpret them in terms of like the things, the headline that you want to write is like the compiler, X compiler is so bad, you know, it's like completely buggy piece of trash and it's like really hard to arrive at a conclusion like that. So, um, when I look at like the rate of bugs that I was able to find in AMD GPU versus x86, um, the LLMs reported a lot of bugs, but you then sort of have to ask like, what's the severity of these? How likely is it that someone's going to hit it? And that's really hand-wavy and hard to say. Um, my— and also I haven't honestly looked into the AMD GPU bugs as much as I've looked into the x86 bugs because I'm letting the AMD people work on the AMD stuff. Um, in terms of the x86 bugs that I found, I found one, maybe two, like very high severity bugs where if code hit it, which is like a little bit iffy, is, is code going to do it? But it's like reasonable that code might hit it. And if it did, it would be really bad. Um, the one that we link in the article is a bug where if you construct an atomic operation in a certain way, the compiler will split it into two non-atomic operations, which is like really bad. But part of what makes it so bad is that actually most of the time it'll work fine. If you don't have contention on this particular line in the CPU, you would never notice that you even had this bug. But, you know, 1% of the time it's gonna do the wrong thing. So like we did find that, but that was kind of the worst one that we found. Whereas like in AMD GPU, we actually did find a number of miscompiles, But again, I don't wanna like extrapolate from that to say that like the AMD GPU backend is actually buggier. It's just what we were able to find with the tools that we have.

Jordan Nanos5:34

Makes sense.

Justin Lebar5:35

Yeah.

Jordan Nanos5:35

And maybe can you talk a little bit about how you actually went about this in terms of writing the fuzzer with an LLM and then having the LLM read the outputs itself?

Justin Lebar5:45

So, you know, my first instinct was to do the thing that I've done before, which was to write a fuzzer. So in that sense, you're, you're writing a program that does two things. First of all, this program has to like generate well-formed inputs, well-formed random programs, because you can't just like throw any arbitrary program at the machine and expect it to like do the right thing. You can have undefined behavior in your program among other things. So you have to somehow like generate well-formed programs that also are weird enough that they might trigger a bug. Um, so that's one thing the program does. And then the second thing that the fuzzer program has to do is it has to have some way of checking equivalency. After compilation versus before compilation. And usually the way to do that is to run the programs somehow, but you could run them actually on a GPU, which is what I was doing. You could also just like, uh, interpret the programs by like looking at the output from the compiler and going line by line and writing an interpreter for that. Um, so that was the first thing I did for both Nvidia and for AMD, and I had, uh, the LLM write the fuzzer for me. And what, What was remarkable to me about this was just like how much better the fuzzing, that experience was as compared to just a couple months before. I know I'm not saying anything new, like we all know that the models have gotten better, but I think maybe a less well-told story is that not only have the models gotten better, but the harnesses have also gotten better. When I first did this, and that was in December to January 2025, 2026, a lot of the time that I spent was literally going into Codex and saying like, good job, keep going. You know, like kind of repeating myself over and over. It would like find one or two bugs and it'd be like, I'm done. And I'd be like, no, you're not. And now like the harnesses all have this /go feature and you just type that and it just works. In some sense, it's not like the most complicated thing in the world. I could have written that myself on top of Codex back then if I had really cared, but now it's built in and it's tested and supported. Um, so that was one thing, but also just like the quality of the code that the LLMs were generating was much higher. And so I was able to just kind of vibe code the whole thing and that hadn't worked when I had tried in January.

Jordan Nanos7:54

So what, what do you think the big unlock was for the LLM? Was it the fact that you were able to do this faster or you just like wouldn't have even wanted to pursue this because you didn't have the time or the motivation in the past, or it was like more enjoyable?

Justin Lebar8:06

I think all of the above, honestly, like these are all related things. Um, you know, whenever you're embarking on a project to like, that has an unknown payoff, I'm looking for bugs. I don't know what I'm going to find. Uh, you always wanna be a little bit skeptical of like the amount of time that you're going to put in. Um, the fact that it was like, and in fact, when I was negotiating this project with Dylan, he sort of said the same thing to me. He's like, yeah, do it if it's gonna take you a couple days, but if it's gonna take you a couple weeks, like, you know, we should really think about it. Um, and I was like, I don't know how long it's gonna take me, so I'll spend a couple days on it. And then a couple days I just had these long lists of bugs, right? Um, so that was the motivation for fuzzing. But then, like, inevitably with fuzzing, you hit kind of two problems. Problem number one is that, like, it becomes harder and harder to expand the universe of programs that your fuzzer will generate. Um, because again, you're sort of always hitting against these limits of, like, undefined behavior or other ill-formed programs. So that's one problem. Problem number two— maybe there are three problems. Problem number two is that as you do this, it also, like, you're generating random programs and you're doing, in some sense, an uninformed search. And so you would expect that, like, in the beginning you're going to find lots of bugs, but then as this universe of valid programs expands, it takes longer and longer to find things that are at the edges that actually trigger bugs. So that was another problem that I was hitting. And then the third problem you get is that, like, after you have a list of, like, you know, 40 bugs that the fuzzer's found, it becomes increasingly hard to teach the fuzzer not to keep finding the same bugs over and over again. So you have to, like, go into the fuzzer code and say, like, don't generate this pattern because this is a known buggy pattern. Well, now your program that's generating random programs gets increasingly complicated as you find more and more bugs. And, like, are you really able to exclude X pattern, Y pattern, especially if you don't really understand what is and isn't the buggy pattern? So with both the Nvidia fuzzer and the AMD fuzzer, I started, it just started slowing down and the Nvidia fuzzer kept finding like the same bug over and over again and it was hard to exclude it. And when I say it was hard to exclude it, I mean it was hard for the model to exclude it because I didn't try. I just kind of asked and it wasn't doing a good job. So then I just had the, you know, idea. And again, this isn't like, this isn't a really novel idea, but I just had the idea like, why don't we try reading the code? I mean, it's exactly the same thing that Anthropic was doing. With this Project Glasswing to find security vulnerabilities, right? I'm just, was using it to find compiler bugs. And to me, that was, it was kind of shocking how well it worked. It was also shocking how expensive it was. But actually, I mean, the update to the article, which I think I just added yesterday or the day before, is like, so we published on day X, and day X plus 1, Um, and Anthropic releases Opus 4.8, and in addition, they update Claude code to have this new UltraCode mode. And so, and they claim that UltraCode is kind of for exactly this thing, like big orchestrated projects. So I tried to use that to find bugs, and my initial, like, read on this is that it's way better. Um, and in particular, it's way more token efficient. So I was able to do big scans over LLVM using UltraCode. That used about a tenth of the tokens as I was using before. I would say the quality of the bugs is actually higher, the number of bugs is lower, but I'm okay with that because I was throwing away a lot of the bugs that I was finding before anyway. How much of that is different— how much of that is due to 4.8 versus UltraCode mode is really hard to say. Um, my guess is that a lot of it has to do with UltraCode, just sort of based on the other, like, the the fact that I don't feel like 4.8 is a huge intelligence step up from 4.7 in, in terms of what I can see. So therefore, like, this big step I would attribute to the harness. But again, I, I really have no idea.

Jordan Nanos12:03

Okay, yeah, really interesting. I, I got a lot of questions that come to mind there. Maybe the first one is to take a step back. If prior to the article update and the 4.8 release, would you have had in your mind a way to estimate I know this is hard, but the amount of time it would've taken to do a project like this without the assistance of the coding models, and then possibly to extend that to the new model and the UltraCode feature.

Justin Lebar12:31

Project like this, meaning like read the code and look for bugs and actually find meaningful bugs.

Jordan Nanos12:37

Maybe start with the writing the fuzzer itself, running it, and then reading the code itself.

Justin Lebar12:43

Yeah. Well, we'll start with reading the code, which basically is impossible. I mean, I don't think that any— I, yeah, it would, it would be impossible and it would fry your brain. If you really, really, really cared, you could try, but actually you wouldn't do it this way. You would, I think the only way really to do it would be to write a formally verified compiler, which people have done, but, so that's like the CompCert project, for example, but is, not really used in production. I think that's really the only way you could have done it. Um, so, but then in terms of writing the fuzzer, yeah, it would've taken me weeks and it took me days instead. Um, again, it's a little hard to say like, what is the quality difference in the fuzzers? Like, I didn't look at the new one, but it found a lot of bugs. Um, and especially for like the fuzzer for the closed source Nvidia compiler, at some point, that's kind of all you can do is find bugs. Like, you can't, you can't really fix them, and you also can't really evaluate, like, how many root cause bugs are there. And it's even kind of hard to evaluate, like, how bad is this bug, because in the end you don't really know what is the buggy pattern, you know. You just know, like, I was able to trigger it with this particular code. It could be much bigger than that, or it could be really, really specific, ultra-specific to what you wrote. You don't really know. So that's why I hesitate to like make big comparisons, right? But like, it was really easy to find a list of like 40 bugs that are like easily reproducible. And every one of the ones that I looked at, I looked at maybe 4 or 5. Every one of those was like, there was no undefined behavior. Like it was very clearly a bug. Like you have a series of, I think it was like max and subtract. You do like max, subtract, max, subtract 4 times, and then you get the wrong answer. Like, it's just kind of very clear that there's something wrong.

Jordan Nanos14:27

I think the most interesting or telling statement that you just made is the fact that like, it would be roughly impossible to take the approach of reading every single line of output from a compiler and then identifying the, uh, not even line of output, like line of code inside the compiler, right?

Justin Lebar14:43

Like, and even if you read them all, are you going to be smart enough to notice that there's a bug there? Most of this, most of the bugs that I've gone through and I've gone through in the x86 backend, more than 30 of them. I'd say almost all of them I never would've noticed. Even if you had said there's a bug in this 100 lines of code, I'd find it here. I like never would've been able to find it.

Jordan Nanos15:03

So does this leave you feeling optimistic or pessimistic about the state of open source and closed source compilers and then the state of the models as far as like how they can contribute to this sort of, well, this sort of like infrastructure that underpins a ton of software that we're using right now?

Justin Lebar15:19

I feel super optimistic. So long as someone is going to be willing to spend the tokens, even at $10 grand, it's not that expensive for, you know, even a small company like Semi Analysis, much less a big company like OpenAI or Google. But are they actually going to do it is the other question. So in that sense, I feel really optimistic because we are finding real problems. Even just when I've been writing the fixes, of course I have the model like check I mean, first of all, I have the model write the fixes itself, but then if I go in and change it, which is happening about 50% of the time, I have the model check my work and it frequently finds bugs in my work. So in this sense too, it's like dramatically improving the quality of the code that we publish. Um, but like part of the reason I wanted to publish the article was to get other people to do the same thing because it does still take human time, um, to filter through the bugs. And also it just like takes human time to convince some boss to let you spend time and money on this project. But if we do, yeah, I think it can dramatically improve the quality of the software that we're shipping.

Jordan Nanos16:28

So I don't know if you have plans to work on this project at all going forward, but is there any sort of like next step or ideas that you have that have sprouted because of the work that you did with FuzzX?

Justin Lebar16:37

Yeah, so one obvious next step is like actually fixing the bugs. I've fixed, the LLVM people have been really great at turning around reviews, so I've fixed most of the, what I've been able to identify as the high priority bugs. And again, I should say what I've been able to identify, what the models have identified for me as high priority bugs in x86 backend, um, that it, that it found. Um, so that's been really great. Uh, I think one interesting next step that I haven't tried would be to apply the same read the code, um, approach to the Nvidia closed source compiler, but the code that you're reading there is the assembly language as opposed to the like You know, the source code. I would be really fascinated to try that, especially with a more powerful model like Mythos, but maybe even 4.8 would do a good job at it. I literally haven't tried. I didn't want to like blow Dylan's token budget on it again.

Jordan Nanos17:35

Yeah, I got, I got, I'll keep my comments to myself about people feeling free to spend tokens currently. Token maxing, man. It's a virus. Okay. So if there's anything that, I don't know, listeners or like the people who potentially are reading the repo or trying to do a project like this themselves, you think should be taking away, or the more general audience that maybe is listening and learning about what a compiler is for the first time, like what do you think the key takeaways are for people like that?

Justin Lebar18:06

I think this, the takeaway for me is less about compilers than about like software development in general right now. We know that anytime you make a new tool for finding bugs, you're going to find bugs. So it's not surprising that LLMs are doing this or that my fuzzers are doing it. But I do think that the rate at which I've been able to fix them has been a lot higher than it would be otherwise. So the rate at which I've been able to triage them has been a lot faster than it has been otherwise. I mean, A tool that finds, you know, a thousand flags, a thousand pieces of code is suspicious, and like one of them is an actual bug, like that's not actually that useful. It's a huge amount of work to go through it. Um, and the LLMs so far have been quite good, especially 4.8 with UltraCode. So, you know, we're talking like just a couple days old at this point. Um, so I think like part of the callout I would make to people is like, if you have a codebase that you care about, um, it's worth throwing an LLM at it and seeing what it finds. If you have the budget to spend $10,000 on it, like, you should. And if you have the budget to spend $1,000 on it, that seemed to be what it was costing me to throw UltraCode at it, which, like, is actually extremely reasonable. Um, but, like, I don't know, I, I, maybe this is the compiler engineer in me talking, but, like, I always feel I feel like bugs that you find yourself are much, it's much better than a user finding them, you know? And on the other hand, like, there often isn't an incentive to do that because nobody gets handed a medal for, like, preventing a fire. Instead, you get handed the medal for, like, oh, the customer filed these 10 high-priority bugs and I, like, got right on it. But I think we have an opportunity to, like, really improve the quality of software that we're shipping if we choose to use the tools, and if we choose to spend the tokens on it, I'm hopeful that people will do that. And that's part of why I wanted to sort of write this article.

Jordan Nanos20:05

Yeah, that's certainly optimistic from my point of view with the discourse right now being something along the lines of slop AI, slop cannon producing a bunch of buggy software. But from your point of view, the kind of fundamentals of this, if used correctly or used for a specific purpose, would allow people to ship less buggy software. Like, that's just kind of— you just need to focus on the right things, I guess.

Justin Lebar20:32

Yeah. I mean, again, I try to be scientific about this and not make larger claims than I have evidence for. I haven't seen the evidence that, like, the AI slop cannon is actually buggier than the human slop cannon. Um, you know, people often say, like, how can I possibly trust that this model code doesn't have bugs in it? Oh, it's Nice that you've never written a bug yourself. I do agree that right now, like, the AI is slop in that it's not as beautifully designed as I think my code is beautifully designed. That's also true for code that anybody other than me writes, but in general, I think, like, it's worse than what I expect from, you know, like, my best collaborators. So, in that sense, it's true. It's like, and when I say it's not as beautiful, I mean, like, I don't enjoy reading it as much. But I haven't seen the scientific evidence one way or another arguing that it's actually buggier. And, and it could be that like the AI slop is harder for LLMs to find bugs in. It could also be that it's easier for LLMs to find bugs in it, or it could be equally hard. I, I also have no idea.

Jordan Nanos21:38

Yeah, super interesting. I think certainly more bloated, but I agree it's a toss-up whether it's actually producing more or less bugs. I mean, per line of code, maybe more per hour or per day that you're shipping features, probably less or something like that. 'Cause at least I can fix 'em more or fix 'em, do different things and fix 'em quickly. No other questions jump outta mind for me right now. Is there anything that you think is left unsaid about this article or this project you wanna cover?

Justin Lebar22:10

I'm excited to see what other people do with these tools, and I hope that people go and use them and like also just report on them because what I did was like extremely unscientific. Um, it was a case study. It's a story. I did X, uh, we're going to get better data if more people do it and talk about their experiences. Even that of course isn't scientific. It's, you know, the plural of anecdote is not data, but I'd love to see more anecdotes about this and then maybe we even try to approach it scientifically, but I'd love to see, you know, us focused on finding bugs, not just that are security bugs, but in databases and in other compilers and in web browsers that aren't security bugs. I think there's a lot of potential for this kind of thing.

Jordan Nanos22:53

Agreed. Yeah. I hope you, uh, I hope you continue to work on this and that you continue to provide these anecdotes in the open, develop in the open. It's been really fun following this project from the sidelines, just learning a little bit along the way. Really cool stuff.

Justin Lebar23:06

Thanks. Good talking to you.

Jordan Nanos23:08

All right, take care everyone. Thanks for listening. Nice job, Justin.