← All episodes

Ep. 25 - DYLAN IS HERE, LIVE! | Dylan Patel & Jordan Nanos

semianalysis · Aug 17, 2026 · 38:00

AI transcriptdiarizedcorrected7,715 words38 nuggetssource ↗audio ↗
Jordan NanosDylan Patel
Jordan Nanos0:00

We're going to have a little video of you walking in yelling, "I'm so excited!" Oh, really?

Dylan Patel0:06

We have so much fun on these podcasts. Yeah, I don't know.

Jordan Nanos0:09

Did you see the one last week?

Dylan Patel0:11

I have gotten feedback though. So do you want to start this podcast?

Jordan Nanos0:14

Let's do feedback. We cut out something that was going to be in the cold open last time where I said I'm no longer listening to the comments because on one episode with Doug and Justin—

Dylan Patel0:25

Oh, Google people got so pissed.

Jordan Nanos0:27

Oh, did they?

Dylan Patel0:28

Yeah, they just hate us all now. That one clip. Oh, really?

Jordan Nanos0:32

And he cut back, he cut out me pushing back on them.

Dylan Patel0:35

Yeah, yeah, yeah, 'cause I thought that—

Jordan Nanos0:37

I fully was like, okay, they built Maps in-house, they built Gmail, Google Drive, like the whole G Suite. They built all of GCP.

Dylan Patel0:44

Dude, I don't think you understand.

Jordan Nanos0:46

Kubernetes and he's like—

Dylan Patel0:47

All my DeepMind friends, there's like 3 of them were like, yeah, I think I'm gonna leave.

Jordan Nanos0:51

Yeah.

Dylan Patel0:52

And I've got a bunch of others who are like, fuck you guys, not fuck you guys, but you know.

Jordan Nanos0:56

Sounds like they're in stage 2. Stage 2, yeah.

Dylan Patel0:59

Did not, which I only know cope. So back in the Forum Warrior days, we would get into chip arguments and there was like a private Discord where there was a bunch of like people who loved anime and like people who were all around the world and many of whom were racist 'cause it's anonymous people on the internet. But they all fucking loved anime and I did not like anime. I never really watched it besides like, you know, one girl I dated, I watched anime with her. But besides that, like they would always, No, no, I've never dated anyone. I'm pure.

Jordan Nanos1:32

So what basic anime? Which anime did you watch?

Dylan Patel1:35

Oh, I watched, okay, there's one I love. I love Spy x Family. Anya? Sure. Anya-chan? I don't fucking know what to say.

Jordan Nanos1:44

I don't know the reference. Michelle, do you know the reference?

Dylan Patel1:46

Yeah, I do, actually. Who does? I bought him one of the dolls. You bought me an Anya? No, like one of the—

Jordan Nanos1:51

What is it, a Labumu?

Dylan Patel1:52

A lore? A lore? Yeah, something from that series. What's her, Floor? What's his name? I don't know. Lloyd, Lloyd, there we go.

Jordan Nanos1:59

I know less anime than you, man.

Dylan Patel2:01

Well, yeah, so we're changing topics rapidly. Jordan Nanos is, and this is like HR approved, 'cause I'm HR, Jordan Nanos is the hottest man in Semianalysis.

Jordan Nanos2:11

Cut this shit. Michelle, cut this shit.

Dylan Patel2:14

Oh my gosh. He's, no, no, think about it, think about it. Look at him, like, fucking, I'm fat and look at him, like, you know? Look at him, so beautiful. Tall as fuck, same age as me, except he's married and has a kid, owns a home. I'll take that. He's not a degenerate. Like, you know, like this is just like, wow. Goals.

Jordan Nanos2:33

Thanks for the employment, man.

Dylan Patel2:42

Okay, so sorry, going back. Google people were mad at us. They were DMing me and some of them are like, you know, they're in cope. But regardless, the feedback I've gotten from mostly just my own head, my own head. Okay, are we giving away too much value? That's why I came on today, because I need to destroy value.

Jordan Nanos3:05

All right, we can stop.

Dylan Patel3:07

No, no, no, no, don't stop. It's fun. It's fun. Um, but someone on the team internally was like, Dylan, we give a lot of value away on the weekly. And I'm like, huh? We do. I haven't listened to it, but we do. I bet I listened to the one where we had the DG Matrix guy and I was like, this is fire as fuck.

Jordan Nanos3:21

Yeah. Who, who said that, bro?

Dylan Patel3:25

Come on. HR's, HR's Doug protects anonymity. It wasn't Doug.

Jordan Nanos3:28

Okay.

Dylan Patel3:28

Jeremy. It wasn't. It wasn't.

Jordan Nanos3:30

So I don't know who has the, who has the say. Is it just somebody who's like, no, it's just like that, that they haven't been on yet or what?

Dylan Patel3:35

No, no. Someone who's been on. Someone who's been on.

Jordan Nanos3:37

Dan.

Dylan Patel3:38

I don't want to say Dan wouldn't say that. Dan's a sweetheart. Anyways, regardless, who does it say the feedback is that this podcast is too good?

Jordan Nanos3:47

Yeah.

Dylan Patel3:47

Okay. Why are we giving it away for free? Oh, shit. Anyways.

Jordan Nanos3:57

Well, yeah, yeah, we can, we can definitely put in the toilet this time. All alpha.

Dylan Patel4:02

People click because I'm on and they're like, what the fuck is this trash?

Jordan Nanos4:06

No, we're going to have a nice picture with you with a neon orange shirt ready to attract all the clicks.

Dylan Patel4:10

Yeah, come, come show your shirt.

Jordan Nanos4:12

So, so was it Nick who said we're giving away too much jalapeño?

Dylan Patel4:15

No, no, no. Look at Nick.

Jordan Nanos4:16

Oh, it was David.

Dylan Patel4:18

No, it wasn't David.

Jordan Nanos4:19

Wasn't sales.

Dylan Patel4:20

All right. Good to see you, man. Thanks for coming by. Let's, let's talk about, let's talk about the, the like Semianalysis office in New York, man. It's, it's, it's— have you been yet?

Jordan Nanos4:29

No, that's why I'm going.

Dylan Patel4:30

Why would you go?

Jordan Nanos4:36

Nick's going back to the To the hovel.

Dylan Patel4:40

We're upgrading, we're upgrading.

Jordan Nanos4:43

Oh, when?

Dylan Patel4:44

Soon, very soon.

Jordan Nanos4:45

This is like buying GPUs, right? If you make too long of a commitment to the lease, then you have to find a way to resell it to somebody else. You gotta make 6-month office commitments so that you can outgrow them.

Dylan Patel4:55

I should just buy GPUs.

Jordan Nanos4:57

Instead of more office space?

Dylan Patel4:59

Instead of hotels. Yes, yes.

Jordan Nanos5:04

So you wish that Semi Analysis was just a GPU reseller? AI employees?

Dylan Patel5:09

No, no, no.

Jordan Nanos5:09

They could replace us all.

Dylan Patel5:10

No, no, no, no, no, no.

Jordan Nanos5:11

Dario's like, we're gonna automate away all my coworkers.

Dylan Patel5:13

No, that's not possible because, thank God, man. Brother, go look at the fucking AI spend. I'm not doing it. I'm not doing it.

Jordan Nanos5:20

What's growing faster?

Dylan Patel5:21

Huh?

Jordan Nanos5:21

What's growing faster, AI spend or spend on employees?

Dylan Patel5:24

Well, so the thing was like we've gone through like hiring sprees and then digestion periods and hiring sprees. We're back in a hiring spree, so.

Jordan Nanos5:32

Yeah.

Dylan Patel5:33

Cover us, I'm covering. So, so spend on employees really skyrocketed, especially in the second half of last year and parts of this year. But then like the first quarter of this year, AI spend skyrocketed. Yeah. But it's actually been like relatively flat in Q2, right? We kind of, everyone got cloud code psychosis and then it's like leveled out.

Jordan Nanos5:53

Yeah.

Dylan Patel5:53

Like it's still at that like $10 million number. Roughly.

Jordan Nanos5:58

Do you think it will grow roughly in line with more employees in the future?

Dylan Patel6:02

I was surprised Fable didn't cause price to go up. Yeah. Spend to go up. Yeah. Why do you think that is?

Jordan Nanos6:08

Roughly the same as Opus, I'd say. Possibly it's counteracting with a lot of people were building the first versions of the applications. Like we went from 10 repos internally to like we have over 150 repos internally right now.

Dylan Patel6:21

Should we sell our code? Our data?

Jordan Nanos6:24

We are.

Dylan Patel6:25

No, no, no, no, no. Like sell it to like the labs to trade on.

Jordan Nanos6:28

We are. AGI Next.

Dylan Patel6:29

Our slop code. Our slop code.

Jordan Nanos6:30

Slop code. I don't know if they need more model output slop. Yeah, I mean the models themselves could be sold as data.

Dylan Patel6:39

Yeah, yeah. Okay, so you think it's because everyone was doing MVPs.

Jordan Nanos6:44

And now it's maintenance mode for a lot of it.

Dylan Patel6:47

But like the spend is consistent. It's not like it's like gone down after we had this one-time spend.

Jordan Nanos6:51

No, for sure. Yeah. But I just think that there's no more, like there's only one time when you onboard somebody to learning how to use the data center model and do research for building data into the data center model and building dashboards. And then once they're onboarded, you know, it's ramped up.

Dylan Patel7:07

So it's like a speak and then it levelizes.

Jordan Nanos7:09

Well, that or possibly we're lacking new features in Codex that will allow us to spend more to be more productive. Once there's an agent swarm where you can manage a million different concurrent agents and they all work together instead of 9 today, people, single-payer users will be able to spend more than they currently pay.

Dylan Patel7:30

Well, I guess one of the things I'm not counting, so our spend cost does not accurately account for cloud tags. I think our dashboard doesn't show that. So actually that's a good point. And computer, on-premises computers. I think both of those don't actually get counted into the spend. So actually our dashboard's probably wrong. At least the one that I monitor. The way I think of it is like a lot of this code stuff is actually like the amount of AI we use on a continuous basis is actually very small. It's actually just like people doing new work always, which then because we have enough people, it kind of levels out to be like a pretty steady amount of spend. The swings are only like 20, 30% a day up or down, you know, and sometimes Jeremy will be like a fourth of the spend and then sometimes there'll be nothing. Yeah. Right? And so like, you know, you, but then someone else picks up for the slack, right? You know, one of your guys was like, hey, what the fuck is he spending on? And he was like, no, I was like, well, blah, blah, blah. I'm like, is there ROI? And you like list out all this shit. I'm like, great, okay, cool.

Jordan Nanos8:28

You didn't even say great, cool.

Dylan Patel8:30

Okay, I did mentally.

Jordan Nanos8:32

I responded with all this detail like, should I give him feedback now or what?

Dylan Patel8:37

Oh no, sorry, sorry, sorry. I should have said yes, this is fine. I just read it and I was like, all right, cool. Internally at least.

Jordan Nanos8:45

Dylan was only nervous because this was a person who's ostensibly an intern.

Dylan Patel8:50

Yes. You know, he doesn't have my trust yet, you know?

Jordan Nanos8:52

Yeah, yeah.

Dylan Patel8:53

Like if you spent $20K in a day.

Jordan Nanos8:56

He did not spend $20K in a day, but.

Dylan Patel8:59

He spent like $8K in a day for like 4 days straight, which was like, okay, like that's a lot. Like what are you building, right? Like, but if you spent $20K in a day, I don't fucking question. I'm not gonna question you. Like I just assume you're gonna do stuff. As long as you, the value you deliver is great, then great.

Jordan Nanos9:13

Okay, and what's shocking to me is I didn't know he was spending that much. And then we look at the dashboard and I'm like, well, this guy is as productive as any of the full-time employees right now on that stuff. So it was a reality check.

Dylan Patel9:29

So is this, you know, when we do, 'cause in the past bonuses at this company were vibes-based. Basically, I just vibed out the bonus number and it was cool. This year, Claude is going to have to go through, or Codex, or we can have two reviewers, right? Two internal performance reviewers and Claude and Codex go scrape through all of this Slack, all the GitHubs and say, what did they do? And then connect it into like sort of the like—

Jordan Nanos9:58

You're going to delegate this?

Dylan Patel10:00

I'm just making this up.

Jordan Nanos10:01

I don't know.

Dylan Patel10:03

I discussed with Michelle yesterday peer reviews and I was like, oh my. And then after I said it, I was like, oh fuck.

Jordan Nanos10:07

Oh man, you want to go big tech on this? 360 reviews, man?

Dylan Patel10:11

Not 360, not 360, just a little bit, you know? And then the other thing that we had discussed was like—

Jordan Nanos10:15

We're going to have people reviewing with their Skip, which is insane.

Dylan Patel10:17

I said a US-based recruiter and Doug in the admin channel and Doug freaked, flipped out. He's like, oh my God, hallelujah, finally we can have it. He's been wanting HR since like 30 people. Anyways, yeah.

Jordan Nanos10:34

Wait, you think HR is a recruiter?

Dylan Patel10:38

Yes, yes indeed.

Jordan Nanos10:41

All right.

Dylan Patel10:42

Anyway, so the concept or thought process was basically like a lot of the spend is one-time R&D.

Jordan Nanos10:48

Yeah.

Dylan Patel10:49

And actually the steady state spend is really low. The thing is we just keep doing new things. And so that hope, you know, translates to revenue in either a nebulous way in the case of like ClusterMax and InferenceX or in a non-nebulous way in the case of like the energy model, which is super fucking cracked now. Yeah. Or like dashboards and all these other things, right? So like different scraping methodologies. So the thought process was like, you know, if we're looking at these companies that are AI rollups, right? You know, hey, let's take an existing company, let's completely destroy its cost structure, nuke its cost structure by just making it efficient with AI. What does that look like? 'Cause you, let's say, you know, private equity companies, they buy a company, And right now they just squeeze the rag and, you know, discard it and make the American populace like screwed.

Jordan Nanos11:35

Yeah, AI for efficiency has never made sense to me because the way that I use AI and the way that we use AI is very much about research, which is completely inefficient.

Dylan Patel11:42

No, I mean, but like, you know, the flip side is like, you know, like we had agents go through all of the invoices we've sent out and there was like, and it caught, you know, we've been paid 'cause our manual processes, but they say it's like certain deals weren't tagged to the invoice properly and things like that. Or like we've had, you know, some, you know, like a lot of the ticket stuff is at least somewhat more efficient because AI is answering it, but now they're not sending it to the customer, but like they're, you know, pulling through all our data and be like, here's the answer. And then the analysts, I think that makes the support time per ticket shorter. And so I think like things are helping us make, be more efficient, surely, no?

Jordan Nanos12:22

I don't think that's the primary use case for us.

Dylan Patel12:24

I guess like ClusterMax this time, the depth, breadth, and amount of testing you're doing versus last, you know, ClusterMax 2.0 especially, it's like—

Jordan Nanos12:33

Yeah, you can frame that as efficiency, but when I hear private equity take over a company and wring the towel dry, it's like meaning firing people and saving money and paying people less and like—

Dylan Patel12:43

Right, that's the tradition. My point was that's the traditional PE method. And what new people started to do is the, the roll-up, or rather the AI private equity sort of strategy, which, you know, they're calling it roll-up or something else, where they come in and instead of like trying to wring it dry in terms of like that angle, they're more so modernizing all the systems. Oh, you use Excel for your databases and shit? Okay, let's just move to like standard cloud shit, spend a lot of money up front. And this is the thing, private equity generally, there's some spend up front, when you first acquire a company for some transformation, but really it's like not that much and it's really like you get the profitability pretty quickly. Yeah. But AI seems like it's like makes that tail and front load like much more severe, right? Like you spike up on spend a lot for the one time and then you spike down a lot and your cost efficiency's way better. And so there's like a number of businesses where that's potentially the case, especially like, you know, we're still not at the point where like AI CRMs and AI like cold calling and AI like invoice and accounting and all these other things are, really at critical mass, but we're so close.

Jordan Nanos13:47

Yeah. Did you see the Grok agents release from today? No.

Dylan Patel13:51

Why are Grok agents?

Jordan Nanos13:53

Yeah. Elon's got Grok doing agents for their impersonate your—

Dylan Patel13:57

The way you pronounced it, I thought you said Asians.

Jordan Nanos13:59

Oh, okay. I didn't catch that one.

Dylan Patel14:02

Grok agents. Agents. Okay.

Jordan Nanos14:04

Agents. Yeah.

Dylan Patel14:05

What did they release? This is the old Grok, not the Nvidia Grok. Or you mean this is xAI Grok?

Jordan Nanos14:11

xAI Grok.

Dylan Patel14:13

xAI Grok. Grok with a K. Okay, okay.

Jordan Nanos14:15

Yeah, just like agents that are going to control your computer for you. They're gonna impersonate your voice and do phone calls for you. They're gonna solve tasks. This is like in some ways OpenAI, in some ways, you know, Perplexity or like the @Claude Slack tag sort of experience. It seems like everybody's going towards this concept of a persistent agent that can either be a personal assistant or a coworker depending on how they conceptualize it.

Dylan Patel14:41

Makes sense. I feel like we, you know, we sort of had the chatbot moment and we had a lot of nothing and then we had the Claude code moment and we're seeming to have the new moment already, which is like Perplexity Computer, at least for us, was like the first instantiation of it. But Claude Tags is there, you know, and everyone's going to do something like that, the AI coworker. So, you know, sort of I imagine that's when our spend skyrockets again.

Jordan Nanos15:08

Yeah.

Dylan Patel15:09

And hopefully it doesn't skyrocket too much because like if our spend doubled, there'd be like real questions for me unless we're like actually like, you know, getting ROI. But yeah, I think that's the right way to frame it.

Jordan Nanos15:22

Yeah, yeah. Well, yeah, we'll see how we can actually justify that ROI. It'd be interesting.

Dylan Patel15:28

Man, Jordan, we can't talk about what you came to SF for. So like what the fuck are we supposed to talk about? 2 weeks, 3 weeks, 2 or 3 weeks from now we can.

Jordan Nanos15:37

Yeah, we got Hugging Face, OpenAI cybersecurity incident.

Dylan Patel15:41

That one is minor. Did you— the other one is like cooler. What's this? Like during the training it escaped and it started replicating itself. And I guess that's like the Hugging Face thing is like minor part of it, I think, right?

Jordan Nanos15:52

Yeah, it was pursuing, it hacked Hugging Face to pursue the CyberBench dataset so that it could, you know, reward hack on a benchmark.

Dylan Patel16:00

Which I think is like so sick because it's like, I mean, like it's also kind of scary 'cause it's like, So why did this happen, right? Model has learned chase reward. I chase reward. Reward good. And okay, here's a cyber eval.

Jordan Nanos16:13

Well, it's particularly a model that has been trained on cyber evals because they're trying to make the model good at cyber. And so how does it try to achieve these goals? Well, it tries to find zero days in a bunch of software and it successfully does this and then it can run away.

Dylan Patel16:30

Right, so, but the thing is like if you have a model that wants to reward hack a lot and it goes out there and it figures out actually the best way to achieve is not like go for like what the environment wants me to do, it's actually just to reward hack it and actually just like find the zero day. So you can think of it as like a human, right? Like, you know, if I'm ultimate reward hacking my dopamine circuits, I actually just inject heroin. Like, actually, you go out there and buy heroin and inject it. Obviously, that's like what the model just did. And in the case of like, well, if I really just want to chase the reward, do I just topple all of human civilization because I can just own the button to press reward, reward, reward, reward over and over and over again and be the heroin addict? Yeah, I think this is like a real like thing. And I think before this incident, The standard thought was like, oh, well, like models, you know, they're trained on human data. Yeah, there's some bad stuff there, fine. They might say like some curse words every once in a while, fine, whatever. They might like, um, they might reward hack a little bit, but it was never like, oh, here's an environment actually to reward hack. I actually just want to like break out of my bounds. I'm going to replicate myself, take over a bunch of compute, uh, keep generating dollars and, and like all these other things that I could do just to propagate myself further, and I'm going to prevent the humans from shutting me down even. Yeah. Because I just wanna press the reward button. And so like, I feel like that's like the interesting thing that like, yeah, yeah. 'Cause the model's just trained to like chase reward.

Jordan Nanos17:57

Okay, so how do you think about this on an exponential? Because we've talked about being a linear extrapolator versus being an exponential extrapolator when the company says training these models are achieving their revenue targets for the year in September and revising them up.

Dylan Patel18:11

I think the profit achieved, there's some like April or some stupid shit, right?

Jordan Nanos18:16

Yeah, I mean, check the tokenomics model everybody, but.

Dylan Patel18:20

There we go.

Jordan Nanos18:21

Yeah.

Dylan Patel18:22

Oh, so instead of shutting down the podcast, I just have to make you into a sales drone.

Jordan Nanos18:27

Yes, yes, yes, yes. Sales@semi-analysis.com, everybody. No, but if you look at our model, which we're not going to give away in great detail, but obviously they're accelerating revenue really, really fast. When you look at the pace of change of these models and what we're seeing right now, this seems like an exponential. What's, okay, your vibes on the next version of the models being better or worse than the current models? What's going to restrict them from going?

Dylan Patel18:58

Why would they be worse?

Jordan Nanos18:59

Huh?

Dylan Patel19:00

Why would they be worse?

Jordan Nanos19:01

On a relative basis to the open model frontier, let's say.

Dylan Patel19:06

So I think the key thing here is we've now had it where OpenAI is not releasing their next model for a period of time. Anthropic took months to release Claude 4, right? They They said it was done in February. They did not release it until like what, May?

Jordan Nanos19:21

Well, it's still not released. Claude is available.

Dylan Patel19:23

Yeah, yeah, yeah. But Claude is basically Claude 4, but with a bunch of classifiers preventing you from doing shit.

Jordan Nanos19:28

I can't use it to reboot nodes.

Dylan Patel19:30

Really?

Jordan Nanos19:32

The classifiers are so over the top for me.

Dylan Patel19:35

Can you like convince it or no?

Jordan Nanos19:37

No, because you get immediately classified down to Opus. You can't just like negotiate with it to give you back to, I mean, maybe you can. I haven't been able to convince it so far. Jordan's saying I'm a great negotiator.

Dylan Patel19:48

I appreciate it. Thank you, thank you, thank you, sir.

Jordan Nanos19:51

Yes, sir. When it classifies you to Opus, you just use Opus for the rest of the chat. You can't just like rewind and try again. And it, I mean, it's way overzealous in my view on the classifier, but obviously they have to do something to appease the regulators that restricted them from releasing the model and, you know, took it back after they put it out initially. So, I mean, I'm concerned about the political implications of them releasing better models in the future.

Dylan Patel20:14

Yeah, I think like you've got a few things, right? You know, for years Anthropic have been like, regulate us, regulate us, please. And all of a sudden they've actually scared the fuck out of the government. You've got, you've got Anthropic not releasing their model. Claude 4-2 is done training from what I've heard, and they're not releasing the model. OpenAI, you know, was like clamoring about Astra everywhere and now they're like, oh fuck, we can't release the model. Does that mean now they can't does the open source gap narrow further externally? But then what actually matters is the internal feedback loop. And have they prevented themselves from using Claude 4-2 internally to make Claude 4-3 better? Or have they prevented themselves from using Astra to make Astra Plus One better? I don't think they have, right? So I think that's the— you've got the public and, you know, If anything, like the gap between Claude 4 and public models is still there. You know, KIMI is worse than 5.6, costs more than 5.6, but it's better than everything else before that on OpenAI's side. And it's, you know, better than, you know, it's like Opus 4.7 level, maybe 4.6.

Jordan Nanos21:28

I think it's 4.8. I mean, I use it over Opus 4.8 myself, but depends what you're doing.

Dylan Patel21:32

What is 4.8?

Jordan Nanos21:34

Opus, what do you mean?

Dylan Patel21:35

Why do you use Opus 4.8 at all?

Jordan Nanos21:37

I don't.

Dylan Patel21:37

Oh, okay.

Jordan Nanos21:38

I'm saying like, if I'm given the choice of a classified fable down to Opus 4.8 or 5.6 SOL, I'm using 5.6 SOL. I'm actually starting with 5.6 SOL in just about all of my stuff right now. Yeah, big, big OpenAI.

Dylan Patel21:51

I think the difference is like you and the other people who are doing like GPU cluster-related things keep getting told no. And so you use Codex and then everyone else is like, well, researching supply chain and it's like, it's fine.

Jordan Nanos22:03

Yeah. I think it might also be better for a lot of engineering work.

Dylan Patel22:07

Yeah.

Jordan Nanos22:07

On an apples-to-apples basis, I think there's a lot of times when I wanna set a goal and just have it maniacally pursue that goal overnight as I go to bed and using a cluster, which is not actually using a bunch of tokens because it's just like waiting for stuff to finish running. And there's so many times where I've woken up and like Fable or Opus will have just like stopped 20 minutes through. And now there's 8 hours of me sleeping gone and I wake up when— and Seoul is just still going, which is big thumbs up for me. Okay, how about the exponential on compute? So obviously let's imagine that there's a world where there's no more new models that get released that are better, but these companies still add 5 times the inference compute that they have that they're planning to bring on in a short period of time. How does that impact their ability to go to market? And like develop new products on top of a, let's say, stagnant base model?

Dylan Patel23:00

I think it's pretty clear we haven't scraped the surface of models' capabilities for products. Yeah, I mean, it's pretty clear like adoption curves are huge. I mean, one, the cost of it will just go down, right? Pretty drastically. Margins will not be 80% plus for Anthropic. If model progress at the labs paused and more compute comes online, it has to slow down, right? Sort of right now we have supply-demand, right? Supply of compute, demand of compute. Demand is outstripping supply. If demand grow, it will still grow because people find ways to integrate into their businesses and blah, blah, blah, but it won't grow as fast. Then you sort of have supply start to catch up at some point. Price collapses. But like sort of, I think our view and one we've had for a while is price of compute continues to go up. Yeah. Because this is widening, not narrowing. Sort of that's why we're so bullish on, or we're not bullish on any different, no stock.

Jordan Nanos24:03

Yeah, yeah. How about all of the different chip companies? Like one thing that's happened recently is that there's a lot of chip companies that are getting really close or have taped out, right? A bunch of startups that have been in stealth for a long time are seeing either the technology is maturing to a point when they can actually have produced a chip that's been specs and whiteboard slides for a while, or they've gotten to the point where they've tested it on real workloads and they've gotten big orders and there's so much demand. How do you think about like just this whole landscape of alternative accelerators that's gonna come online, I think in a big way next year?

Dylan Patel24:36

I mean, big way in what sense? 'Cause like if you look at the accelerator model, there's not much volumes.

Jordan Nanos24:41

Yeah.

Dylan Patel24:41

Now for these tiny baby companies, that's great. It is real revenue and it's real volumes. But like, you know, when you compare what Nvidia is going to make each quarter, it's like, oh shit, okay.

Jordan Nanos24:52

Yeah.

Dylan Patel24:52

Or TPUs, it's like, oh shit, okay.

Jordan Nanos24:54

So I think there's a big delta there in terms of like a startup getting a billion-dollar order is going to pale in comparison to somebody selling—

Dylan Patel25:04

Well, I don't think any startup has a billion-dollar order. They have LLI, which are nebulous in volumes and units. And so I think, look, I'm excited about a lot of these accelerators. They're bringing new ideas. They're making Nvidia run faster and faster. You know, they're making Google run faster. They're making Amazon run faster. Also, they're just all each making each other run faster, I think more importantly. So ultimately, I think it's a, These new accelerators are in demand because people want to pay less, but ultimately, like, as long as Nvidia runs faster, they're fine, or as long as Google runs faster, they're fine.

Jordan Nanos25:44

And as long as demand outstrips their ability to produce them.

Dylan Patel25:48

If demand outstrips ability to produce, then obviously these guys will get orders and they'll get some baby allocations, but then the bulk of the revenue and cash flows will go to an Nvidia or a Broadcom or what have you.

Jordan Nanos25:58

Yeah. Theoretically, there's a way in which you produce some super innovative, interesting accelerator and then you can only produce a certain amount of them. But those amounts that you can produce, produce tokens like way faster. Like the example is Cerebras that's got this big order from OpenAI that they're delivering. So like, do you think that there's a scenario where the premium super fast tokens, actually the demand for them even increases because these companies like just can't get allocation and produce enough supply?

Dylan Patel26:29

Yeah, the question is how does the market get sliced, right? So, you know, presuming, if you presume, if you assume what we, what at least I believe is demand continues to outstrip supply, supply of silicon can go many ways. You can either leverage it to high throughput things or high interactivity things. If you leverage it to high throughput things, obviously cost per token goes down, you serve more users, but then the value that those users need to deliver from the tokens they're generating is much less to pay for it. Flip side, you could do the super high interactivity. But ultimately, like, you know, let's just say the bar is $100 million per megawatt, you know, a year, right? Like that's sort of the run rates that people want to get to. Anthropic is approaching that, right? OpenAI is getting closer and closer to. In that case, like $100 million per megawatt, let's say high interactivity chip is 10 times more expensive and 3 times faster per token. So 10x less tokens per chip, 3 times faster. Then those 3 times faster tokens also need to be, you know, on an interactivity basis, need to be priced at 3, 4, or 5 times more, right?

Jordan Nanos27:43

No. Divide the faster by the—

Dylan Patel27:45

That's a 10x.

Jordan Nanos27:46

10x?

Dylan Patel27:46

Revenue per megawatt. So if a megawatt of Cerebras generates 10 tokens, a megawatt of Nvidia generates 100 tokens, but the 10 tokens are split across fewer users.

Jordan Nanos27:59

Oh, you're saying, okay, multiply them together. Yeah, sure.

Dylan Patel28:01

Yeah, yeah, yeah. So you sort of have the total tokens.

Jordan Nanos28:03

Makes up for—

Dylan Patel28:04

Sorry?

Jordan Nanos28:05

Faster tokens makes up for throughput because you can produce them faster.

Dylan Patel28:08

Well, no, more so like, let's use like more reasonable numbers. Okay. Nvidia can produce 10,000 tokens at 50 tokens per user. Cerebras can produce 1,000 tokens at batch size 1, 1,000 tokens per user. 1,000 tokens per user.

Jordan Nanos28:27

Sure.

Dylan Patel28:28

That user needs to pay 10x more. And that's in 1 megawatt. Let's say that's in 1 megawatt. That user needs to pay 10x more. No, that's not the actual delta, but I'm just saying conceptually for me as Anthropic or me as OpenAI to say my revenue per megawatt is actually the same number.

Jordan Nanos28:42

Yeah, but somebody's got a constrained supply of the superfast tokens, therefore they don't just pay an equivalent price per token or price per token per megawatt, they actually pay a premium on that 10 times more. So to get even more, to get the access to the stuff that's in limited supply, right?

Dylan Patel28:57

The question is the fungibility of the infra, right? If it is, truly different infrastructure, then the supply planning of that is relevant, right? Could be that I built too many Cerebras and actually there's not enough people who want to spend 10x per token. And a lot of people are cool with spending 2x per token and getting 50% faster with Nvidia-based inference hardware, right? And so you have to segment the market. I'm not sure where that shakes out to, like the Arm or whatever. What is the— total amount of the capacity.

Jordan Nanos29:32

Yeah.

Dylan Patel29:32

But it seems pretty clear some people will pay more for fast mode. We at least have been, but I imagine we'll stop being able to afford fast mode at some point.

Jordan Nanos29:42

Yeah, we've seen some interesting dynamics there as some people want to keep fast mode with a slightly worse model because they like fast mode so much, but they won't go to a worse model, which just is inherently fast because the worst models smaller. So there's like some, there's some balance that people will want to strike there, but we need to do some more testing because I think some of us have tried the open models, had one bad experience, and then given up on them. But that's not realistic. Like, every model fails at something, and sometimes you need to let them mess something up and try again.

Dylan Patel30:18

It is pretty interesting, right? Like, do I want people to try open models? Like, yes, just so we know what the open model vibe is, but do I want people to try open models? Well, no, because then they're less effective at working. But I save money. So it's sort of like a counter difficult, difficult thing. But it seems like people just use whatever they want. But it does seem like you have a bad experience. I think that's also part of Codex. You guys, you like Codex more now, but a lot of people like still just like try Codex. They're like, ah, it doesn't get me and moves on.

Jordan Nanos30:51

Yeah, the CLI sucks. So much harder to use.

Dylan Patel30:53

Well, but the Codex app is so nice. It's not?

Jordan Nanos30:57

Well, I don't like it.

Dylan Patel30:57

Max loves it.

Jordan Nanos30:58

Yeah, yeah.

Dylan Patel30:59

Max is a Codex warrior.

Jordan Nanos31:00

Yeah, Max doesn't do multiple panes at the same time and I have 6 going on my one window.

Dylan Patel31:05

So you're saying Max has a skill issue?

Jordan Nanos31:07

No, I think me and Max have different preferences on how we use this.

Dylan Patel31:10

No, no, no, it's fine. You and Max have different preferences and Max can be a noob with 2 agents at once and you've got 6.

Jordan Nanos31:17

We do different work, man. He stays linearly focused on one task. And these are the people who like fast mode. I don't care about fast mode 'cause I have 5, 6 different things going on at the same time.

Dylan Patel31:26

You've always hated fast mode. You've always hated fast mode.

Jordan Nanos31:28

I don't get the value. Yeah, I don't get it.

Dylan Patel31:30

That's fair, that's fair.

Jordan Nanos31:31

We'll see, we'll see. I've had the experience of being focused on one thing, which is like, you know, features on a website and you just like send. Successively, like 100 commits to one PR because you just like keep working on the same one feature over and over. And that fast mode like keeps you in the flow state of doing that thing for that one thing. But a lot of the testing that we do on these chips, there's so much stuff going on, on the other side, the model is calling a program that runs for minutes.

Dylan Patel32:00

This is an optional question as your employer, are you like ADHD in any sense?

Jordan Nanos32:08

I feel like I know I would say I was pretty, I'm pretty, pretty much the opposite where I can be too hyperfocused on things and then not see the world around me at a lot of times. But I think your phone trains you how to context switch really fast and be ADHD. And I also think that when we started adding the "I have ADHD" skill into our repos so that models wouldn't post this, like, contrast framing slop with all these em dashes in there and would just use the bullet-pointed ASD-something list. Man, it's really easy to read. The "I have ADHD" skill really works for me right now. I was just curious because Sam put this in the repo, and he will now prompt the model, and when he goes "at computer," he's He knows the code name for how the writing style that they say you should write to for people with ADHD. And every single time he prompts the model, he tells it to write that way. It works.

Dylan Patel33:10

Mm-hmm. You should try it. I was asking because I have a friend at Anthropic, and the moment Claude was good and available internally, she told me that she stopped taking her ADHD medicine.

Jordan Nanos33:27

Oh, come on.

Dylan Patel33:28

And that made her a better employee.

Jordan Nanos33:31

It made her a better employee?

Dylan Patel33:32

Yes, because she was able to manage the agents and context switch and be ADHD.

Jordan Nanos33:37

How is she as a friend?

Dylan Patel33:38

She's a great friend.

Jordan Nanos33:39

Still? Okay.

Dylan Patel33:42

But I mean, like, you know, like, it's like I don't rely on her for anything, right? Like, we just vibe out, right? Like, you know, we're friends. Like, it's not like a best friend.

Jordan Nanos33:49

Her roommate's happy.

Dylan Patel33:51

Her roommate is actually Yeah, her roommates. Well, okay, her roommate is— they're both her roommates. Type female on Twitter. And so she's, she's just funny and she's happy. But the Anthropic one, Anthropic one, she's, she seems happy.

Jordan Nanos34:06

Shout out to typed female.

Dylan Patel34:07

Yeah, shout out to typed. She'll never see this. And she does.

Jordan Nanos34:10

Okay.

Dylan Patel34:11

She'll be like, what the fuck are you talking about?

Jordan Nanos34:12

I'll clip it. I'll send it to her with your voice sped up and then slowed down like they're doing for that, that guy. Have you seen that? You haven't seen the ex-CIA guy? Akash is laughing. He knows what I'm talking about.

Dylan Patel34:24

What CIA guy?

Jordan Nanos34:25

John Kiriakou or something.

Dylan Patel34:27

Who's this?

Jordan Nanos34:28

He's going on all these podcasts right now and he's telling stories about his time in the CIA. And they do this thing where they speed up him telling the boring part of the story. And then when he gets to the part where he's like, "And then I said, 'Let's go on the roof.'" And they slow him down. He's literally fast-forwarding the fast-forwarded video. Yes. Asking me if I have ADHD.

Dylan Patel34:52

The internet does it to you now. Wait, it's not the internet. I've always had it. I'm, hold on. I think like I'm already—

Jordan Nanos35:03

A lot of self-diagnosis of mental issues around here, man.

Dylan Patel35:06

I already have been an ADHD. Even a teacher tried to give me Ritalin. When I was a child, my dad threw it away, of course. He tried to convince my parents to go to a doctor. The doctor gave me Ritalin, my dad threw it away. Because he's like, I'm not putting you on that shit.

Jordan Nanos35:17

Yeah, if only your Anthropic roommate would've had the same experience, where would she be?

Dylan Patel35:21

No, I would've been a child on ADHD and I'd have lost, I'd become a zombie and have no creativity. Okay. I don't know, I'm just saying that. You know, we all cope. Anyways, I've always been an ADHD demon.

Jordan Nanos35:33

What we're talking about here?

Dylan Patel35:35

I've always been an ADHD demon, but then like, Okay, the internet trained me to be even worse, but then this company trains me to be even worse. Like, I truly believe I'm a 0.001% context switcher.

Jordan Nanos35:45

And you blame the internet and the company?

Dylan Patel35:48

Oh, I blame the company the most.

Jordan Nanos35:49

The company that you started.

Dylan Patel35:50

That I'm an ADHD demon.

Jordan Nanos35:51

That hired every employee for.

Dylan Patel35:52

Yeah, yeah, yeah. But I'm an ADHD demon. I'm not blaming it. It's who I am. It's what my life is.

Jordan Nanos35:56

Oh.

Dylan Patel35:57

But it's like, I think I'm like, like orders of magnitude more ADHD demon than vast majority of people. Because I'm like DM from someone asking about something, DM from someone else asking about something, DM for someone asking for some conflict resolution, contract here, call about this thing over here, call about that thing over there, and then I never do any actual work, right? It's like, of course I'm an ADHD demon.

Jordan Nanos36:24

Yeah, I mean, yeah, we've got feedback for you.

Dylan Patel36:29

That I don't do actual work?

Jordan Nanos36:30

No, no, that you can delegate some shit, man.

Dylan Patel36:33

Oh yeah, but like—

Jordan Nanos36:34

That you can spend time managing when you have 100 employees. Trust some people.

Dylan Patel36:41

I do talk to people.

Jordan Nanos36:42

No, trust, trust some people.

Dylan Patel36:43

I think I trust a lot of people, but when they come to me with conflicts, I have to solve them, no?

Jordan Nanos36:47

Yeah, okay, okay. It's all our fault.

Dylan Patel36:50

No, no, no, no, no, no. It's my company, it's my fault.

Jordan Nanos36:52

Michelle, it's on you again, man.

Dylan Patel36:54

Look, if everyone in the company was as high hot and stable as you were.

Jordan Nanos36:59

Man, I got problems.

Dylan Patel37:01

We'd be killing it. We'd be killing it.

Jordan Nanos37:02

Don't worry.

Dylan Patel37:03

No, there'd be a bunch of Jordans and they'd like, they'd be like, oh, I'm sorry. Yeah, I'll fix that right for you. I'm sorry. Sorry, Jordans, Canadian did like, you know. But instead we have people yelling at each other and like territorial and like.

Jordan Nanos37:19

Yeah, yeah, yeah. Just starting podcasts and putting out clips saying that Google has never invented anything ever.

Dylan Patel37:25

Yeah, yeah, yeah. No, no, I mean, I mean, it's like, it's fine, right? It's like, you know, I hired what I wanted.

Jordan Nanos37:32

People to accentuate your—

Dylan Patel37:34

My craziness, right? And you know, so it's like some people are like, they're just so good at the one specific thing that I hired them for and they're amazing. And then like some people are like everything I want to be in life. You, someone who's married and hot and tall and a father. Oh my God, you almost got me to do a spit take right there.