← All episodes

Ep. 020 - Anthropic vs OpenAI Usage, Margins, Meta Compute, Future of MSL (Tokenomics) | Crystual Huang, Max Kan, Joey Brookhart, Jordan Nanos

semianalysis · Jul 18, 2026 · 50:46

AI transcriptdiarizedcorrected9,995 words69 nuggetssource ↗audio ↗
Jordan NanosMax KanCrystual HuangJoey Brookhart
Jordan Nanos0:05

Hello everyone. Welcome back to Semi Analysis Weekly. We're here with episode 20 with the Tokenomics team. Things are moving fast. We are recording this on Wednesday, July 15th. I assume by the time this episode comes out, there's gonna be a lot more releases of models, disclosures from these companies about their profitability, revenue, uh, how many active users they've got. Regardless, we're going to record a point in time right now, talk about token budgeting, Meta's compute strategy, MSL futures, the release of Fable, SOL, and Anthropic's profit margins. So excited to dig in. With me today, we've got Max. How you doing, buddy?

Max Kan0:44

Doing great. Thanks for having me.

Jordan Nanos0:46

Welcome. We've got Crystal. How's it going, Crystal?

Crystual Huang0:48

Hey, Jordan.

Jordan Nanos0:51

And we got Joey.

Joey Brookhart0:53

Hey, Jordan.

Jordan Nanos0:53

Cool. So we're going to start with token budgeting. Crystal, over to you. We got lots of conversations that we've been having with enterprises over time and just asking them like, what's going on? At SemiAnalysis, we're still token maxing, but others are moving into the time of austerity. Can you take me through a little bit about what you found with this article?

Crystual Huang1:13

So there's a lot of like enterprises right now that are cracking down on their budgeting because they think their employees are spending too much and a lot of people are getting scared because they're like, oh, what if I don't have enough tokens to do my work? I don't think that's necessarily true then, and it applies to most people, right? Because a lot of people aren't even getting close to that limit. There's probably a few handful power users at each of these companies which will really be affected. But other than that, I think it's been— it hasn't been as big of an impact, I think, on most people's workflows as it has been portrayed on social media and stuff. There's been a lot of different strategies that companies have been imposing, whether it's like on a per person basis or like a monthly budget for the entire company from what we've seen. And I think the per-company basis obviously probably makes a lot more sense than like on a per-person basis, just given that some users will generate more like use from it than others.

Jordan Nanos2:05

And how do you see people actually enforcing this? Like clearly you can burn through a budget faster if you're going with the super ultra premium Max thinking model, but in some cases when you're actually enforcing a budget, you take different approaches, right?

Crystual Huang2:18

Yeah. Yeah. So it dep— I've seen some people actually, so I was talking to some people and they are at like a smaller company and they have a much smaller budget. So the way that they do it is they try to optimize how they're using the like more expensive tokens for like Anthropic or OpenAI. So they use like a cheaper model to like process a lot of it and like condense it down to like a smaller prompt. And then they use the more expensive models or they use something that their company's not counting. So there's a lot of like discrepancies in terms of what models actually count. Some people, or like some enterprises are counting their own models into the budget, but some enterprises aren't. So if there's like, if they aren't counting it, a lot of people will tend to like really push those to the limit, use those a lot, and then use like the ones that actually cost money.

Jordan Nanos3:05

And you're seeing people, maybe Max, I can bring you in on this one. You've, you've been seeing a lot more people burning tokens for coding rather than other use cases, right? Like that seems to be the market that's growing the fastest.

Max Kan3:18

Yeah, I think coding, like software engineering in general, is just by far the most token-hungry use case. And I actually think this is why a lot of the token budgeting discourse is pretty— like, I think a lot of the budgets themselves are pretty uninformed. Like, I hear people saying that they want to make sure their sales guys aren't using Opus or Fable to write emails, so they should definitely be using Sonnet instead. And it's like, dude, like, generating your email is basically free. From a token's perspective. Like, it does not matter if you use Opus or Sonnet to write this email. Like, I have no idea why this is the policy you're enforcing to try to like reduce your token spend. I've also heard lots of stories, like certain companies where only the engineering team is allowed to use like Claude Code or Codex, and then, you know, everyone else only gets like Copilot or something like that. I think that's pretty short-sighted, and it's honestly, you shouldn't have a caste system where your engineers are at the top and everyone else is just forced to use like dumber AIs. You should really be giving everyone the opportunity to try these tools and figure out how you can become more productive with them.

Jordan Nanos4:19

Yeah. Joey, are you using these models right now? How would you feel if I got access to it and you didn't?

Joey Brookhart4:24

I would probably be— I do use, yeah, some access, not as much as like Max and other people on at SemiAnalysis, but that's probably just the nature of our work or my work right now. And that's changing a little bit. But yeah, I'd be, I'd be pretty pissed cuz I still use Opus 4.0 a lot, like to pull in, to do more like agentic research tasks that will pull in a lot of like different data sources, do more like research type tasks. Um, but yeah, I'd be, I'd be pretty mad.

Jordan Nanos4:49

How do you think about the ROI question? I think a lot of people, the, the whole motivation behind token budgeting is clearly that somebody is not seeing a return on this investment, either that or they're just like not seeing it yet. But I think in a lot of that research you're doing, you're seeing that there's a lot of benefits, right? To using it.

Joey Brookhart5:05

Yeah, I, I think like with the market being so coding focused today, so probably 70% plus of ARR at Lab ARR right now on the API side is in coding. And then two, like from what we've heard, it's very power user, a power company driven. So it's pretty widespread, like coding. There's no like major big company for Anthropic, like Meta is probably 3 to 5% of total Anthropic ARR, but it's heavily driven by power users. If you're at a point where you're a power user at one of these companies, I think the ramp data is like the top 1%, you know, the 99th percentile of companies, they spent $100,000, you know, per year on AI per employee. And if you're, if you've gotten to that point, like where you're spending that much, you're obviously getting pretty good ROI. And I think like our conversations with a lot of enterprises were, these top engineers could go, you know, well above, you know, they're, they're limited, but you know, whatever their budget limit was. But if you were someone who was spending, you know, an insane amount of money, you know, you were getting budgeted already pretty quickly if you weren't seeing ROI. Um, and I think, you know, there's always projects that you'll try that don't get ROI, you'll scrape that, you'll scrap that and then move on. But I think it was, yeah, it was pretty obvious that because the market is so coding-focused, like people are getting ROI, they continue to spend. I think we continue to hear that net new ARR at both Anthropic and now OpenAI with like Codex and 5.5, 5.6, like since late March, spend is broadening. There's clearly a lot of like API token spend today. Yeah, there's a lot of organizations like in the, you know, that aren't traditionally, you know, doing a ton of, you know, development type work or focused. And they're, you know, they only allow their like middle back office workers to use, you know, Copilot 365 or Sonnet, you know, things like that. But I think the market will continue to broaden out.

Jordan Nanos7:03

Um, do, do you guys, did you guys have a take on the balance between API usage in applications versus like coding plans when it comes to people's individual usage? Like, do you ask them when you're talking to them if they are using a Claude Code subscription versus API paying per token?

Max Kan7:24

Well, I mean, anybody who can use a subscription definitely should use a subscription just 'cause it's so subsidized. It's just, you know, OpenAI and Anthropic are very wise to this. And so when they have their enterprise plans, like SemiAnalysis is on a, you know, Anthropic enterprise plan, they just, they don't let you use a subscription. They force you to pay per token 'cause they know that's where the margin is. But I think in general, like anyone who can use subscription does use subscription and it's just the people who need more limits than that, that pay for API.

Crystual Huang7:51

I've even talked to like a few people at like startups and stuff and they're not on like the team or the enterprise plan at all. They just buy 5 however many like Claude or OpenAI subscriptions they need a month, charge it to the company card, and just keep doing that. If they hit the limit on however many accounts they already have, they just buy another one for the month because it's just more cost effective that way for them to operate than it is for them to actually use the API billing.

Max Kan8:15

Yeah, if you're a startup that you're like young enough to where you don't care about like security or any of the other enterprise guarantees, like just absolutely buy, you know, 10 subscriptions per person instead of, you know, paying for API credits. Like I think the recent study we did on this, we had a decently viral tweet said that for Codex, I think $200 a month, you get like around $12K worth of API credits. And for Anthropic, it was like $200 a month for around $8K. Like obviously that's a no-brainer for anyone that is able to pick between those two.

Jordan Nanos8:49

Can you dig into that a little bit more, Max? Like when you're talking about the gross margin profile of these businesses and then you're considering the subscription plan versus the API, like what does that, what does that look like? I'm gonna bring up the chart on screen that I think is the key one from that tweet thread.

Crystual Huang9:09

Yeah.

Jordan Nanos9:10

Yeah.

Max Kan9:10

So essentially, obviously when you hear that like a $200 a month plan can generate $8,000 worth of API credits, the subscription is going to be much lower margin if the average user has anywhere near 100% utilization. And so really the key question is sort of like, what is the actual average utilization across our entire user base? And this chart that Jordan shows here is like kind of the break-even percentage for the various plans. We can see for the Claude Max 20x plan, it's like just 10%. For the Pro and 5x, it's like a little better at 20%. OpenAI, because they are sort of even more generous with their subscription limits, is slightly lower at 11.4% for the Plus and Pro 5x and just 5.7% for the Pro 20x. We're pretty confident that the average utilization, especially for these like 20x plans, is much higher than 10% and 5.7% respectively for Anthropic and OpenAI, which would mean that like forget being, you know, worse margin than API, which we think is like likely 85% plus for Anthropic at least, it might just be like negative margin in general for these 20x plans. And so of course, like these businesses want to move as much volume as possible to API. I think maybe Joey, since you wrote the newsletter on Anthropic's business, this would be a good segue to talk about why we think Anthropic is potentially in a stronger position than OpenAI.

Joey Brookhart10:42

Yeah, and I think too, like, once you move on to enterprise subscription plans at Anthropic, there's no usage included in your subscription fee. It's all, it's all done at API pricing. Yeah, I think at Anthropic what we found was really interesting is because the business, you know, today is so much more, and this is, this is changing. But because it's, you know, 80% plus of ARR is on the API side, and that is, as Max showed, is really high margin, you know, they've gotten to a point where they're profitable to today. You know, and some of that is on a non-GAAP basis. So excluding stock-based compensation, but like operating profit was profitable in Q2, you know, could be profitable to the tune of $1 billion plus in Q3. Some of that's a function of they've just grown so much. I don't think they can hire enough or like plow enough back into training as they'd like. But because of that margin, that kind of gross profit advantage that they have today, they can plow as much back into training as they want and can kind of extend their model advantage and model capability advantage. And we think that's something they should obviously try to do. So just because it's, you know, very profitable for them today. You know, it doesn't mean they should go run at, you know, 30% EBIT margins right away or as quick as possible. Show that this is a very profitable business model, but continue to invest as much as you can into training as possible.

Jordan Nanos12:07

Yeah. Can you, can you talk about the mix between consumer and enterprise, assuming that consumer is those paid plans and enterprise is pay per use on the API? Like when you, when you look at these things, clearly the bulk of like Twitter users are using the subscription plans, but then you get like one really, you know, big customer who's spending so much more money on the API tokens that it kind of drowns out everything else. And that, I guess, makes it a good business.

Joey Brookhart12:36

Yeah. And because Anthropic's really focused, they have been historically focused on enterprise. So it's 90, you know, it's 90% plus enterprise. It's, you know, versus ChatGPT has gone into, you know, started having a consumer up in, you know, back in Q1, 60% of ARR was in consumer and only 6% of the 950 million weekly active users end up actually paying for a plan. And most people pay for the $20 a month plan. It's a much different business mix because then, you know, OpenAI has to subsidize and pay for, you know, not subsidize, but there's a cost to serving those other, you know, 900 million weekly active users that are free. So it lowers their gross margins on a blended basis by about 20 points. You know, that's changing. So I think, you know, we think personally with some of the rise in OpenAI's API business since late March with Codex and then 5.5 and 5.6, that the 60/40 consumer enterprise split, you know, has flipped now to 40, you know, 60 consumer versus enterprise here in Q2. And it's, you know, from what we're hearing like in the channel, especially out of, you know, token-as-a-service businesses like Bedrock and Foundry, is that, you know, OpenAI has recognized this and it's starting to shift. And even, you know, ARR, like net new ARR is starting, you know, to look a little more even between Anthropic and OpenAI.

Jordan Nanos14:01

Yeah, and it's kind of shocking, like with the release of 5.6 SOL, the Guys like Thibault on X who are tweeting about the daily active user counts are saying that they're celebrating going to 7 and then to 8 million active users. But to be clear, they've got over 900 million weekly active users on the free tier of ChatGPT. Like they've got a lot of room to grow to just get people using the number one paid product. Maybe they got a few extra users using other paid products, but I would assume the bulk of people using the paid products are using Codex, right? So that's not a great conversion rate at this point of, in terms of who's actually using the full paid product from the, you know, base of 900 million.

Joey Brookhart14:52

Yeah, it's funny. You actually see it too in the paying rates for like even the consumer version of Claude. Like Claude is like 50, like 9%. Of free users end up paying. So it's a much more like focused user base. I think if we think of those 900 million weekly active users, a lot of it's like Google search type replacement, things like that. And then, you know, obviously because of that consumer focus, they've been working on advertising a lot. Consumers traditionally monetized, you know, people aren't willing to pay. The retention curves are typically pretty poor, um, and how, you know, people, how many people are paying, you know, 3 months later, 6 months later, 12 months later. Ads is a tough business model. It's tough to scale, but, you know, maybe that's the eventual, you know, they've put out some pretty crazy targets for 2030, you know, advertising revenue. But yeah, it's, consumer's a tough business. Enterprise is, you know, obviously given the margins on tokens for API, like enterprise is the place to be. I think Anthropic had a pretty good lead because they were so leveraged to coding. Early, you know, the first 6 months of this year and it seems like it's shifting to more of a kind of equal two-horse race here in mid-July.

Jordan Nanos16:04

Yeah. Crystal, what are you seeing in that two-horse race? Like, are people that you're talking to using Anthropic, Codex, both?

Crystual Huang16:14

I feel like a few weeks ago when I was talking to people, people were pretty heavy on using Anthropic Claude code specifically. And if I asked them why, they would say that's just what I've been using, right? And I'm used to using it now. But now as people in like the consumer section start using like ChatGPT a lot more than like Claude, that effect in like the consumer sector kind of starts trickling into like enterprise a little bit where like these people who were previously like Claude Code, like power users, didn't touch Codex at all. They're kind of like, okay, what if I like consider tinkering with Codex? So even though consumer is like a really small portion of the pie and the margins obviously aren't as great as enterprise, like, and historically speaking, OpenAI hasn't done as well in the consumer market as like Anthropic has, but like they have the, that traction in the consumer space that eventually like it's going to trickle in and we're already kind of starting to see that in like just this past week already.

Jordan Nanos17:13

Max, what's your take? OpenAI versus Anthropic, where do things stand for you?

Max Kan17:18

I think I remember last time I was on this podcast, we were talking about how it was, it was right when 4.5 came out., and we were talking about how things, uh, were looking really dire for OpenAI at the start of the year. Um, I think ARR growth was slightly flat in like March and April, which is like the alarm bells were definitely ringing, you know, in Sam's head when that happened. But then we said that 4.5 would likely be an inflection point, like this was actually a really good model, you know, probably Opus 4 level. And I think that prediction has played out even like sort of more quickly and more optimistically than we would have guessed. Like Some people here are saying that 5.6 is as good as Fable, sort of just like straight up, even though it's only half the price. And we also likely think it's like a much smaller model, which would be sort of a testament to OpenAI's, I guess, like training abilities. We think that like, as Joey mentioned earlier, net new ARR, like month on month is likely comparable between OpenAI and Anthropic now, which is pretty shocking. Like people, people were counting OpenAI out a little bit. I think Jordan might have been one of those people, but they're, but they're back, guys. They're back.

Jordan Nanos18:23

I am nothing if not flexible. It's always been about the model, guys. Like everybody's like, when are you gonna stop hating on OpenAI? As soon as they ship a model that's good and it's good. I mean, for the things that I want to do. Yeah. I'm, yeah, basically default OpenAI at this point, despite the application not doing what I want it to do with multiple tabs and stuff like that.

Max Kan18:45

Wait, are you using the Codex app or like the CLI? CLI. Which, which is your form factor? CLI. I feel like they don't want you to use the CLI, dude. Like they, they very actively want you to use the app. What is, uh, what is pulling you back?

Jordan Nanos18:57

Tabs. I need like multiple tabs going. Yeah.

Joey Brookhart19:01

Uh, okay.

Max Kan19:02

I think, I think the way you're supposed to do this, 'cause I agree, like the first time I tried using Codex and I was like, every single time I open a new chat, it's just like lost in the ether on like, you know, the left-hand sidebar of all those tabs. I think how you're supposed to do it is like you have one pinned thread, like per project you're working on, and then you just like ask that thread to spin up subthreads anytime you have like a discrete ask. And so you actually like never even interact with the vast majority of threads, like in that left sidebar. It's just like your couple pinned threads that you use.

Jordan Nanos19:32

Yeah, that sucks. Second of all, Second of all, I'm in remote SSHs all the time and they're remote and in the thing sucks. So yeah, both of those are, I'm very against it.

Max Kan19:49

Makes sense.

Jordan Nanos19:50

I think they are working on this, but yeah, I'm, I'm just, I'm not reading the code to be clear.

Max Kan19:57

Thanks for clarifying.

Jordan Nanos19:59

Just that, yeah, last time you were on, I, you know, we had that discussion, but you know, still have to run some things, still have to run some things from the terminal and see some output in the logs at some point. And this is just I'm just like not part of the standard Codex thing, but I'm not the primary market. Like I understand that they're targeting Vibe coders around the world, you know?

Max Kan20:20

Yeah.

Jordan Nanos20:20

We need to help people make more B2B SaaS is like the core business model of this. And that's, I'm not the target market and that's cool. I'll keep using the CLI. It's all good.

Max Kan20:31

Actually, I wonder if like internally most of their software engineers are using the app or the CLI. At Semi Analysis, for sure. I think all of our real engineers are like CLI, you know, diehards, right? Like they'll never give it up.

Jordan Nanos20:41

That would be a good question. You gotta ask them that. Yeah. Yeah.

Max Kan20:45

Well, next time we, I see some OpenAI people, I'll ask.

Jordan Nanos20:48

Sounds good. Yeah. Okay. Let's, uh, let's change gears and talk a little bit about who's gonna come in third. There's been some releases, uh, that we've covered a little bit. xAI is back. They've got Grok with Cursor bolted on and things are going well. Meta has Llama Spark out in the world. They've also stacked up a whole bunch of compute and they've teased launching a NeoCloud, which of course xAI was first to do. So what do you guys want to talk about first, the models or the compute for these guys?

Max Kan21:25

Let's do compute. I think it's more interesting.

Joey Brookhart21:27

I was going to tell you to walk through the models.

Max Kan21:28

I think, yeah, I mean, I think quip TL;DR the models is like, they're not bad, but at this point I don't think it matters unless you can deliver like a true frontier model. And they're obviously not true frontiers, so who cares?

Joey Brookhart21:41

Do you think anyone will use the— Max, would we want to use their API for like product?

Max Kan21:47

I don't think anyone's going to use the API until it's a frontier model. Okay. Yeah.

Joey Brookhart21:51

So when they, when they have an API, their token as a service business will be, will be fire. It'll be fire. 5% margins too.

Max Kan21:59

And I mean, that's only if it's like better than whatever the best Anthropic and OpenAI model is at that time. Which I think easier said than done. I think like Grok 4.5, Meta Llama Spark 1.1, they're only useful as like, uh, just like proving points along the scaling curve. It's like, okay, these guys, unlike Gemini, they're still like kind of in the game right now. Um, you know, they can train like adequate models. It's not quite frontier yet, but there is a non-zero chance they can catch up, I think is the way to view them. I will personally be shifting, uh, zero of my token usage to Either Grok 4.5 or Llama Spark 1.1.

Jordan Nanos22:37

Okay. But you said the other name, which is shockingly in my view, in 5th place right now, Gemini. Do you share that view?

Max Kan22:45

100%. I think they're like clearly in 5th place and like, unless Gemini 3.5 Pro is like better than the, you know, industry chatter that we're hearing, I think they're going to stay in 5th place. Maybe forever.

Jordan Nanos23:01

Maybe forever, Max. Okay. Talk to me about compute. So staying in 5th place forever would kind of be related to the quality of the model, but the quality of the model is kind of tied to how much compute these guys have.

Max Kan23:14

Yeah. Yeah. And I think, so I have to say that when I first saw the xAI-Anthropic deal where Elon rented like basically all of Colossus 1 to Anthropic, I think I incorrectly viewed that as Elon giving up when in reality, like I think the really important question is like, do you have the ability to claw back all of your compute, assuming that you have proven that you can truly reach the frontier by scaling up your compute? I think that's sort of like the view that Elon is taking with like Cursor and xAI. He's essentially saying like, I'm going to leave you guys with enough compute to sort of prove that you're capable of reaching the frontier. And in the meantime, I'm going to monetize all the additional compute by like renting it out on these like extremely high margin, you know, 3x, maybe 4x market rate rental deals. I'm going to make sure there's this like 90-day clawback clause included in every single one of them such that if you do really well, I'm just going to claw back all the compute, give it to you so you can, you know, make digital gods. And I think that is like a valid strategy for someone who still truly thinks they can like build RSI. And I think you can really think of where it's like OpenAI and Anthropic are monetizing their compute by just like selling inference. Elon is doing these short-term compute deals instead. I think this is fine. I think on the topic of Meta Compute, like, you know, there are the rumors that they're going to become a NeoCloud. I think as long as they do these like xAI-style deals with the clawback, or they do like, or they just like give it all to Rexus, or they like do token as a service, like any monetization method. That allows them to, in theory, claw it all back if MSL proves they're doing well. I think that's like an excellent smart decision on Meta's part. Probably allows them to be more aggressive with buying compute in the short term, and it still like gives them a chance to build RSI in the long term. I think the issue with Google and why they're in 5th place, aside from their model like Gemini 3.5 Flash being bad, is that they have signed all these long-term compute deals where they cannot claw back the TPUs that they're giving Anthropic, right? That's just like committed for the long term. And I think it indicates like a lack of conviction on their part that they think they can build RSI. And I think, you know, Noam Shazeer, John Jumper, like all these really high-profile guys at DeepMind leaving is sort of because of this fact, like they are AGI-pilled, their leadership is not AGI-pilled, and they think that the only way they can do— the only thing they can do in this situation is to go work somewhere else.

Jordan Nanos25:44

Wow. Okay. Hot takes. So the two things I want to dig in on there. The first is related to the clawback. So this is a valid strategy, but you can basically only do the clawback once and then nobody's ever going to buy from you and trust you again. Or do you think you can actually oscillate between these two?

Max Kan26:04

Well, I think in an ideal world, you only need to do it once because you only do it if you're confident that like, I have, I'm going to be able to reach the frontier by continuing like scaling, right? But I also think it's possible you could do it more than once. I think in principle Anthropic is like viewing this as a short-term deal. And I think if Elon said like, hey, actually I want it all back 6 months from now. And then a year from now he's like, JK, we actually still can train a frontier model. If Anthropic is still ripping at that point and it's in desperate need of like compute capacity, why would they say no?

Jordan Nanos26:37

Yeah, I guess financial motivations can make Friends out of enemies. So second thing is just on the topic of the business, maybe Joey, I can bring you in here. What's a better business, selling tokens of your frontier model or close to frontier model, or selling compute access on the short term at a premium over the market?

Joey Brookhart26:59

Selling your own tokens is probably the best model. I mean, even like even selling someone else's tokens, if you're AWS or Azure, is a pretty good business model as well. If you can participate in some of the upside revenue share, but if you can rent a above, you know, X market rates, it's not a bad business. I think before, like in, like from the Meta Compute perspective, it's been floated for a while. You know, Meta Compute has been floated since like Q3 of last year and it gives Zuck a backstop to say, okay, if MSL isn't successful, we have means to essentially make decent ROI on our capex spend. And it allows them now to spend a ton more on capex in '27. Potentially 28, because there's a market and, you know, you've gotten these market signals that Anthropic can pay, you know, 3x market rates and, you know, you can do a small deal here. And that's in, you know, the background. Investors can feel comfortable that, you know, if MSL is deemed, you know, unsuccessful, that, you know, Meta could have a pretty nice, you know, single client or multi, you know, few client like Neo Cloud business just serving the labs.

Jordan Nanos28:06

Yeah, but I mean, I thought this chart that you guys had in the Meta Compute article was striking where you can see that the xAI plus Google deal, they're monetizing this compute at 4x the market rate. So I think in the past, everything you're saying makes perfect sense, but I would challenge a little bit that like this near-term xAI GB300, Red Cell deals, aren't a really good business. Like there's clearly high margin in this, right?

Joey Brookhart28:38

Yeah. I mean, that's extremely high margin. I think like when you think about the, you think about like the ARR per, per like gigawatt, I think a lot of people would be happy to get $48 billion of like actual token ARR, like of inference ARR.

Jordan Nanos28:49

So I don't know, that deal doesn't seem very economical for, for Google, but you know, they're telling a, you think they're, MetaCompute's, them telling a story to the market that they're going to be able to replicate this big xAI thing, which is probably a one-time thing unless demand keeps going so exponential that there's such a compute crunch that everybody's fighting for access to the GPUs. And while we have them and we can monetize them at like 4x the market rate from NeoClouds, where NeoClouds, I think we already, I think our house view is that NeoCloud is a pretty good business on its own at like $12 billion per gigawatt, approaching $50 billion per gigawatt. It's like, Yeah.

Joey Brookhart29:27

And like, and the labs have gotten so good at the ARR per megawatt or gigawatt, like numbers gone up a ton. So token throughput's been pretty impressive. Like just, you know, the models as well. And then too, how they've monetized them, you know, the new models price higher. So like if you look at like Anthropic, like revenue per megawatt was like $16 million per megawatt last year, or $16, you know, billion per gigawatt. And now You know, that's more than doubled in 3 quarters. And I think there's some people that think that can double again in 2 to 3 quarters.

Jordan Nanos29:58

So it looks like on that trend, I've got this chart on screen right now. So total token as a service market, 2 things jumped out at me. One, how fast this thing is growing. Like I thought the market was pretty big in Q4 of last year, and that looks puny compared to our forecast that you guys have on this chart. Towards the end of this year, like way more than doubling, as you said. But the other thing that strikes out, that jumps out to me is just how bad Microsoft Foundry is doing here. Can you— Yeah.

Joey Brookhart30:32

And that'll be revised up with the recent OpenAI API success. But I think up until, because OpenAI, you know, they were, you know, Foundry was 90% OAI and, you know, OAI was so consumer focused and API. Didn't really, this is not the latest version of this chart. Yeah, Foundry's, you know, doing a little better now. Maybe Vertex is doing a little worse given, you know, Matt's or Max's thoughts on Gemini in that chart. But yeah.

Jordan Nanos31:00

You gotta subscribe for the most recent tokenomics forecast.

Max Kan31:04

These companies— There's always— We can't be giving this away for, you know, for free, dude. You gotta pay for the tokenomics forecast.

Joey Brookhart31:10

We were nice enough to leave the axis on that chart this time. But yeah, I think we're like, we're of the view and like when we talk to enterprises, you know, token as a service is a very popular choice for consuming tokens. You know, AWS and Azure have massive businesses with Fortune 500, Global 2000 companies that, you know, spend 9 figures plus, you know, on, on cloud a year. And if I can go to an existing vendor and have more model optionality, increase my ELA or my credits and then burn that down, it's pretty attractive for both AWS and Azure. And I think like, you know, if it was 5% indirect like sources, like, you know, this token as a service for like Anthropic 6 months ago, like that's probably, it's probably 20% now of your business.

Jordan Nanos32:04

Crystal, what's your take on people using the Big 3 hyperscalers, token-as-a-service stuff versus the smaller startups like a Together, Base10, Fireworks, anything they can get on OpenRouter?

Crystual Huang32:19

I feel like as we see more of the bigger enterprises, so like I feel like a lot of the financial services industry still hasn't unlocked like and use AI to the level that they probably can and should be. And that's where like Azure and Bedrock are going to be benefiting from because most of them probably already like buy their services, right? And then once they do get like internal approval or whatever to be using these AI tools, like they'll just buy it through whatever existing channel they'll have already. Whereas it's like probably going to be like, that's where the bulk of the market pretty much is.

Jordan Nanos32:56

Makes sense. Yeah. Max, how about you? What do you think of token-as-a-service from like the startups, how fast Together, Fireworks, Base10, they're growing compared to the hyperscalers that are kind of tied to the big labs, but maybe not actually selling a lot of tokens to the startups of tomorrow?

Max Kan33:17

Yeah, I mean, I think Together, Base10, Fireworks, these are all super impressive businesses. I do actually expect like open source token volumes to grow slower than frontier token volumes, but that's only because I think frontier token volumes are going to like absolutely explode. I think open source token volumes are also going to explode just to a slightly smaller degree. And I think like, it's kind of funny, I feel like if you talk to the VCs investing in like Together, Fireworks, and Base10, and you ask them for the explanation why, oftentimes it'd be with like inference is going to be like the largest market ever. And so it doesn't even matter if like these companies can only capture like a super, super tiny slice of this like extremely large market. That's good enough for us. And I think that thesis is like honestly more or less right. I don't expect these guys to ever be doing like more volume than an OpenAI or Anthropic or like nowhere near that. I, I actually think like this is in terms of like global token volume, they might be like the highest percentage they've ever been, like they ever will be rather today. But they'll still be like good businesses in the future, I think. Yeah.

Jordan Nanos34:21

How about the rumors that they can hitch their wagon to some of the bigger labs? So like xAI starts to win a little bit and then Fireworks grows because they're exposed to Cursor and actually have some, I don't know, exposure to that growth.

Max Kan34:38

Yeah. Well, I don't think that's gonna like last in the long term. I think the only reason why Fireworks got that exposure is because originally Compose was post-trained on Chemi, right? I think if you're a frontier lab that has made like a new model from scratch, there's no reason to believe why Fireworks' kernel engineers would be better at optimizing that model than your own engineers. So I don't think they can get like any share of that market in the future. Okay.

Jordan Nanos35:05

So why do the hyperscalers get a share of that margin with the big three?

Max Kan35:11

I mean, because they just have like, and I'm sure Joe can speak more on this, but they just have like the enterprise distribution that the NeoClouds don't. Like Anthropic is not giving up this margin to Bedrock so that Amazon engineers can like make Claude run faster on, you know, B300s or whatever, or it's just so, yeah, yeah, sure, Trainium, but it's just so like all the existing customers that are reliant on Bedrock. Can also use cloud models. Makes sense.

Joey Brookhart35:39

Yeah. Yeah. I think the deal, I mean, there's, we hear more, we don't hear anything, but people tell us, you know, they think the deal could get re-struck. 'Cause it's obviously, you know, it was very beneficial for AWS. I think we've written on AWS and their margins and how they monetize, you know, versus just selling the bare, you know, compute for Anthropic to run inference on, you know, to get that 30%, you know, or 20 to 30% revenue share. That just falls down to the bottom line is pretty attractive. AWS, like Azure, lesser extent GCP, like obviously have massive customer bases. People's cloud estates are there. The enterprises are very comfortable, security and compliance and everything now with cloud, even your more regulated industries are. So it's a natural place for them to want to buy. But to see that much of the economics for Anthropic is obviously, yeah, I think a lot of that was because when the deal was struck and kind of the incentive to make Trainium work on that, you know, they were able to, to run with those terms and make it pretty favorable. Mm-hmm. So yeah, we'll see what happens if those revenue share deals can continue.

Jordan Nanos36:45

Okay, let me make one more attempt at the bull case for these token as a service companies that aren't the hyperscalers. For AWS and Google, obviously a big portion of the benefit of getting token as a service from them is that their engineers are are working on Trainium and TPU, and that might be different than the experience that the lab has with the GPUs where they have more experience like running this themselves. There's a whole class of chip startups that are coming to market right now. And an obvious way in which they come to market is by partnering with these token-as-a-service companies where, hey, the, the biggest Frontier Labs aren't going to spend a whole bunch of time optimizing for, you know, the 7th best chip startup that's coming to market right now. But if they do, and that chip startup strikes something really nice for a given model that makes it a lot more cost-effective or a lot higher performance to run instead of Nvidia, then they've got a shot at doing something super unique. Do you think that's a, you know, potential future in terms of where these kernel engineers who have learned a lot end up going and spending their time over the next few months or years?

Max Kan37:51

This is a good point and it's reasonable in the short term, but if any of these new chip startups actually reach sufficient scale, I think the labs will just dedicate teams to making their model run really good on their chip. I think there's just, I think there's no world in which you have a new accelerator that's meaningfully better than Nvidia and it's also being sold at large volume that OpenAI is relying on together to serve their model on that chip. And so I'm just working directly with that company to, you know, develop the first-party capabilities to run their model on that chip. Yeah.

Jordan Nanos38:29

You think it's gonna go the way of Cerebras where the chip companies to be successful are effectively gonna have to become a Neo cloud for themselves and yeah, exactly. Relationship with the frontier labs. Okay. Makes sense. Maybe we can finish by talking a little bit about MSL with Meta, particularly Max. I just loved the, like, crash course on RL that this article turned into, not necessarily how it works from a perspective, but how it works from a business perspective, like where people buy data, how they build these environments, what sort, what the market looks like. Can you kind of give a, an over, like previously in the world where everything's pre-training, it's just like whoever has a, you know, frontier class team and the most compute can just train the biggest model and win. They follow the scaling laws and go there. But now there's a, scaling law related to RL? How is this playing out in your mind at a high level?

Max Kan39:26

Yeah, yeah. I think it's, I think A, it's like really important for everyone to understand that reinforcement learning or RL is like probably the most important scaling law for improving model capabilities today. And there are a lot of people who believe that the only thing that's stopping the models from being able to do like literally anything a human can do on a computer is having sufficient RL environments. This is sort of like data that lets the model like try to complete white-collar tasks itself and tell, and it can sort of like repeatedly try and clean the task until it fully learns how to solve it. And there's like an entire new industry slash like supply chain of these RL environment startups whose like entire job is to convert real-world economically viable tasks into these RL environments that can then sell to the labs and allows the user to improve their models. Yeah, you see like most of the main players on this chart Jordan has pulled up here. I actually think like some of these ARR numbers might be slightly understated. I think it's like pretty consensus that the total sort of data budgets at the frontier labs this year, so this is like primarily OpenAI, Anthropic, Google, Meta, and xAI, but also the long-tail companies like, you know, Meta, sorry, Amazon, Microsoft, like Thinking Machines. I think all those summed together is going to be like well over $10 billion. That's sort of like a 10x relative to last year. It's very possible it 10xs again in '27. And this is just like a hugely important market for improving AI capabilities. This is probably like maybe the only market in the world where in customer demand, I guess maybe compute is the other one, but like in customer demand, it's just not even a question for all these startups. Like they will never have a contract turned down from the labs because it's too expensive. It's just a question of like, can you scale up, you know, creating high quality data fast enough? And if the answer is yes, like we will pay any price for it. So yeah, this is like, I guess maybe as one other like side note, one of the reasons why Anthropic's models are the best at coding today, um, or at least like they definitely were before FiveSixSoul came out, is that they were by far the most aggressive from buying coding data from all these RL environment startups. Um, I think some of the other labs are starting to like realize this and catch on, uh, but it is definitely like a super important industry everyone should be aware of.

Jordan Nanos41:58

Can you dig into a little bit of the process to create some of these tasks? You kind of ran through this in the article by describing how Meta has moved thousands of engineers doing this work and also dispelled the notion that this work is like meaningless soul-crushing stuff and is actually, you know, pretty both economically valuable. And I don't want to steal your thunder, but you had a nice line on that one.

Max Kan42:24

I think I said it was both potentially more economically valuable and intellectually stimulating than like your average big tech job. So like, yeah, I think a lot of people think like they hear the phrase AI data and they still think like, oh, we have some random people in the Philippines who are like drawing bounding boxes or like labeling text as like NSFW or something. And that's just like, like, dude, the models have already fully solved that. Your data is only valuable if it's doing something that the models don't already know how to do. And so what this means in the case of software engineering is that in order to make a good, like, software engineering task today, it typically needs to be something that would take, like, a really good human engineer, maybe like a full day of work to solve in order to be something that, like, the model can't already munch on itself today. And so if you want to create this data, like, Essentially, you need like a really good human engineer to sit down and like think of an example problem that he would actually want to do. You then need him to like create a verifier, like usually a set of integration tests, maybe along with a rubric that can check if the model actually successfully completed this like day-long task or not. Obviously, this is easier said than done. And then you also need this engineer to like write a prompt for the model that like asks it to do this task, but also in— it sort of like needs to fulfill these like two competing factors where it needs to be like 100% fair and then like unambiguous what you want the model to do because you can't sort of like incorrectly fail the model during training because it like successfully did what your prompt asked for, but your prompt was just not specific enough. And so like the model didn't know it actually had to do some extra thing. But at the same time, like you need your prompt to be realistic and natural sounding and representative of something like a human would actually ask an AI to do in the real world. And it's like very difficult to get this balance right. We've heard for like the highest quality coding tasks, like the labs are willing to pay well over 5 figures for a single task. And, you know, this is obviously already entering into the realm of like how much you would pay a decent engineer for a full week of work. And so I think that's just sort of just like dispel any myths about this being like easy, you know, mind-numbing work. To all the listeners out there who are looking for a new job, like if you are really good at creating RL, like tasks, you can make 7 figures plus annually at this point. So maybe, maybe consider that as a new job option.

Jordan Nanos44:52

Yeah. Maybe we have some listeners excited. Can you maybe actually say like one click lower, do you have personal experience with striking that balance between easy for the AI to do and not impossible for the AI to do? Like what sort of intuition could you give to a listener about what that means?

Max Kan45:12

Honestly, it's like, it's always changing. So, so back in the day, like, I actually did sell some RL environment data to the labs myself. Don't do it anymore. But it was like much easier to create data even just like 8 months ago when I was still doing it than it is today. I feel like today it often looks like you need to identify a specific failure mode that like you are aware of in the model, and then you need to create like RL tasks that specifically target that failure mode. And really the only way to know if it's like at the right difficulty for the model model, it's like you have the model try solving it 10 times and then you see how many times you're successful. Um, you sort of just like repeat that iteration. So it's like, I, I don't know if I can really provide any like blanket advice of how to find the right difficulty other than if this is like your first time trying to do this, your first thought is almost certainly too easy. Try making it 10 times harder, like try that, and then maybe, maybe it'll be like in the right difficulty level.

Jordan Nanos46:10

Interesting. Cool stuff. Yeah. Okay guys, we've got a whirlwind tour. Anything you think has, uh, that I've missed as we've gone through token budgeting, Anthropic's profit margins, MetaCompute, MSL? What, uh, what have we missed? Crystal, what have we missed?

Crystual Huang46:28

I don't know. Nothing I can think of.

Jordan Nanos46:31

Just been too busy watching the World Cup here. We got that. We got Wednesday, July 15th, right when we're recording this. We just got to watch England get knocked out as Argentina stormed from behind for a nice 2-1 victory.

Max Kan46:43

That was crazy, dude. There's like a bunch of British people in the office right now and they were just like all depressed downstairs. It was crazy.

Jordan Nanos46:50

No, Joey, how about you? What's, what's on your mind as we wrap up here?

Joey Brookhart46:54

Nothing. Token spend is up and to the right right now, so it's good. You know, I'm excited for hyperscaler earnings starting next week. We've got Google, you know, Amazon and Microsoft the week after. It should be pretty good, you know, on the top line.

Jordan Nanos47:10

Okay, let me go around the horn. We'll close by getting a vibe check from everybody. Joey, what's your vibe like on the market? On the market?

Joey Brookhart47:18

We're like, I like, I don't like, as long as LAB ARR is going up like, and at a good pace, like it doesn't decelerate, I, I think like, I think things are fine. Yeah. And like right now it's going up. OpenAI is catching up. Like vibes, vibes should be good. That's not what you're seeing in semis, you know, the last few days. But that's just summer, you know, summer momentum. It doesn't work. So it'll come back. People come back to the office from vacations and semis will rip end of year and, you know, things will be good.

Jordan Nanos47:46

Yeah. We're gonna see an acceleration after the summer pop. Joey's gonna shoot 84. And see ARR go up and be happy.

Crystual Huang47:53

Okay, Crystal, how's your vibe? It's going— I have to say, I think it's going good. Same thing Joey said. I need to leave San Francisco before the AI bubble pops. So very quickly moving away from San Francisco before it's too late. Crystal, you think it's a bubble?

Joey Brookhart48:09

Yeah. Max, we should publicly talk about our bet here. Yeah. Yeah. Joey and I— Yeah, we have the loser. So Max chose the over-under of Anthropic ARR at $400 billion to end 2027. 2027. Yes, 2027. And we're at like, you know, we're in probably—

Max Kan48:34

you can't give them the actual number right now, Joey. They gotta subscribe to the tokenomics model.

Joey Brookhart48:37

Yeah, we can figure it out.

Jordan Nanos48:40

But you know, okay. But does somebody have Who took the over? Who took the under?

Joey Brookhart48:46

I took the under. I took the over.

Max Kan48:49

Jeremy and Joey actually both took the under, but we haven't, we haven't decided what like the actual bet's gonna be for yet. No, the bet is like—

Joey Brookhart48:55

I think at some point this year we'll figure it out.

Max Kan48:56

Did you guys—

Joey Brookhart48:57

if you lose, you have to write—

Jordan Nanos48:58

did you guys see our pot at Raise last week? I had— no, I pulled out a Canadian 20, Rake pulled out 200 Singapore dollars, also some rupees. We had Dylan pull out some euros, like we had 5 different currencies going on the table. So I, whatever you guys are betting, that's funny. I will throw in Canadian currency and take the over with Max because we are— let's go. We're exponential extrapolators here.

Max Kan49:22

Yes, yes, exactly. Exactly.

Joey Brookhart49:24

The loser has to write a newsletter post of why they were wrong.

Max Kan49:28

Okay, that's good. That's good.

Joey Brookhart49:30

No, that's what it was.

Max Kan49:32

I didn't realize we agreed to this. Like, I must have missed that in the Slack thread.

Jordan Nanos49:35

Are we doing like Frontier Lab fantasy here? We need to pick a model for a given week and set up our team.

Joey Brookhart49:41

We'll bet on Jeremy's spend. Jeremy's bet on Jeremy's— we bet on Jeremy's weekly spend and we'll have like an over-under and over-under, get everyone involved. No, it's like, isn't Meta building an internal like Polymarket or Kalshi? Like, we'll do that for semi-analysis. Someone can write code like Polymarket.

Max Kan49:57

Dude, Meta needs to shut down that effort right now, dude. Like, what are they doing?

Joey Brookhart50:03

Well, we'll have a— we'll have a Polymarket for betting on Jeremy's token spend. There we go. I like it, guys.

Jordan Nanos50:10

There's no way that market could be—

Max Kan50:12

No way. Totally fair.

Jordan Nanos50:17

Hey, Jeremy, I've got 6,000 rupees riding on you. Need you to hammer the data center model down for this week, buddy. Okay. Well, guys, I appreciate you taking the time. Hopefully the listeners enjoyed it. Nice, uh, default a little bit at the end. Let's, uh, let's all get back to work. Keep tracking those tokens. Yeah.