← All episodes

Open Source Wins, AGI Is Here, and Scorsese's AI Toolkit with CEOs of Cerebras & Black Forest Labs

all-in · Jul 10, 2026 · 1:03:57

AI transcriptdiarizedcorrected11,139 words53 nuggetssource ↗audio ↗

Synthesized from 53 insights · Jul 12, 2026

The Compute Gold Rush: Demand Outstrips Everything

AI infrastructure is uniquely selling out before it's built — this isn't speculative hype but booked demand straining global power grids.

  • •Cerebras is sitting on a $25 billion backlog and must deploy at a 'blistering pace,' with demand far outstripping the ability to build and fill data centers, per Andrew Feldman.↗↗
    quote
    “The irony is, unlike many sort of exciting times in technology, They're trying to capture yesterday's demand, right? The demand is way outstripping our ability to build data centers and to fill them with hardware. All right? And so we have a $25 billion backlog.”
  • •Unlike past tech booms, providers aren't chasing 'if you build it they will come' — demand is already booked and the scramble is to keep customers from leaving, Feldman argues.↗
    quote
    “And we are not alone in that. OpenAI, Anthropic, you go through this list of— Google wants more data centers. Microsoft wants more data centers. AWS wants more data centers. All of these players are not chasing sort of, if you build it, they will come. They're chasing the demand is booked.”
  • •Data center buildout has gone global, reaching non-traditional markets like Kazakhstan, Tajikistan, Georgia, and Armenia alongside the US, Nordics, and Middle East.↗
    quote
    “Right. And what we're talking about now are data centers. That are in the next several years going to use more power than the previous 50 years on Earth took. Wow. Right. We're talking about individual buildings the size of football fields that have more power coming into them than mid-sized cities. And they're being built. They're being built across the US. They're being built in Canada. They're being built throughout the Nordics. They're being built here in Paris and throughout France and Europe, in the Middle East. In nations that sort of weren't front and center in anybody's mind previously. You know, Kazakhstan, Tajikistan are building out, Georgia are building out data centers of size, Armenia. Everybody's sort of focused.”
  • •Individual AI buildings are now the size of football fields drawing more power than mid-sized cities, and Feldman claims data centers built in the next several years will consume more power than the entire Earth used in the prior 50 years.↗↗
    quote
    “Right. And what we're talking about now are data centers. That are in the next several years going to use more power than the previous 50 years on Earth took. Wow. Right. We're talking about individual buildings the size of football fields that have more power coming into them than mid-sized cities. And they're being built. They're being built across the US. They're being built in Canada. They're being built throughout the Nordics. They're being built here in Paris and throughout France and Europe, in the Middle East. In nations that sort of weren't front and center in anybody's mind previously. You know, Kazakhstan, Tajikistan are building out, Georgia are building out data centers of size, Armenia. Everybody's sort of focused.”

Breaking Moore's Law with New Silicon

Cerebras argues its fresh architecture has escaped the GPU's aging trajectory — critical because fast inference is now the bottleneck for reasoning AI.

  • •Feldman claims Cerebras 'crushed' Moore's Law and expects to exceed 2x performance gains in the next 18 months, far outpacing the traditional doubling cadence.↗
    quote
    “Doubling about every 18 months. Got it. And we crushed it with this chip. And we've carved out a whole new trajectory. And my view is in the next 18 months, we'll be way over 2X.”
  • •Early-stage architectures have far more optimization headroom, while 20-year-old GPU designs must rely mainly on fab node shrinks for gains, per Feldman.↗
    quote
    “And so, uh, I, I think that, uh, early in an architecture you have room to, to do much better than what was traditionally Moore's Law. Now, if you've got a 20-year-old architecture like the GPU, it's much harder, right? You, you have to rely on things like smaller geometry, right? Going to the next fab node. But in a newer architecture, you have a huge amount of room still to learn about the work that is being presented and make optimizations that give you huge gains.”
  • •Reasoning is computationally intensive inference; fast chips make it tractable rather than crippling it with long wait times — a single Cerebras chip costs roughly half a billion dollars to make.↗↗
    quote
    “Right. And so fast compute makes this sort of work fast and sort of tractable. It doesn't cripple it by taking a huge amount of time to get a good answer. And so it's exactly the fact that this reasoning consumes a huge amount of tokens internally. That allows a blisteringly fast machine like ours. And I brought one. Oh, you got— I'm never far without, you know, when one costs half a billion to make, you bring it everywhere with you.”
  • •Companies build their own chips less to outperform vendors than to control their destiny, echoing hyperscalers' lessons from over-dependence on Intel x86, Feldman notes.↗
    quote
    “No, I think nobody likes being dependent. And I think some of the lessons learned by the hyperscalers of the x86 world is they were dependent on Intel. And some of the lessons learned by the GPU makers was they were dependent on a small number of hyperscalers.”

Reasoning, Recursion, and the 'AGI Is Already Here' Case

Feldman argues unlimited tokens equal unlimited reasoning, and recursive loops are producing exponential gains that already clear old AGI benchmarks.

  • •Feldman contends AI has already surpassed any 20-year-old definition of AGI, having blown past the Turing test.↗
    quote
    “by any definition we had 20 years ago, we've hit it. Right. I mean, if you think about, oh, there's a Turing test, blew it away.”
  • •Unlimited tokens means unlimited reasoning — running reasoning models 24-48 hours at 15x Cerebras speed could compress weeks or months of thinking into a single day.↗↗
    quote
    “It does. What does that mean? Yeah, I mean, if you run these for 25 or 48 hours, you get amazing things now. And what if by using Cerebras, we were 15 times faster, and then you ran it for 24 24 hours, right? And you've got weeks or months worth of thinking.”
  • •Recursive learning loops produce exponential, not linear, improvements with no clear ceiling yet, and AI compresses inter-generational learning that once took 20-40 years into thousands of equivalent generations.↗↗
    quote
    “And that we're just beginning to see that now. You ask it a question, you learn from the results, you ask it to do it again. The results get better and more information's added. Your answer gets better. You ask it to do again, it covers more material. And these sort of loops are producing sort of not a little bit better answers, but vastly better answers.”
  • •Models like GPT-4o and o3 are shifting from needing precise prompts to understanding user intent, and tools are beginning to build themselves in a self-reinforcing recursive loop, per Calacanis.↗↗
    quote
    “Right. And if you have a chance to play with 4o or o3 from OpenAI, increasingly What you don't have to get the prompt just right. You don't have to be a prompt whisperer. Instead, you ask it and it says, well, here's some things. And by the way, maybe you wanted the chart to go two ways. You wanted it aligned in a bar. And it's like, well, that's exactly what I wanted. I didn't ask for it, but that is better. And so it's understanding intent.”
  • •The bottleneck for superintelligence may shift from intellectual problems to human coordination — organizing people to execute what AI prescribes, Feldman suggests.↗
    quote
    “And well, that's right. And when are the problems no longer sort of intellectual problems and they're now people problems?”
  • •Feldman frames the upside boldly: AI offers a realistic shot at ending cancer deaths and enabling personalized tutoring agents — a superior teaching method known for 1,000 years but never scaled.↗↗
    quote
    “Yeah, we have a shot with this technology so not our children nor anyone they know dies of cancer. I mean, say it like that. There will be some dislocation in the economy. Sure, there will be. There was dislocation when cars came, and it was a bad deal to be a guy who shooed horses or built carriages. But you got to also, against that, make your tea of the cons and the pros. There's a shot that our children, none of them nor the people they love, will die of cancer. And that's one thing that we can work on with this technology and we will have great purchase on. And I think you begin listing those and then it's a more thoughtful discussion.”

Open Source Wins and the Bifurcated AI Market

Routine enterprise work runs fine on open models while only hard problems need frontier labs — but the US faces an open-source sovereignty gap versus China.

  • •Open-source models handle the large volume of routine enterprise tasks (G&A, data processing) while frontier closed models from OpenAI, Anthropic, and Gemini are reserved for genuinely hard problems, Feldman argues.↗
    quote
    “Right, there's minivan time. And I think that as the sort of sophistication of the user grows, right, you're gonna have hard problems and those are gonna be frontier model problems. They're gonna be OpenAI problems, they're gonna be Anthropic problems, they're gonna be Gemini problems. And behind that, they're gonna be a lot of ordinary problems, right? I mean, if you think about a company, you know how much time is spent cutting things out of Workday and getting it in a different cell for, yeah, right?”
  • •Enterprises are maturing from unconstrained token usage to strategic allocation — heavy compute for high-value users, cheaper open-source models elsewhere.↗
    quote
    “Strategic. And it's exactly the same. I think at first people opened up and said, everybody, as much tokens as you want. And in enterprises, there's no open loop. We don't give sort of any resource unconstrained to people. And now we're jumping on and saying, whoa, all right, these guys should have as much as they need. They're enormously productive over here. We can use maybe an open source model, maybe a cheaper model over here.”
  • •The US needs more domestic open-source models; today the only options are OpenAI's OSS-120B or Chinese models, creating a sovereignty gap, per Feldman.↗
    quote
    “We are seeing that for sure. And I think OpenAI made a good call releasing OSS 120B. Some months back. That was a good open source model. But I think in the US, we need more domestic open source models. We need to give the world a choice. Right. If they want to run open source right now, it's OSS-120B or Chinese models.”
  • •Regulated industries and foreign governments are demanding on-prem, open-source deployments for control and to avoid data leakage — Cerebras runs a diverse set including GLM, KIMI, Qwen, GSK's proprietary models, and UAE partners G42/MBZUAI.↗↗
    quote
    “We need to have this on-prem, domestically, and we'd like an open source version where we have a little bit more control.”
  • •Nvidia is reluctant to champion its own open-source models because doing so would put it in direct competition with its largest customers, per Calacanis; Feldman adds fierce lab competition (Anthropic vs. OpenAI) is net positive.↗↗
    quote
    “that was— you cut me off at the pass. My understanding was Jensen was like, hey, we don't even want to talk about these open source models we have because our customers, right? We're now going to be competing with Sam, Dario, Elon, Sergey. Like, do we want to be in that position?”

AI as a Cybersecurity Double Threat

The same models that expose critical vulnerabilities in hours also add safety-guardrail latency that faster chips can mask — and a massive breach is inevitable.

  • •An AI model found tens of critical vulnerabilities in Palo Alto Networks' software within an hour; CEO Nikesh Arora had to halt everything for 6 weeks of emergency patching.↗↗
    quote
    “Right. And that's when you know, right? I mean, Nikesh leads maybe the leading security software firm. Right. And when it finds in an hour, right, tens of critical opens, you're like, whoa, this is a powerful tool that we need to think. And maybe you show it to a group first, right? Maybe you— I don't know what the right thing is, but—”
  • •Cerebras discovered in the last 6 weeks that AI safety guardrails add measurable latency, and fast chips reduce that penalty, making safety less costly to implement.↗↗
    quote
    “The guardrails have an impact. One of the things that's Fast does is it makes the guardrails less painful. And so that is— we discovered that in the last 6 weeks. Yeah. Is that the very guardrails can add time and make it feel slower. And so fast chips like ours can really help that. But so they're racing against competition. They're racing against their own sense of greatness.”
  • •Feldman warns a massive AI-related data breach is not a question of if but when, invoking Warren Buffett's reinsurance logic — prepare response plans in advance.↗
    quote
    “Right. And it's like Warren Buffett talked about the reinsurance industry, that you know something bad is going to happen. You don't know when. Yeah, but you gotta save up for it, right? You put money away for insurance, but there will be a tornado, there will be a massive earthquake. I mean, we, we know this and we can do our best to plan, but there'll be a massive breach and there'll be— and we have to steel ourselves in advance and we have to think about it, think about the right response at the time and sort of prepare ourselves for a future that is in specific unknown, but in general we're pretty sure it's good, something's going to happen.”

Generative Media Reinvents Hollywood Economics

Black Forest Labs' Flux models are collapsing production costs and shifting AI's role toward human-in-the-loop creative iteration, not push-button movies.

  • •Black Forest Labs' co-founders invented Latent Diffusion — the fundamental algorithm behind all deployed image, video, and physical AI generation — by compressing natural data into efficient latent representations, the same principle as JPEG and MP3.↗↗
    quote
    “I'm splitting my time to a certain degree. We started a company 2 years ago. Me and my co-founders, as you said, like we've worked on Stable Diffusion in the past. Before that, we invented like an algorithm called Latent Diffusion. Which is basically like the fundamental algorithm behind all of like generative models that are being deployed for image generation, video generation, even like physical AI now.”
  • •A Bitcoin movie was filmed on a soundstage using Flux for all scenery instead of green screens, cutting a $150 million traditional budget to $30 million; startup launch videos that cost $100K-$250K now take a week or two.↗↗↗
    quote
    “I was talking to her at an event and she was telling me, um, it was the Breakthrough Prize Uh, Yuri Milner's event, and she was telling me she just did a Bitcoin movie, and they did it on a soundstage without green screens, but all the actors just worked in like a soundstage, and then all of the scenery behind them was being done by generative AI. That's a real movie. That's a $30 million budget movie. She said it would have cost $150 million if they had to build sets, and the film would have never been greenlit. Are you starting to see people use that in production, not just in the backend and the ideation phase, but actually in production yet with your tools?”
  • •Director Martin Scorsese used Black Forest Labs' models to visualize and iterate on scenery to communicate creative vision, exemplifying AI as pre-production storyboarding akin to Ridley Scott and Spielberg's sketches.↗↗
    quote
    “Um, I think it was really this idea of like, he has clearly, um, a vision in his head of like a scene or a scenery where like maybe a new movie, um, will be shot. And he's trying to explore that and kind of like, we basically looked at the scenery of like a village in Eastern Europe somewhere and he was describing it. We saw some outputs, we iterated on the outputs. And ultimately, I think, and that's what he said in the end, is like getting like the mental picture of something out of your head and communicating it in a visual way by making like these images or the series of images. Is something that just makes it easier to communicate and convey an idea of what is actually in your head. And I think that's one of the very interesting and powerful ways to use this technology. And I think ultimately—”
  • •Speaker C argues the most valuable use is human-in-the-loop creative iteration, not a fully automated pipeline to produce complete movies.↗
    quote
    “It's also interpreted in different ways, but then visual information is so rich, so rich, like an image or video, there's so much signal in it. And it's just like another way of communicating. And I think that's like one of the beautiful things that this technology ultimately enables. And I think like to your question of making like full movies with, I don't know, like a video generation model, for example, I'm not sure if that is like the ultimate goal. Maybe it's like interesting to plug this into like some kind of a gigantic workflow and make like a very long video. And I think that's really cool to explore. But I think ultimately, like the real interesting use cases, they come when you have like a human in the loop who iterates and uses it as a medium. And I think this is at least like a perspective that I take that makes it interesting. And that this is most often when the most interesting outputs arrive or are actually being made.”
  • •Black Forest Labs builds custom models with IP holders as a commercial offering; Calacanis sees a future where IP owners license universes to AI-empowered fans, citing 'Star Wars Stories Untold' getting millions of views (OpenAI's Sora-Disney deal has since lapsed).↗↗↗
    quote
    “I think it is, look, I think like the most interesting use cases of this, like if you think about like content creation is in generating something, making something that hasn't been there before, right? Like that's a fundamental, like interesting aspect of this technology. And then I think like, yeah, when it comes to IP, what we implement, for example, on like our public-facing tools is You cannot generate certain IP with these models, right? And I think that's something that is a sensible approach. And then yes, we do work with certain IP holders to develop models together with them. Some of them based on our open source models, some of them based on our more powerful proprietary models. But I think that is a very attractive value proposition.”
  • •Founded roughly two years ago by former Stable Diffusion creators, Black Forest Labs has crossed 100 employees and is hiring in Freiburg, Germany and San Francisco; quality has leapt from 64x64 pixel images to multi-minute high-resolution video.↗↗↗
    quote
    “I'm splitting my time to a certain degree. We started a company 2 years ago. Me and my co-founders, as you said, like we've worked on Stable Diffusion in the past. Before that, we invented like an algorithm called Latent Diffusion. Which is basically like the fundamental algorithm behind all of like generative models that are being deployed for image generation, video generation, even like physical AI now.”

One Model to Rule Them All: Multimodal Convergence and Robotics

Speaker C's boldest thesis is that the same generative model making movies can become a robot's brain, unifying digital and physical AI.

  • •The convergence of image, video, audio, and action prediction into a single multimodal model is a new paradigm deployable on real-world robots, per Speaker C.↗↗
    quote
    “It basically makes use of this principle that you can compress natural data such as images, such as video, such as audio into a much more like efficient representation and then train a transformer model on that. And I mean, this is the stuff why like, you know, like JPEG, MP3 and all of that works. And we basically translated that into like a neural algorithm a few years ago when we were still like PhD students in Munich actually. And then built on top of that, we built Stable Diffusion. And then on top of that, yeah, the generative models that we are developing today. And of course, like the technology has advanced. But we are now tackling, I would say, models that are really made for understanding the whole world around us. Multimodal visual models pre-trained on images, videos, audio data at the same time. And we are now entering a new paradigm, which is combining that with something that's called action prediction, such that you can actually use the same model to make images, to make videos, to make audio, and to predict actions, which means you can ultimately deploy it on a robot in the real world.”
  • •Pretraining on video gives implicit understanding of real-world physics, enabling action prediction and robotics from the same model.↗
    quote
    “I think that's, yeah, I think that's like a really good way to think about it. It's like an intuitive way to interact with the world, right? Like, I would say there's like these complementary forms of intelligence. Ultimately, there's like intuitive intelligence and then there's like a deep reasoning layer. Now, ultimately, you need for like a kind of like complete form, you need both and you need them to interact. And I think like we've been approaching it more from like the intuitive side. Images is like a very natural way to approach this whole field because it's not as computationally intensive as let's say video, right? But now, yeah, I think like we're combining it. It's converging into like a multimodal model. And yeah, we see like exactly like pretraining on videos gives like implicit understanding of the physics of interactions with the real world. And then you can get stuff like action prediction, like robotics out of the same model.”
  • •World models, world-action models, and computer use are all converging on the same underlying architecture, making generative models a unified platform across domains.↗
    quote
    “Production workflow, right? But I think when I look at multimodal generative models as a whole, I think what really excites me is you can use the same kind of AI model to make a movie and deploy that as a brain on a robot. And I think this is so interesting. And I don't know, there's some thoughts around trying that in the digital world, right? Which would be, for example, computer use. Remains to be seen if that is actually something that works or not. But I think the technology is so powerful and so versatile. And it's just moving into that. In all the talk on world models, world action models, all of that, it's basically all the same. And I think that's what's making it so interesting and what I find most exciting.”
  • •Robots today need only a few hours of fine-tuning on top of broad visual understanding to adapt to specific hardware and tasks, with in-context prompting of robots as the key research frontier to watch.↗↗
    quote
    “I mean, ultimately, I think you would want to go to a place where you could prompt a robot in context, right? As you can do with a language model, basically just tell it, hey, go and, I don't know, pick up this glass with the, I don't know, orange juice or whatever. Make a cocktail. Yeah, exactly. We're not there yet, but I think this is one of the goals. And I think how these models are deployed currently is there's a lot of different hardware, different robots that are running in factories that all have some different kind of action representation that you need to kind of tune the models towards, right? So in practice, what you do is you have all this visual understanding in the models and then you need only a very little bit of a few hours of fine-tuning data to adjust the model on that specific task. And I think the goal would be to kind of move away from that towards as much in-context as possible. But it is a little bit of a research problem. I think that open source.”

Stock read-through

CerebrasBullishCEO touts $25B backlog with demand outstripping supply, claims to have broken Moore's Law with expected >2x gains in 18 months, and runs diverse models enabling faster reasoning.↗
OpenAIMentionedDiscussed as a frontier closed model provider whose OSS-120B is a rare US open-source option; its models run on Cerebras and it competes healthily with Anthropic.↗
Black Forest LabsBullishMaker of Flux/latent diffusion, crossed 100 employees, hiring, and partnered with Martin Scorsese to visualize film scenes.↗
AnthropicMentionedCited as a frontier model provider for hard problems and as a healthy competitor to OpenAI benefiting the ecosystem.↗
GoogleMentionedReferenced among frontier model providers and as an example of competition sharpening incumbents.↗
Palo Alto NetworksMentionedCEO Nikesh Arora found tens of critical vulnerabilities via an AI model in an hour, forcing six weeks of emergency patching.↗
NvidiaMentionedJensen reportedly reluctant to promote its own open-source models to avoid competing with major customers like OpenAI and Anthropic.↗
DisneyMentionedDiscussed re AI-generated fan films licensing IP and a lapsed OpenAI/Sora character licensing relationship.↗
GlaxoSmithKlineMentionedRuns its own proprietary models on Cerebras infrastructure.↗