(Preview) Nvidia’s Answer to Capital Constraints, Google’s Attrition and Direction, Q&A on AI Writing, Vision Pro, Vibe Coding
- Ben frames the binding constraint on the AI buildout as money itself, after compute and power. He leans on the railroad analogy because the fundamental issue in 1873 "is the world ran out of money" — a year ago bubble talk was dismissed since hyperscalers paid out of free cash flow, but "we sort of blew through debt in like nine months," with debt raised in the second half of last year and first half of this year "probably soon to be approaching, like, a trillion dollars" as issuers' balance sheets get "sketchier and sketchier."
- NVIDIA's new $500B+ financing platform with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR is a pitch to reclassify AI as patient-capital infrastructure. Jensen Huang's case is that "you're all thinking about AI wrong" — GPUs run longer than you think and CUDA improves them over time, so the asset class fits pension-fund money that classically buys toll roads. Ben's zoom-out: the case is being made now "because all the short-run capital's been used up."
- The A100 proof point — six-year-old, fully depreciated chips that CoreWeave says are contracting at higher rates than before — doesn't prove what Jensen wants it to. Ben's mechanism: the shift to water cooling means GB200s and the upcoming Vera Rubin can't slot into old air-cooled data centers, so those facilities may be stranded and A100s may persist because "there's no replacement for them." In abundance, "the old compute's gonna get retired very quickly."
- Today's scarcity reflects pre-2024 decisions, and correlated signals are how boom-bust cycles happen. Everyone sees demand exceeding supply at the same time: "It might be a 5X signal, but 10 companies invest, so you end up with double the capacity that you need" — so the current supply-demand environment is not representative of two, three, or thirty years out.
- Ben believes the infrastructure can pay off, but is bearish on certainty of timing. Unlike a railroad, AI can have digital-good scalability — "none of that stuff quite works now... it's working pretty well, and it's accelerating unbelievably rapidly" — but "it's not enough to be right, it's about timing," and the risk is "an air pocket where we run out of money" before the spend cycles back as profit.
- The micro story: LLMs helped send NVIDIA's stock to the moon while diminishing its CUDA moat. The developer platform shifted far above CUDA — "no one who's writing an AI application today is using CUDA"; apps sit on OpenAI/Anthropic APIs or Bedrock-on-Trainium, fully abstracted from chips — so Ben agrees with Andrew Sharp's read that the financing platform is partly defensive as cost-sensitive customers push toward Google TPUs.
1. The real question isn't compute or power — it's what happens when you run out of money
- This is a mailbag episode, opening with an emailer (Andrew, not the DC one) asking whether NVIDIA's newly announced financing platform — Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR mobilizing "over $500 billion of third-party capital" — echoes the re-securitization of mortgages into CDOs that set up the subprime crisis.
- Ben's macro framing runs through the railroad analogy he "reluctantly" linked (even Satya Nadella cited it): the fundamental issue in 1873 "is the world ran out of money." A year ago you couldn't call this a bubble when companies paid from free cash flow — "once we start getting into debt, then we need to have a conversation." Then: "we sort of blew through debt in like nine months," a sum "probably soon to be approaching, like, a trillion dollars," raised by great businesses whose balance sheets are "getting sketchier and sketchier."
- The hosts' victory lap — they'd predicted in Madison a week earlier that lending would tighten. "Good job by us."
2. Jensen's pitch: AI is long-run infrastructure that deserves long-run capital
- The untapped pool is patient capital — pension funds whose classic investment is a toll road (Ben detours through the "doctor plan," the catch-up pension structure for late-starting high earners, to explain the mechanics). Pensions in theory would have been a good match for railroads: money that must exist long-term but pays out slowly.
- Huang's post argues "you're all thinking about AI wrong" — it's a long-term investment: GPUs run longer than you think, CUDA makes them better over time, and data shells are 30-year assets. Ben's zoom-out: "it's like, yeah, because all the short-run capital's been used up."
- Ben sees a "beautiful symmetry" with Google's equity issuance, which he'd compared to Berkshire using high-margin See's Candies cash flow to buy BNSF — a lower-margin business throwing off high absolute, predictable cash.
3. The A100 evidence is real but not representative
- Ben flags the choreography as no accident: Huang makes the case, then CoreWeave's earnings tout A100s — a six-year-old, previous-generation chip — "contracting out at a higher rate than before," fully depreciated, pure profit. On the surface, "a pretty good argument. There's just a couple problems."
- Problem one: the shift to water cooling means GB200s and the upcoming Vera Rubin can't slot into old passively-cooled data centers (the H generation may have been air-cooled or half-and-half). Those facilities may be stranded, so A100s may stay in place because "there's no replacement for them" — fine in a compute-scarce world, but "not representative of what your expectations should be for GPUs going forward."
- Problem two: today's supply reflects 2024-and-earlier decisions (two-year lead times), when markets were "freaking out about CapEx" — and the spenders were wrong only in not spending more. But when everyone gets the same signal: "It might be a 5X signal, but 10 companies invest, so you end up with double the capacity that you need." If GB200s become abundant, "the old compute's gonna get retired very quickly."
4. The bull case isn't insane — but being right isn't enough
- Ben's distinction from the railroad: there's no way to accelerate a railroad's revenue — brutal terrain, land development, finite trains — whereas AI can have digital-good scalability, especially in the possibility of AI writing its own programs or being "set loose on a company" to create agents. "None of that stuff quite works now, but... it's working pretty well, and it's accelerating unbelievably rapidly." He mentions pushback from an emailer calling him a Luddite for taking six months to vibe code, and says the criticism is kind of valid.
- The unavoidable math: more supply depresses prices; the bet is demand accelerates even faster. And debt can't fund things forever — "at some point you need to actually make money." Ben's bottom line: "I believe this stuff will pay for itself. The question is will it pay for itself in time to avoid, like, an air pocket where we run out of money?" Andrew's translation: "a whole bunch of bag holders."
5. The micro story: LLMs shifted the platform above CUDA, and this deal is partly defense
- Andrew's read — which Ben endorses as "the NVIDIA-specific question" — is that as everyone gets cost-sensitive and Google brings TPU infrastructure online, NVIDIA wants to encourage buildouts using NVIDIA hardware and software.
- Ben's history: pre-ChatGPT GTCs threw every parallel-computing library at the wall; he recalls GTC 2024 as oddly boring because LLMs, while sending the stock to the moon, "were bad for NVIDIA" — the developer platform moved far above where NVIDIA sits. "No one who's writing an AI application today is using CUDA"; apps run on OpenAI or Anthropic APIs, or Bedrock on Trainium with a Chinese open-source model, "totally abstracted away." The moat "has been tremendously diminished."
- The earned-it caveat Ben insists on: NVIDIA almost went under building CUDA when nobody understood why, bottoming out as recently as October 2022 — three weeks before ChatGPT, when Ben wrote "NVIDIA in the Valley." "They have earned every dollar they've gotten through 25 years of taking massive risks."
Full transcript
Hello, and welcome to a free preview of Sharp Tech. Hello and welcome back to another episode of Sharp Tech. I'm Andrew Sharp, and on the other line, Ben Thompson.
Ben, how are you doing?
I'm doing okay, Andrew. I feel a little bit in a funk. There's been some travel going on. It's kind of dreary outside. The Brewers are terrible. I'm trying to figure out what is causing what.
But it's okay.
But here we go.
We'll make it happen. That's right.
You know what I feel? I feel FOMO because we were together in Wisconsin last week, and I feel like we could have put a call in to Mark Walter to see whether he was interested in selling the Lakers.
To us.
Sounded like an asset he needed to move pretty quickly.
Yeah.
Maybe we would've gotten lucky, could've beaten Kushner to the punch. Alas, here we are, humble podcasters once again.
Well, the big question then—not to dive into a totally random aside—but Josh Kushner—
Mm-hmm.
Not Jared—Josh Kushner is now one of the owners of the Los Angeles Lakers. Thrive Capital is kind of on the cutting edge—the new generation of VC companies. They're doing very well for themselves.
Sure.
I do think their largest holding is OpenAI, so maybe the real bubble concern now is whether anything happens to the Los Angeles Lakers if everything goes sideways.
Well—
We'll have to keep an eye on it.
God willing, that would be one benefit of the bubble bursting, so let's see what happens. For now, Ben, we're gonna do all mail on this episode, and I'll tell you why: the last 2 episodes we've recorded, we've gotten so deep into various conversations that we've hit hardly any mail. So we'll try to remedy that today, and we'll start with an article you wrote this week.
Are you telling me I need to not monologue so much?
That's right.
Keep it short.
Be on your P's and Q's.
Keep it super short.
Let's hit as many of these questions as we can.
Yeah, we'll see how it goes.
We'll see. NVIDIA announced partnerships this week with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR—a super team—to establish an independent financing platform designed to mobilize over $500 billion of third-party capital to support the build-out of AI infrastructure over time. That's NVIDIA's announcement.
Part of that plan, as I understand it, involves shifting GPU depreciation risk away from traditional lenders in a bit of innovative financial engineering that I hope you can explain for me, because I'm still a little confused about what the plan is there.
Hey, this is American greatness at play. We can invent a very expensive thing to spend money on and invent incredibly convoluted ways to pay for it.
Sure. Great.
Exactly.
God bless America. So, Andrew, in response to the article you wrote about this on Tuesday—
Wait, is this Andrew in Washington, DC? Just to clarify.
This is a different Andrew, although—
Okay.
Look, this Andrew also has lots of questions about what this actually entails. Andrew asks, “Can CUDA really generate earnings growth at a rate that outpaces depreciation of the GPUs? I'm being very unscientific about this, but it feels to me like there's an order-of-magnitude difference in there, and not in CUDA's favor.
“The conclusion of your article on Tuesday carries echoes for me of the re-securitization of mortgage instruments into CDOs and credit default swaps that created the conditions for the subprime loan crisis and the global financial crisis. Do you see any parallels?
“In seeking to expand the breadth of available capital, is Huang creating the preconditions for a subsequent cascading collapse? Perhaps more interestingly, is there a feasible alternative, or is this just the way the bubble expands?”
So what do you think, Ben? Take it in whatever direction you prefer.
Well, if you let me take it in whatever direction I prefer, we may look up an hour later and not have gotten very far toward the mailbag. There's a macro question about AI infrastructure generally, and then there's a micro question about NVIDIA specifically. Both are at play in what happened this week.
Okay.
1. AI Runs Out Of Money
At a very high level, this is where people reach for the railroad analogy. I sort of reluctantly link to it. It's such a good analogy this week, but only because everyone's talking about it.
Mm-hmm.
So I had to cite it: even Satya Nadella brought this up on his call. I'm not anything special here.
Everyone's reading the same book.
That was just an acknowledgment—
Yep.
That this is not—well, not just that, but people have been talking about the railroad thing for a few years now. The book just came out this year, which is fuel on the railroad analogy fire.
But where the railroad point is interesting is that the fundamental issue in 1873 was that the world ran out of money.
Mm-hmm.
We've talked about running out of compute, and we've talked about running out of power, but the issue at hand here is what happens when you run out of money? That sounds like an incredible thing to say, given how much money there is in the world, but we talked on this podcast even a year ago—not that long ago—about how you can't really call it a bubble when these companies are paying for this out of their free cash flow, right? What's the spillover—
Sure.
That we're worried about?
What's the risk they're assuming in that scenario?
That's right. It's like once we start getting into debt, then we need to have a conversation. The crazy thing is we blew through debt in 9 months. The amount of debt that was raised in the second half of last year and the first half of this year is in the hundreds of millions, probably soon to be approaching a trillion dollars.
Mm-hmm.
It was raised by big companies with great balance sheets, or great businesses, I should say. The balance sheets are getting sketchier and sketchier.
Money-printing businesses. So they're real—
Right.
Businesses.
At some point, you run out of people willing to give you money.
Totally.
And—
I mean, we talked about this a week ago in Madison, where we were discussing how the lending environment will tighten, and hyperscalers—
Good job by us.
Yeah.
Yeah, good job by us, foreshadowing this announcement. But there's still lots of money out there.
Mm-hmm.
There is money that traditionally goes to large, long-running infrastructure projects because that money itself is a long-term liability. The classic example here is the pension fund.
Mm-hmm.
You're paying into your pension over time. Your employer is paying into your pension over time. I actually know a surprising amount about the mechanics of this because, for one-person businesses, pensions are actually the best possible retirement plan.
Ah.
For one-person businesses, you could contribute a much greater amount than with a traditional retirement plan before taxes, and shift your tax liability window—all these things that go into it. It's actually called the doctor plan because doctors are the most frequent users.
Okay, yeah.
What happens with a doctor is that you're in school for a very long time, so you start making money relatively late. But once you make money, you usually make a fairly decent amount of money. So it's a catch-up plan where you can put way more money into retirement—
All the money—
That's right.
You weren't saving in your late 20s as you were toiling through school and residency.
That's right.
Okay.
It's interesting because it's a hangover from old-school pension plans that aren't really in favor anymore. But that's money that has to be there in the long run, but it doesn't have to be paid out for quite a while.
Mm-hmm.
These are the sorts of investments that money wants to go into. A toll road is the classic pension investment, where you're putting a lot of money to work, but the predictability and understandability of the long-term payback is very clear.
Yeah.
And it’s going to pay back over a very long time, and you’re going to make a lot of money in the long run, but you have to have very patient capital because pensions, in theory, would’ve been a good match for, say, railroads, right?
Mm-hmm.
Because the problem with the railroad is you build it, and you might not really get your money back for 30 years. And this is the beautiful symmetry, because I think this NVIDIA deal is symmetric with the Google equity issuance, in which I wrote about Berkshire Hathaway and its shift from See’s Candies, a very high-margin business, using that cash flow to get into BNSF Railway, which is a lower-margin business, but the absolute—
Stable.
—the cash that’s thrown off—
Predictable.
—is very high.
Yeah.
Right. And the analogy there is, to what extent is Google making the same shift? I think that’s a very pertinent point to this NVIDIA thing, which we can circle back around to. So you have this long, patient capital that is a very good alignment for long-running investments.
Mm-hmm.
2. Jensen Recasts AI As Infrastructure
And what you had in this post by Jensen Huang is him trying to make the case that you’re all thinking about AI wrong.
Hmm.
It’s not a short-term investment. It’s actually a long-term investment. And if you put NVIDIA GPUs in, they run for a very long time—longer than you think—and we make them better with CUDA over time.
This is sort of building on the hyperscalers’ argument, which is, look, the data shells, the actual buildings, are 30-year investments. We’re only buying GPUs right when we need them, so they’re kind of aligned but a little misaligned in that regard. The case being made here is that this is a long-run investment that deserves long-run capital.
If you zoom out, it’s like, yeah, because all the short-run capital has been used up. That’s sort of the case being made here. Now, is the case valid?
Yeah.
Is—
Well—
—that sort of the next question—
—the lenders’ concern—
Sorry, Ben in Madison wants to email and say—
Do we buy it?
“Hi, guys. Is this case valid?” Yeah.
Well, in terms of the invalidity, or potential invalidity, one of the concerns is that the GPUs that any of these companies—any of these infrastructure companies—are buying from NVIDIA burn out before the patient capital can realize the upside.
Or not just that, but NVIDIA comes out with new GPUs—
Right.
—that make your own GPUs obsolete.
They’re obsoleted. Exactly.
Right?
And so—
So—
—NVIDIA’s trying to guard against that risk, correct, and try to allay some of those concerns?
Yeah. NVIDIA’s trying to do a lot of things, most importantly preserving its competitive position and margins.
It’s kind of an interesting point, a big talking point that Jensen Huang raised, and that was repeated on the CoreWeave earnings call. I don’t think it was an accident that these happened back-to-back. Jensen Huang comes out and makes this case. Then CoreWeave comes out and says in its earnings, “We have A100 chips that we are contracting out at a higher rate than before.”
And they’re working great. Yep.
I think that’s absolutely believable. It better be—they said it in their earnings, right?
Yeah.
It makes sense. Compute is in such demand. There’s already installed compute, even if that compute is 6 years old. I think the A100 hit, you know, in 2000—
Yeah, it’s a previous generation, for anybody who’s not clear.
Right.
But it’s still being utilized.
So on the surface, it’s a great case. It’s like, look, people are out there saying GPUs only last 2 to 3 years. Actually, here’s an example of a chip that is 6 years old signing contracts right now.
Still comes with demand, yep.
Those contracts are worth more than what the contracts were previously. They’re actually increasing in value. And by the way, these are fully depreciated assets. All the cash they’re earning is pure profit. This is a long-term asset.
Hmm.
On the surface, it’s a pretty good argument. There are just a couple of problems.
Okay.
3. Water Cooling Strands Old GPUs
Problem number one: A big shift that has happened in the last couple of generations has been a shift to water cooling, which requires entirely new kinds of data centers. You can’t just take your GB200s or the upcoming Vera Rubin and slot them into the old data center.
Mm-hmm.
They actually need water cooling, and this requires entirely new ways of putting servers together. Facebook had this whole open-source, open-data-center thing. It had this concept that it could manufacture data centers very rapidly. It was this 2-story sort of thing—I think it was 2 stories, or whatever—but it all depended on passive cooling.
Okay.
So one question I have about the A100 case—and I think the H100 generation might also be air-cooled, not water-cooled, or maybe it was half and half—is whether the reason those are staying in place is because there’s no replacement for them.
Hmm.
You have data centers that are built around a particular assumption about cooling. New GPUs don’t fit that assumption, so that data center is actually stranded.
They’re stuck—
So, sure—
—with the A100s for life because of the way—
Yeah.
—the data center was built.
That’s right. So on one hand, in a compute-scarce environment, absolutely, they can keep selling them.
Mm-hmm.
4. Compute Demand Could Overshoot
But the A100 is not representative of what your expectations should be for GPUs going forward.
And it’s not necessarily—
Because—
—dispositive as to the question of whether this will still have utility—
That’s right.
—in a market.
This doesn’t undo it. The fact of the matter is that A100s are being sold for more than they were before because compute is scarce. But that gets to the next question: The available compute today is a function of decisions that were made in 2024 and before.
Mm-hmm.
Right? It takes about 2 years to bring these online. And, of course, back then, the market was, for the record, freaking out about CapEx.
Yeah.
Everyone who spent money on CapEx was right. Actually, no, they were wrong. They were wrong because they didn’t spend enough on CapEx. They should have spent more in 2024.
But everyone has these signals. Everyone’s talking about how demand exceeds supply, but everyone’s getting the same signal at the same time. A reasonable concern from the market is, okay, if one company was getting this signal, then yes, it can invest appropriately. If 10 companies are getting this signal and they invest, do we overshoot?
Hmm.
This is how the boom-bust cycle happens: Everyone’s getting the same signal. That doesn’t mean the signal is a 10× signal. It might be a 5× signal, but 10 companies invest, so you end up with double the capacity that you need.
Right.
That’s another concern: The supply-and-demand environment right now is not necessarily representative of the supply-and-demand environment in 2 years, 3 years, or 30 years—however long you want these long-lived assets to be considered over.
And that scenario—
Yeah.
—would involve several companies bowing out of some of these infrastructure build-outs and the race to the frontier. Is that right?
What it would entail is that your A100s are not going to be getting contracts if there are a gazillion GB200s available.
Hmm.
Right?
Okay. Yeah.
They’re available as a function of there not being compute. If there’s an abundant amount of compute, the old compute is going to get retired very quickly.
Yeah.
So again, I’m not saying the argument being put forward is wrong. There’s a lot of weight being hung on these A100 contracts that I’m just saying are not necessarily going to be representative in the long term.
Hmm.
5. AI Demand Faces A Timing Test
The pushback is that we are so short on compute, we’ve barely scratched the surface of what these things can do. Actually, it’s not just that in 2 years we’re not going to have a surplus; we’re still going to be in a shortage.
And by the way, that might be true. The extent to which the possibilities are barely being tapped as far as AI—particularly once we get to purely autonomous functionality, where you don’t need to have a human in the loop—the bull case is not insane. And it’s not like a railroad.
This is the distinction from the article—the railroad article. There’s just no way to accelerate the revenue-generation potential of a railroad.
Railroads, yep.
You have to actually—
It’s closer to a toll road.
That’s right. You have to actually build it across brutal terrain—
Yeah.
—which takes a very long time. Then you actually have to develop the land that you got for it. The land has to build up productive functions such that it starts using—
You need trains.
—in the physical world, things—
And routes.
—are slow.
Yeah.
That’s right. And even then, say you instantly had total saturation all over the railroad, you could only run so many trains.
Mm-hmm.
You have to build trains. Whereas with AI, the scalability capability, if this stuff starts working, gives you all the benefits of any digital good, right? What’s the idea of software? You write software once. It’s instantly, infinitely duplicatable. It can be used everywhere.
Yep.
There are aspects of that to AI, particularly when you think about the concept of AI improving itself, AI writing its own programs, AI being set loose on a company and creating agents on its own—
Mm-hmm.
—that figure out all the functions of it. Again, none of that quite works now, but it’s working pretty well, and it’s accelerating unbelievably rapidly. I think we have an emailer in here saying that I’m a Luddite because I took too long to vibe code, which I’ll push back on in a little bit. But it speaks to the point that I’m sorry, my 6 months was too slow for you, and it’s kind of a valid point, right?
Yeah.
The speed with which this is moving is a very real thing, but it’s not a slam-dunk case at all. Also, there’s a real tension: bringing more supply to market will depress prices.
Mm.
Now, you can argue that demand is so high that prices will still go up because demand will accelerate more than supply.
Yep.
But they’re not going to go up as much as if you did bring more supply to market. This is just a math function. The price depends on how much supply you have. It also depends on how much demand you have. The bet is that demand is going to—
Be insatiable, yeah.
—not just increase faster than supply, but increase even more, such that it doesn’t matter how much NVIDIA produces; the price is going to go up. And maybe that will be the case. But the other question is just this timing question. This gets back to the amount of capital in the market. In the long run, you can’t be funding stuff with debt forever. At some point, you need to actually make money, and that money gets cycled back into buying new stuff.
Mm-hmm.
And that, I’m sure, is going to happen. But, like we talk about with stock picking, it’s not enough to be right; it’s about timing.
Yeah.
The big question with these capital issues is, I believe this stuff will pay for itself. The question is, will it pay for itself in time to avoid an air pocket where we run out of money?
Run out of money and leave a whole bunch of bag holders. Sure.
That’s right.
6. NVIDIA Defends Its CUDA Moat
Well, and one other question before we move on. There’s an element of this that read to me, in reading your article, as sort of a defensive move from NVIDIA as Google brings all this infrastructure online, and you’ve got 2 dominant AI players. As everybody becomes more cost-sensitive, there’s going to be an increasingly urgent push to get on TPUs as opposed to NVIDIA chips. So NVIDIA wants to facilitate building out with NVIDIA hardware and NVIDIA software. Does that make sense? Did I read that correctly?
Yeah, so that gets to the micro question—the NVIDIA-specific question. This is a question, by the way, we’ve been talking about for a few years now.
Mm-hmm.
I think it was GTC 2024, so it was about 15 months after ChatGPT had come out, when NVIDIA was truly a stock aflame.
Astride the world.
That was the—
Yep.
—that was the GTC where Jensen Huang was at the SAP Center in San Jose, the hockey arena.
Yep.
And it’s like a rock star thing, right? It’s like the—
I think he may have also signed someone’s boobs at that GTC.
That was actually—no, I think that was in Taiwan when that happened.
Okay.
But I might be wrong.
Either way, same era. NVIDIA—
Yeah.
—just owning the universe at that point.
And I remember that was kind of a boring keynote in a way that NVIDIA’s GTC keynotes were not boring.
Mm-hmm.
Because before ChatGPT, they knew they had this incredible computing capability, this highly parallel—what are the things you can do with it? CUDA lets you program it more easily. I wrote an update years ago where someone was like, “How can NVIDIA announce all this stuff? Why can these keynotes be so cool?” Especially because NVIDIA loves doing keynotes. They do keynotes every 6 months, or actually less if you include CES and things like that. Jensen’s up on stage every 3 to 4 months.
That’s true.
How does he talk about so many new things? The reason is that it’s all the same thing. Everything is just parallel computing using CUDA, and they’re just making all these libraries—
Mm-hmm.
—where they’re just changing a few things, but they’re all the same thing. The reason they were doing that is they were throwing everything against the wall: for every possible application of parallel computing, let’s make a library and see if we can find and get the next market spinning around this, beyond gaming and beyond Bitcoin mining.
It’s funny you say that because, before ChatGPT launched, I remember a GTC that you covered. I don’t know whether I was working with Stratechery at that point, but it just seemed like Jensen was throwing all kinds of crazy ideas at the wall to see what sticks. It was cool. It was like imagining the future. It’s great that he’s got all these ideas. I don’t know how much any of this will actually be real, but he’s clearly thinking about where we’re going to be and how we’re going to be computing 10 years from now. And then ChatGPT blows up maybe 9 months later, and it’s like, “Oh, okay, so this is it.” And NVIDIA’s—
Yeah.
—in the catbird seat.
But the weird thing about large language models is they were obviously incredible for NVIDIA. That’s why their stock went to the moon. They have been on and off the most valuable company in the world. It was also very bad for NVIDIA, and the reason it was bad for NVIDIA is that the play with CUDA is to build a developer ecosystem on top of CUDA.
Mm-hmm.
But CUDA only works on NVIDIA GPUs. So you get CUDA for free, it’s easier to use, and it’s a tremendous investment. NVIDIA almost went under trying to build CUDA at a time when no one understood what they were doing or why they were wasting money on it. And that’s why Jensen Huang will get bristly, particularly when people question their rent-seeking or profit, whatever.
Sure.
It’s like, no, they earned their spot fair and—
He was taking the risks.
Absolutely. And it shouldn’t be forgotten. They have earned every dollar they’ve gotten through 25 years of taking massive risks. And—
And the stock bottomed out several times along the way as they were doing all this.
It bottomed out in October 2022.
Right.
I wrote an article 3 weeks before ChatGPT came out, tracing their bottoming-out history and their—
Mm-hmm.
—search for what was next.
NVIDIA in the Valley. I remember it well.
NVIDIA in the Valley. So, go back to this GTC. I wrote an article at the time called “NVIDIA Waves and Moats.” What was interesting about that GTC was, number one, it was very boring. All the cool stuff kind of got scrubbed out.
Now, Jensen Huang has brought that stuff back. So the last few GTCs, he’s been talking more about other things. Now it comes across as, “Oh, you’re still looking for something beyond the LLM.”
Ah.
Because the problem with the LLM is it shifts the developer platform far above where NVIDIA sits.
Yeah.
All the activity is happening on top of LLMs. No one who’s writing an AI application today is using CUDA.
Hmm.
Some people are, if they’re training their own model and doing some low-level things or non-LLM things. But the vast majority of the energy and all the money and the ecosystem is far removed from CUDA. They have no idea and don’t need to know or care what chips their applications are running on. They’re just on the OpenAI API, or they’re on the Anthropic API.
Anthropic, yeah.
Or they’re using Bedrock in Amazon, and it’s sitting on Trainium, and they’re using a Chinese open-source model. It’s totally abstracted away. And this is why LLMs were bad for NVIDIA. Now, again, all the money they made along the way is worth it, but their moat has been tremendously diminished.
Hmm.
CUDA’s still a moat if you need to do stuff that requires CUDA.
Right.
But the vast majority of stuff and energy doesn’t require CUDA.
All right, and that is the end of the free preview. If you'd like to hear more from Ben and I, there are links to subscribe in the show notes, or you can also go to sharptech.fm. Either option will get you access to a personalized feed that has all the shows we do every week, plus lots more great content from Stratechery and the Stratechery Plus bundle. Check it out, and if you've got feedback, please email us at email@sharptech.fm.