Ep. 034 - The Fight for Fast Tokens, TPU v7, Vera Rubin, and Engrams (AI Supply Chain, InferenceX)
Jordan NanosCam QuiliciBryan ShanAlec Ibarra
- Engrams turn model architecture and memory hierarchy into one co-design problem. They learn 2-gram and 3-gram token embeddings in a hash table, letting the model retrieve familiar meaning instead of rebuilding it through later Transformer layers; the ablations suggest a hybrid of Engram and MoE layers beats either all-MoE/no-Engram or all-Engram/no-MoE. Because token IDs are known upfront, retrieval from DRAM or SSD can overlap with compute—effectively yielding “free layers”—although earlier placement helps the model while leaving less time to hide retrieval latency.
- The design reflects a structural Chinese constraint: less HBM forces more aggressive compression, tiering, and asynchronous offload. The speakers connect Engrams to DeepSeek’s earlier work on reducing KV cache and using HBM, and to domestic accelerators whose bandwidth trails US systems. Their broader call is that Chinese labs will keep moving KV cache and related model state down the hierarchy, while US labs have historically been less forced to optimize around that constraint.
- AgentX shows why KV-cache management becomes decisive in real coding-agent workloads. Replayed sessions captured through SemiAnalysis’s internal proxy have a P50 request near 100,000 input tokens and theoretical cache-hit rates above 98%, because each tool call appends to an already-long conversation. DRAM offload matters little for low-concurrency interactive use, but becomes material around 40–60 tokens/s with hundreds of concurrent agents, offline batches, or reinforcement-learning rollouts.
- Current inference economics can make the newest systems look like “money printers,” provided the workload and pricing assumptions hold. The team’s calculator combines hardware, power, datacenter cost, throughput, concurrency, and OpenRouter pricing; scenarios discussed included approximately a 40.7% margin on GB200 NVL72 and about 50% in a GB300/Moonshot scenario using 30% and 60% of list price. Results remain sensitive to utilization, node availability, software, and traffic.
- TPU V7’s external availability creates the clearest prospective third competitor beside NVIDIA and AMD. In the team’s early testing, TPU V7 occupied nearly the entire low-cost edge of the Pareto curve, with only one point beating it, using an external-TCO input around $1.21. Software is still immature; native Torch support, TPU vLLM, megakernels, V8, and GCP hosting beyond Google’s internal environment could improve the position.
- Vera Rubin looks less disruptive operationally than Blackwell but potentially much stronger economically. At a roughly 100-token-per-second operating point, the early discussion suggested an improvement approaching 3×, while high-throughput gains could be closer to about 1.4×; the roughly 3× HBM3-to-HBM4 memory-bandwidth progress helps explain the memory-bound middle. A hardware-accelerated lookup-table quantization path—8-bit table values, 3-bit keys, roughly 3.125 effective bits—remains an area to watch, with no major application identified yet.
- The debate over fast tokens is unresolved because agent workflows contain both latency-sensitive and latency-insensitive stages. Bryan, Cam, and Alec see limited ROI when several agents run in parallel or 30-second tool calls dominate, especially at prices approaching $300 per million output tokens; Alec also cautions that future workflows may change. TensorRT’s hundreds of tokens per second on NVIDIA GPUs strengthens the case for specialized inference software, but batch-size-1 megakernels, changing prefill/decode ratios, and difficult deployment economics complicate custom-silicon theses.
1. Engrams replace repeated semantic work with addressable memory
Cam’s framing: ordinary token embeddings accumulate meaning as they traverse Transformer layers, but Engrams instead learn embeddings for two- and three-token sequences, hash them into a table, and inject the retrieved representation into the model.
The motivating discussion referenced DeepSeek work comparing next-token prediction capacity across hidden-state depths and KL divergence relative to the final layer. Bryan’s explanation was that Engram moves word-meaning information earlier, so later layers do not need to spend as many parameters reconstructing it.
The key ablation was not “replace everything with memory.” The loss curve improved around a midpoint combining Engram modules with conventional MoE layers, while the all-MoE/no-Engram and all-Engram/no-MoE endpoints were less effective. That is why the team describes the gain as receiving several “free layers,” not eliminating Transformer computation.
Placement creates the central co-design trade-off. Earlier retrieval gives more downstream layers access to the memory and helped the model more in the reported ablations, but it leaves little computation behind which to hide a DRAM or SSD fetch; placing the module later provides more opportunity for asynchronous overlap.
2. Memory scarcity is shaping Chinese model architecture
Token IDs are available before the forward pass begins, so an Engram lookup can start immediately through unified virtual addressing and pinned host memory on the CPU. Unlike retrieval dependent on an intermediate hidden state, it can overlap with other parts of the model, making DRAM—and potentially SSD—a practical memory tier.
The speakers place Engrams alongside DeepSeek’s earlier work on reducing the KV cache and using HBM. Their broader point is that information that can be retrieved early and overlapped becomes a natural candidate for offload.
Chinese accelerators intensify that incentive. Domestic systems may use a Chinese-built HBM-like technology with lower bandwidth, so Chinese labs are pushed toward smaller KV representations, compression, lower memory tiers, and architectural tricks; US labs have had more HBM and what Cam described as effectively unlimited compute relative to that constraint.
3. AgentX makes the KV-cache problem measurable
Cam’s description of an agent trace: a user sends a task, the model emits text and tool calls, tool outputs are appended, and the growing conversation returns to a stateless model. The KV cache “grows at a crazy rate” across every turn.
AgentX uses coding sessions captured through a proxy used by SemiAnalysis employees, removes unsuitable requests, and replays the remaining traces against open-source serving stacks. The team concedes this does not represent all coding traffic, but argues that the basic shape—task, tool loop, added context, repeated prefix—is common.
The striking numbers are a P50 request around 100,000 input tokens and a theoretical cache-hit rate above 98%. Most turns reuse much of the prior context, so routing a request to the right prefill server, transferring KV between prefill and decode, and preserving cache residency all require careful system design.
Offload’s value depends on the operating point. At low concurrency and high interactivity, there is little KV to spill and slower secondary-tier memory adds limited value; around 40–60 tokens/s with hundreds of concurrent sessions, offline batching, or reinforcement-learning rollouts, DRAM offload becomes relevant. Bryan separately said he would like to see cache-hit rates in the 20–30% range for HBM rather than VRAM.
4. Serving margins reward utilization more than benchmark supremacy
Bryan explains that the calculator begins with hourly ownership cost—accelerators, datacenter costs, electricity, and associated infrastructure—then combines those figures with model throughput, concurrency, utilization assumptions, and OpenRouter pricing. It is intended to translate benchmark curves into the profit available from different chips under a specific workload.
On the assumptions discussed, a GB200 NVL72 deployment could earn approximately a 40.7% profit margin. Other margin profiles were shown for Kimi K2 and DeepSeek V4.1 Flash. Jordan also cited a GB300/Moonshot scenario using 30% and 60% of list price that produced about a 50% margin.
The pushback is embedded in the model: 60% utilization means the system is effectively saturated unless nodes are down, and providers may have poor utilization, different software, or different compute prices. Bryan said some providers can still be profitable at 30–40% utilization. Jordan’s point was that providers want to recover rack capital within months, not that every scenario guarantees that result.
5. TPU externalization turns Google into a platform competitor
Alec’s early TPU V7 results put the public-cloud offering on the cheapest part of nearly the entire price-performance Pareto curve, with only one point beating it. The team used an external-TCO input around $1.21 and emphasizes that these are first-day results, before further software optimization.
The comparison is imperfect: TPU’s silicon, programming model, accessibility, and community support differ materially from B200 or B300. Fixed-size systolic arrays can waste compute through padding, and Google’s internal models were historically designed around constraints that open-source models do not necessarily respect.
Software is the obvious catch-up path. The speakers expect TPU vLLM, more megakernel work, and the forthcoming Torch-TPU backend to reduce programming-model friction; v8 is part of the next stage of TPU development, while v7 itself can improve as kernels and public-model support mature.
Bryan’s larger call is “large-scale externalization”: Google will still use a large amount of TPU capacity internally, but will sell some of it. GCP hosting TPUs outside Google’s own environment could open the ecosystem to more customers, while Cam described AMD and NVIDIA as gaining a real third competitor.
6. Rubin, fast decoding, and custom silicon split the workload
The move from H100/H200 HGX systems to GB200 and GB300 introduced an Arm CPU, scale-up NVLink switches, new networking, and liquid cooling; Blackwell’s SM100 also introduced a new FP4 data type. Vera Rubin appears operationally more incremental: no major new scale-up network or Arm CPU, and liquid cooling is already deployed, but with substantially more memory bandwidth.
The performance curve is workload-dependent. Around 100 tokens/s, the early discussion suggested Rubin could approach 3× the performance or users at a given interactivity level; toward maximum throughput, the advantage could be closer to roughly 1.4×. The HBM3-to-HBM4 bandwidth progress of about 3× maps closely to the memory-bound middle of the curve.
Rubin also exposes a new hardware-accelerated lookup-table quantization path using 8-bit values and 3-bit keys, around 3.125 effective bits. The team has not yet seen the specialized application and explicitly leaves open how much use the open-source ecosystem will make of it; comparisons with MXFP4-like scaling and accuracy evaluations remain areas to watch.
Fast-token demand produced the sharpest disagreement. Bryan prefers many parallel agents; Cam notes that a 30-second tool call can nullify thousand-token-per-second decoding; Alec questions the ROI and says the price is high. Alec’s historical caution is that one or two years ago he would not have predicted his current AI workflow, so future demand is difficult to forecast. Cam’s prediction is that fast modes will be useful for some tasks, while batchable work can tolerate more latency.
TensorRT demonstrates the software extreme: fuse much of the model into a giant dataflow-style megakernel and achieve hundreds of tokens per second, including a chart point near 500 on a reasonably large model approaching one trillion parameters, using NVIDIA GPUs alone. The cost is enormous engineering effort, slow model onboarding, and batch-size-1 behavior that favors interactivity over high-throughput serving.
Jordan is therefore unworried about accelerator startups finding initial customers because “demand is completely ahead of supply,” but the group disputes whether that proves durable economics. Building working silicon is now plausible; achieving good yield, packaging servers, securing distribution, operating datacenters, adapting to shifting prefill/decode ratios, and competing with NVIDIA’s software and controlled GPU supply remain the harder business problems.
Full transcript
It’s been about 2 months since the last episode. This week, the AgentX team joined me. We’re talking about Engrams, AgentX, and TPU offload, and we have plenty to discuss. Joining me are Cam, Bryan, and Alec. Welcome to the show, everyone.
Hello. Happy to be here.
Good to be here.
Good.
1. Engrams
A new article came out on September 18, about 2 weeks before this episode. Let’s start with that. It’s about Engrams: offloading some of the model components, or KV caches, to SSD, and what this means for model architecture co-design. What are Engrams? Can you give us a short introduction?
Engrams are basically a way for models to transfer more information into embeddings. The idea has existed for a while, I think. Bryan can touch on more of the history, but, basically, in a Transformer architecture, each token has an embedding. An embedding is what gives a token meaning in a latent space.
As tokens pass through the Transformer’s layers, each layer adds more information to that representation. The tokens are related to one another, and that information flows through the layers into a hidden state.
The Engram idea is basically that, during training, instead of learning all of this meaning across multiple layers, you learn some 2-gram and 3-gram token embeddings. These are stored in a hash table. The lookup is essentially free—not completely free, but close.
If you look at the 2 graphs on the right, the experiments are ordered by depth. On one side, there are all MoE layers and no Engram layers; on the other, there are all Engram layers and no MoE layers. The loss is actually reduced at the midpoint. When there are a few Engram layers and slightly fewer MoE layers, the performance is very good.
That’s shown in Qwen3-8B, I think.
Bryan, could you speak a little more about the history of Engram and its inspirations?
This year, DeepSeek introduced a paper arguing that, for language models, representations of word meaning are built in the later layers. They compare the next-token prediction capacity of hidden states at different layers. There was an earlier paper—I forgot what it was called.
Basically, if you calculate the KL divergence between the hidden states in the previous layers and the last layer, the later hidden states for Engrams have a low KL divergence. This shows that, instead of using the later layers of the model to predict the next token, Engram is moving that information earlier. That’s useful because the model doesn’t have to waste parameters trying to reconstruct the meaning of words. They put the second Engram layer in the second layer, because if you put it in the first layer, you don’t have enough time for the retrieval from DRAM or SSD to return, so there’s very little overhead. But in the other Flash model, the interesting part is that they put the layer in layer 1. There was still room for overlap, but not as much.
One really special thing about this model’s co-design is that, when you’re considering the architecture at inference time, the Engram modules can be asynchronous. They can overlap with other parts of the model because Engram module inputs are just token IDs.
During inference, the model receives token IDs. You can use those IDs to retrieve the hidden-state additions, so overlapping the Engram lookup is very natural. In the original paper, they mention a trade-off between putting Engram layers earlier versus later. Putting them earlier helps the model more, based on the ablations, because earlier Engram layers can bring in previous semantic memory for as long as possible.
But the problem with putting Engram layers early is that there’s nothing to overlap with. You have to wait for the Engram retrieval to finish before continuing. Jordan, can you pull up the chart?
Yes, I’m trying to find it. Don’t worry.
While Jordan finds the chart, let’s talk about how offload works. Offload uses UVA—unified virtual addressing, or unified virtual address, sorry. Basically, this moves the Engram to pinned host memory on the CPU.
Here, we’re comparing the DRAM and SSD offload versions.
Yes, thank you, Jordan. We also tried HBM for Engram. We didn’t discuss it in the article because the cause-and-effect relationship wasn’t very interesting, but HBM is one step in this direction.
Go back, Bryan. The main point is that there’s an HBM controller, and presenting Engrams as a control mechanism around DRAM- or SSD-based offload makes it easier for people to innovate, right?
Yes. Over time, this is a natural progression. A lot of this goes back to last year, when HBM was limited. DeepSeek’s work was about reducing the KV cache for the whole model and using HBM.
I don’t know if Engram was specifically an HBM co-design. It was partially a control mechanism. The idea is that the model can retrieve the information it needs from memory without requiring everything to remain in HBM.
Over the past year, a lot of the conversation around 12-high, 8-high, and 4-high Rubin D-spec has been about American labs optimizing HBM bandwidth, capacity, or bandwidth per bit. Chinese labs already have that limitation because of the chips they can access. If bandwidth isn’t available, you need to do something to improve high-concurrency serving or high-interactivity workloads.
Cam, you’re also an expert in this area. Maybe you can explain it.
Chinese chips are very interesting. In our previous DeepSeek V4 day-zero article, we talked about the Ascend 950 chip design. They use HBM, although they don’t call it HBM. It’s their own Chinese-built version. Of course, the bandwidth is lower, but they have very creative hardware solutions to deal with it.
At a high level, the trend is that the United States has more HBM, while China has limited HBM and limited chip-to-chip bandwidth. Therefore, Chinese labs are trying to make KV caches smaller and compress them more efficiently, then offload them to lower levels of memory.
Engram looks like a natural extension of that because it effectively gives you free layers. As Bryan mentioned, you don’t need to increase the representation, and it can be offloaded to DRAM. Engram can also overlap with other computations because it starts with token IDs.
It seems like Chinese labs are trying to be more efficient. In the United States, labs have effectively unlimited compute and are collecting more HBM because they aren’t forced to make these trade-offs.
That’s very interesting. It seems like a lot of innovation is coming from China.
2. AgentX Benchmark
Let’s talk about the workload implications. There’s an obvious desire to store KV because many requests can be served at the same time, and the requests are very large. Agentic coding and other coding-for-agent workloads need high throughput for long-context serving.
Agentic coding was a great inspiration for developing a benchmark that more accurately represents this workload pattern. We applied it to all the new models and released the initial dataset and results. This is what AgentX is called.
Cam, what is AgentX? Compared with traditional A/K1K, what changes does AgentX represent? And how does the model architecture differ in terms of efficiency? Let’s use this benchmark to show that clearly.
Can you tell me?
Good question. Before, we were doing A/K1K with prefix-only, no caching, and testing at the chip level. Now we’ve changed to AgentX. Basically, this is agent traffic only, isn’t it?
In this setup, you’re looking at an agentic harness using code. What happens is that you send a message, add the output, maybe call some tools, add the tool output to your messages in a streaming fashion, and then keep sending that back to the model. The model is stateless, isn’t it? But as we know, the KV cache grows at every turn at a crazy rate. When you have models with millions of tokens of context length, you need to store a lot of KV cache.
The idea behind AgentX is that we’re not testing chip-level performance alone. We’re testing the whole system and looking at how KV cache moves between chips, how prefill and decode disaggregation work in different circumstances, and how offloading works.
Can you pull up some results on inferencex.com?
Yes, of course. We replay the traces, and the results are really interesting. The cache-hit rate—the theoretical cache-hit rate—is very high. Many people are surprised that it’s above 98%.
What does that 98% mean for your system? If you click on “B300 vLLM,” for example, what does that show?
The system is shown with offloading enabled. For high-end, highly interactive workloads, this is not very important, because offloading to secondary-tier memory is slow. When you want low concurrency and high interactivity—sorry, I said that incorrectly—you don’t need to store that many KV cache entries, so offloading is not especially relevant.
But when you’re in the range of 40 to 60 tokens per second, with very high throughput, offloading becomes relevant. That could be an offline batch assumption, hundreds of concurrent agent sessions, or reinforcement-learning rollouts. In those cases, offloading is very relevant. For HBM rather than VRAM, I’d like to see cache-hit rates in the 20% to 30% range.
I see. So, yes, AgentX is intended to show what an actual workload looks like. I think, Bryan, you may want to talk about the profit calculator.
The hourly cost is why this benchmark is so interesting. For Cerebras, our TCO model calculates the total cost of ownership for the accelerator chips. Another way to think about it is: how much does it cost to run the system for 1 hour?
That includes data-center costs, electricity, chip costs, and so on. For InferenceX users, we’re giving you a number that will be shown soon. This is the cost basis. On the router, depending on the assumptions—how many users you have, what the concurrency is, and what throughput you’re achieving—we combine all those numbers and calculate the profit you can make using different chips.
For example, with a GB200 NVL72 serving one of the highest-end models, I believe you can achieve approximately a 40.7% profit margin. With Kimi K2, depending on the parameters, there’s another margin profile. With DeepSeek V4.1 Flash, you can see the trend change under different settings, and in some cases you can’t make a profit. That’s very interesting because I believe the methodology is going to be useful.
The biggest realization is that, under the open-source assumption, you can achieve real profit margins by running these models. Compute cost is normal at this level, but these systems are essentially money printers.
To break this down, with Dynamo and vLLM on an existing GB300, I’m seeing a 60% reduction—or, sorry, we’re assuming 60% utilization. That’s probably too much. At 60% usage, the system is always completely saturated, unless some nodes are down. For Moonshot, at 30% and 60% of the list price, you can save on GB300 and achieve about a 50% profit margin. This is just a money printer, isn’t it?
In AgentX, let’s use the assumed cache-hit rate and then look at the OpenRouter data. This is very practical.
Let’s go back for 1 minute and talk about the dataset. The agent-coding request distribution is interesting. We discussed how the P50 request is still about 100,000 input tokens, and that has implications for the cache-hit proportion. Some people might say, “That’s real? How do you know?”
We replayed our own data, which is why it’s real. All SemiAnalysis employees have agent-coding sessions intercepted by a proxy. We take those requests, remove the bad ones, do a little post-processing, and then replay them against the original open-source servers.
This does not completely represent all agent-coding traffic, but the actual workload shape is very simple. Everyone has the same basic process: you have a task, you send it to the model, the model suggests something, it decides to use some tools, collects more context, and then writes. You repeat that self-agent loop until the task is complete.
The key point is that, most of the time, the cache is a hit. So how you store and handle the KV cache, how you transfer it, and how you route requests between different prefill servers all have to be handled intelligently.
As you can see, there are a lot of parameters you can set. We use NIXL as the KV-transfer engine. For CPU KV offload, we use vLLM, and we’re only using DRAM offloading. We’re also using the Dynamo router. None of this is trivial; there are many ways to configure it.
But to answer the question, the data represents actual captured coding sessions. That’s why it’s useful.
That’s very interesting. Let’s explore all of this at inferencex.com.
When you ask how the per-gigabyte profit is calculated, we take our original traffic, apply the discounted usage rate, and then divide by the original hardware-cost numbers. If people are using the newest chips and the best models, it can be unbelievably profitable.
If you’re listening to this podcast and you have a few million dollars, you might think about buying a GB300 rack.
No, I’m serious. You want to get your return on investment as quickly as possible, ideally getting your money back within a few months.
For inference providers, they have their own driver kernels, and they work on OpenRouter. Even if you go back to the page, you don’t need to; even at a 30% or 40% utilization rate, it can still work. For some inference providers on OpenRouter, there’s no change in the result.
Their utilization might be very poor, but it can still be profitable. I don’t know what the compute price will be. I’m not a finance expert, but I think this is a normal level. It’s still very profitable.
At the beginning, people had a lot of questions about this. Alec, can you speak to that? I want to bring you in and encourage you to take this one.
Two things are changing. First, the models themselves are becoming more efficient and higher quality. We discussed how new models can charge higher prices because of their quality, but as the cost of serving them goes down, they can earn more money, especially with KV-cache optimizations.
3. TPU v7
Hardware is also getting better, so there’s more competition. You recently wrote another article about TPUs. Maybe the most exciting announcement is TPU racks. Here’s the picture. It looks like you can rent them on GCP. What’s missing from the assumption? Everyone talks about NVIDIA and AMD, but what happened with TPUs?
Alec, speak to us at a high level. What has happened with TPUs over the years?
We tested TPU v7. Google’s own accelerator appears to be available through the cloud, which makes it the first accelerator you can buy that way. We wanted to do an apples-to-apples comparison with the B200 and B300 and see what the performance looked like.
On the Pareto curve, the TPU v7 points are blue. They’re below the others, which is the important part. Except for 1 point, it’s the cheapest option across the board. That’s amazing.
It depends on the TPU accelerator’s TCO. For the external TPU TCO, I believe it’s $1.21. That’s very good. They’ve extracted amazing performance from it.
This is still very early. These are first-day performance assumptions, so software improvements will make it cheaper. This is the open-source version, though. It’s not a major assumption among providers because they’re running it internally, I think. For public access, this is what’s available.
Yes, let’s talk a little about the trend. Obviously, the TPU is very different from the B200 or B300 in terms of chip design and programming model. How much access do people have to it? How much of a community is there, how open source is it, and so on?
When you compare TPU v7—or perhaps the current generation—with TPU v8, which is coming very quickly, it’s like comparing apples to apples and then apples to bananas. There are a lot of differences. Google’s highest priority is to optimize certain things and make models open source, as we’ve discussed. vLLM-TPU is very new, so what is the trend for TPUs in the future? It seems like they’re going to get better from here, right?
Recently, we’ve seen that TPUs and megakernels can provide a major performance boost. It’s exciting to see what comes from that. Google has produced many TPUs over the years, but they were all used internally, with the TPU engineering team providing full support for those workloads. Now customers will be able to use them as well.
vLLM-TPU is in preview. How is that looking?
Two things are becoming clear. First, as we know, we’re seeing the large-scale externalization of TPUs, with v7 and, later, v8 becoming more mature. Google is moving away from keeping all of this internal for its own workloads. It will still use a ton of TPUs, but it will sell some of them.
Second, you’ll see major progress in the open-source ecosystem. Right now, TPU doesn’t have a native Torch backend, right? It passes through intermediate representations in a different way. Torch-TPU is very close to coming out, and that will provide a native Torch backend for TPU and make the programming model easier from a software perspective. You’ll immediately see improvements in the open-source ecosystem. I think that will be a trend for some time.
Bryan, is there anything you’d add?
In the past, TPUs were built around systolic arrays, which prioritize data movement. One problem with a systolic array is that it has a fixed starting point and a fixed size. With smaller workloads, you have to pad them, which wastes compute.
Because TPUs were mainly used internally, many models were built around that limitation. In open source, that wasn’t necessarily the case. This was a limitation for inference workloads, and it held TPUs back. As a result, the full power of the TPU wasn’t really being used.
A year ago—maybe two years ago—everyone at Google and DeepMind was bullish because TPUs gave them more compute. I don’t remember the exact number; perhaps 3 million TPUs had been manufactured. Google also counted the amount of research being conducted on TPUs, which was another reason people were so bullish.
If you attend any machine-learning conference, you’ll see Google and DeepMind papers everywhere. Among prominent people in machine learning, almost everyone has had some period of their career at DeepMind.
Yes, the fact that TPUs are actually being externalized is very important. AMD and NVIDIA will have a real third competitor, and it’ll be interesting to see how things develop over the next few years.
I want to run private inference support in vLLM—or a private fork of the original public repository—on some permanent clusters. That’s really exciting.
Another thing is cloud TPU deployment. For the first time, GCP is going to host TPUs outside Google’s own environment. That will open up the ecosystem, and more customers will use TPUs through GCP for workloads beyond Google’s internal use.
People will actually be able to reduce their costs. If they don’t have to build everything themselves and can go through GCP, they’ll be happier.
Alec, tell us about the GCP console. What are your top 3 cloud consoles? What do you think?
My top 3? I haven’t used many of them, so I can’t really say. To be honest, I prefer the command line. You don’t need to navigate all those crazy user interfaces.
What about Codex or Claude Code? Which is your favorite? Of the 3 major cloud CLIs, which one is best?
GCP, bro. When you’re updating GCP from the command line, you can give it 14 prompts and it’ll do it. It installs a lot of things. It’s free, I think—the software is free.
I think you can now install the whole thing with Claude Code. Or maybe you opened Claude—
4. Vera Rubin
Okay, moving on from TPUs, some new hardware is hitting the inference wires: Vera Rubin. Let’s talk a little about the difference between Vera Rubin and Grace Hopper.
Jensen is sandbagging again, isn’t he?
Of course he’s sandbagging. Look, this is continuous progress. People are asking you to serve interactivity, and this is no longer the case. I don’t think that older setup is used anymore, but the current systems are actually serving interactivity in this range, and it’s 3 times better.
That means you can serve 3 times as many users at the same level of interactivity. You can charge for 3 times as many users at the same price, although you’re not necessarily worth that. If you go to the profit estimator, you can see the effect.
Bryan, can you talk more about the architectural improvements?
There are more FLOPs and more memory bandwidth.
Yes, that’s excellent. One thing I mentioned before on the show is the transition from Grace Hopper to Vera Rubin. In terms of the physical system alignment, there’s a small improvement going from GB200 to Vera Rubin.
When H100 and H200 HGX systems move to GB200 and GB300, they introduce an Arm CPU. For the first time, they also introduce scale-up NVLink switches. For GB300, there’s a new 400-gigabit or 800-gigabit network. Blackwell’s SM100 is a complete redesign; in the past, there wasn’t a new FP4 data type.
Vera Rubin looks like an even bigger change. There’s no major change in the scale-up network and no new Arm CPU. Liquid cooling has already been physically deployed, so everyone should be familiar with it. On one hand, that makes it easier to adopt. On the other hand, you may look at it and say, “Compared with GB300, I’m not getting a major performance improvement.” There may be less inspiration to move to the new system.
Bryan, can you talk about the initial software-porting experience with Vera Rubin?
To back up what Cam said, compared with GB200, Vera Rubin should have token-based performance benefits and cost benefits. Vera Rubin isn’t as big a jump over the previous generation as Blackwell was. I agree with your comment.
We’re still early, but some interesting things have been said about Vera Rubin. We’re about 8 months away. Because of the shortage of VRAM, scale is important. Since the network prioritizes memory bandwidth, labs may be able to use model sparsity and similar techniques to make more GPU memory bandwidth available to users.
That means the capacity of each individual GPU’s memory may not be as important, because aggregate memory bandwidth is what matters. We can’t say much more about the complexity of Vera Rubin yet.
Another interesting new feature of Vera Rubin is its NVFP4 implementation. We’re seeing that for the first time, and it’s very interesting.
For those who haven’t read our article, Vera Rubin’s internal documents show a new hardware-accelerated path for NVFP4. This is already confirmed in the PTX—I forget which version, but we’ve discussed it before. There are also open-source Triton branches and forks implementing NVFP4.
At a high level, they’re doing quantization with a lookup table. If I’m not mistaken, they’re using 8-bit table values with 3-bit keys. If you remember, that comes out to approximately 3.125 bits. NVIDIA provided a special hardware-accelerated path for this in Rubin.
We haven’t seen any applications yet, so it’ll be interesting to see what comes out of it. It may be that there isn’t much specialized usage in the open-source ecosystem so far, but there’s a lot to watch.
What is that—MXFP4, or an MXFP4-like scaling quantization? How does it compare? Jordan, thanks for showing that.
What is it specifically used for? It’ll be interesting to see. You can actually emulate it. In Blackwell, you can emulate it to measure how much efficiency you get with high-accuracy weights. You can use evaluations to check whether you’re losing things like perplexity.
A lot of work is coming out of this. Rubin is very interesting in some respects. Even if there isn’t a major architectural difference in every area, performance itself is the point. On the performance side, we’ve seen a remarkable increase.
These are somehow very early results. It’s like, the faster we are, the better results we’ll get, I think. This is before Blackwell comes out, so maybe we’ll get better results. This is how it seems to me, isn’t it? That’s it. It’s crooked.
So, with the base comparison, what is the multiplier? To see which one is on the curve, the point you choose depends. If you take 200 tokens per second for people to serve, that’s impossibly good, because now Vera Rubin and NVL72 can do that. But in the middle, at 100 tokens per second, making it 3 times better is a different story. In high-throughput cases, 3 times better may not be the best; it may be greater than 1.4. So you’re asking whether a 40% improvement is good for throughput.
For me, it’s interesting. Oh, sorry, Jordan, but this is CPX for me. I’m reminding you that the sound on CPX is too much decreased, but because of the CPX workload, with the prefill processing and the decode phase, it will do that. Our inference benchmark has no CPX. We hope to get CPX results soon, maybe to check how CPX improves disaggregated workloads.
This is different at the ends of the curve: failures, bandwidth trade-offs, and how to optimize. It makes people think. For me, it’s clear: one GPU with Rubin is a big step in memory bandwidth. When you move from HBM3 to HBM4, you’ll get about 3 times the progress in memory bandwidth. Assuming the curve is memory-bandwidth-limited, the big middle parts of the curve seem to match that 3-times progress.
What do you think, Bryan? This is an early result looking at lookup-table quantization schemes, and that exploration hasn’t been done on the GPU. There are some other hardware-acceleration features that haven’t been used so far. On a per-dollar or per-watt basis, or even a dollar-watt basis, Rubin should be more efficient than Blackwell. We’re seeing evidence of that here.
It’s a great chip. The price hasn’t been determined, but it will be priced aggressively. By aggressively, I mean that people are going to buy it because they know it’s expensive. If you get 3 times the performance for 2 times the price, that’s not a bad trade. That’s what everyone wants to do next year.
Going back to the previous profit-estimator calculations, people are thinking, “Yes, I can carry the calculation forward.” It’s clear: “Yes, I want the newest and the best.” NVIDIA has done that again. Jensen’s sandbagging performance again. There are really parts of the curve that are memory-bandwidth-limited, and NVIDIA is telling you, “It should be possible.”
5. Fast Tokens
For the last question, with coding agents—you use many of them—can you speak about your experience using fast tokens? The consequence of these performance curves is that people are going to pay a premium for more than 100 tokens per second. We’ve had some heated discussions on the phone, but I won’t share mine. What are your initial thoughts on fast versus ultra-fast? Bryan, can you start? Are you using Super, Max, or bare? What are you thinking?
No, no, I’m not doing that quickly. My personal use case is that I like working with multiple agents at the same time. I use multiple VS Code windows, and in each VS Code terminal I’m using multiple agents, basically doing different jobs.
How quickly the agent finishes isn’t a problem for me because, by the time I return to the agent, the other things I was spending time on should already be finished. But, Cam, you have a different perspective on ultra-fast models.
6. Profit Calculator
Yes, I do. It’s very strange because my thought is that, as tasks become more agentic, you have these powerful models that can call tools very well, don’t you? They can bring in more context to complete things and call tools. But tools take time to run. Cerebras can serve 1,000 tokens in a second, but tool usage can take up to 30 seconds. In production, I don’t understand that speed because you’re still interrupted by the tool-usage time.
If you’re creating HTML pages all day—100,000 lines of HTML—that could be very useful, I think. These models, in terms of my use, give me a good hint. No, no, I’m kidding. Alec, what’s your comment?
Sometimes I need to do something very quickly or get an answer suddenly, and for things like that, fast mode would be good. But most of the time, I have multiple agents. While they’re working, I can do something else. I have the code, and it doesn’t need to be immediate, I think. It doesn’t need to be merged instantly, you know? That can happen within 1 or 2 hours.
So I’m not using it that much anymore. It’s very expensive. Is there an ROI? I think so. I’m not sure.
That’s what I mean, especially when you’re interrupted by tool-usage time. This is very fast stuff. I’m not paying $300 for 1 million output tokens, right?
Okay, then let me give you the opposite fact. I agree about my current use case, but 1 or 2 years ago, I wouldn’t have predicted that I would use AI in this style today. So what will it be 1 year from now? I can’t confidently predict that fast tokens of any kind wouldn’t be useful. It’s always difficult for me to predict.
Let’s push this to 2 extremes. On one side, there’s a model that’s really big. On the other, if I may say so, the network is really big. Where does it slow down—the network, the device you use, or the endpoint? If the model is where it slows down, then nothing comes back because it isn’t giving you anything. You’ll be disappointed.
Can models that are 10 times bigger but 10 times less efficient still be more useful? Can we imagine a world where they are? We need them to hurry up.
This is a very interesting question. As you scale, you get better results; the scaling laws show that. So the approach to scaling is interesting. That’s a bitter lesson, isn’t it? Scaling computation allows computers to take over.
Recently, I was thinking about why scaling can’t deliver usable token output quickly at the right price. But what is the industry actually doing? For example, Astra 6.1 Loop Transformers are rumored, although that isn’t confirmed. That’s very interesting, isn’t it? They’re showing where the industry is going. It’s scaling, but it doesn’t seem like they’re bringing architectural changes. It’s not a new method, but it is a very interesting innovation.
Yes, even Loop Transformers require a lot of computation. How much computation you do for reasoning isn’t free, right?
Yes, exactly. On reasoning, that’s a good point. If you’re using Ultra Max, or if you like reasoning, that completely affects output speed. For example, if you need to decode very quickly, that’s important.
To be honest, my prediction is that, for some tasks, it will be possible to route to a fast mode, while tasks where low latency is unnecessary can be batched. If you’re bootstrapping a React app, having fast tokens may be useful. But as Bryan said, if you’re running a performance benchmark all night to climb the hill, and the benchmark takes 50 minutes, then it takes 54 minutes and costs $40. Spending a few thousand tokens on decoding and then waiting 50 more minutes is just a waste of money.
So that’s my prediction: it will be very specialized, but it may still be useful for some things.
The other question is intensity. What about models that are really small and really efficient while still being high quality?
Let me give you an example. If you’re someone who bootstraps a React app and you’re using OpenAI, Manus, Operator, or Perplexity Computer, you can get instant responses. You think, “Wow, this is really interesting.” Do you have that experience?
It’s an interesting direction, but personally, I don’t think we’re going to arrive at a world of small models. In the history of computing, if you make something more efficient, someone else adds more things. CPUs and RAM have improved a lot, but websites with libraries have also become much slower and more inefficient.
So I don’t think a world using small models—a true small-model world—will really come true. But I could be wrong. If I’m proven wrong, that will be interesting. It would be good for the environment.
That’s my short tangent. Congratulations. When you’re using all these different agents in parallel, the return time may not really matter to you. Even if it’s 10 times slower, you can still get the result eventually. For me, though, I have doubts. If you go to sleep, wake up in the morning, and still haven’t received a response, that’s not good.
So, in the future, there’s a place for every kind of fast token. Maybe not all of them need to be fast, right? Also, Jordan, in America, even Big Macs don’t have a fixed price, do they? So would you pay up to $100 for 1 million output tokens? I don’t know. I really don’t know.
Of course. I agree. Speed is important, to some extent. Jensen has always emphasized this, and he also said in the keynote that fast tokens are excellent tokens. Obviously, he is inclined to say that because he sells GPUs.
But, to some extent, that's right. If you can get instant output, I think that can be useful in some cases because you get an answer quickly. But I don't know. Time will tell.
Who knows? Let's see. Okay, let's talk about this—we need to talk. That was a perfect lead-in, man. I was going to conclude this, but let's do this. Let's end it.
7. TileRT
No, maybe we should talk about TensorRT. I don't think so, because there are many people talking about ultra-fast models. When OpenAI's DevDay announcement came out, everyone said, “Oh, okay, OpenAI and Cerebras have a deal.” Over the next few years, they are promising a megawatt-scale deployment.
Obviously, Cerebras is impressive, but TensorRT is also very capable. From our perspective, if you look at the models and the optimization TensorRT is showing, they are using similar techniques. We're seeing 200 tokens per second here, and, as shown on the screen, a reasonably large model approaching 1 trillion parameters can reach up to 500 tokens per second on NVIDIA GPUs alone, right?
Yes. For those who don't know, TensorRT is basically an inference engine optimized for ultra-low latency. It is designed to achieve as much interactivity as possible without needing specialized SRAM architectures. These results are bounded by physics, of course, but they show what is possible with highly optimized inference engines.
There is also a result from AMD. Going forward, as you look at different use cases, inference engines will become more specialized. I think that's the trend we're already seeing.
TensorRT is very cool because you don't need a new chip. You're building on the best chips. How does that work? Can you speak a little about it? NVIDIA has SRAM in its GPUs, obviously, and there are many ways to configure the serving setup, write special kernels, and provide model-serving support. Technically, let's talk about it. For high interactivity, this is great. Bryan, how does it work? But for high throughput, this is bad, isn't it?
Yes. With TensorRT, you can potentially see the entire model as one big megakernel. Therefore, there is less of an interactivity problem because the data flows through the model. Companies such as Cerebras built dataflow hardware—dataflow machines. That is a little more difficult to do with GPUs.
With TensorRT, you can run at a batch size of 1. Model releases are actually very slow because getting that megakernel to run requires a great deal of software effort, and GPUs are not particularly friendly to this kind of dataflow. But GPUs are very common, which makes it interesting. They can use specialized hardware to replace some of that, in fact.
For an AI agent to read anything and then enter a giant megakernel seems impossible. How much of this is actually possible? But it works that way: you get good accuracy and high interactivity, right?
Yes, but people were doing this. That's for sure. Of course, it is important for agents because there is a lot to do, and you want to be able to do it easily. I personally believe in it because it can be checked very easily. We've seen many other kernel competitions on GPU leaderboards, and we've seen agents do this.
This is still very new, though. Recently, I've seen these kinds of megakernel ideas coming out on MLX. Cerebras recently released SLS on MLX. There have also been other developments—for example, TPU megakernels. The idea is the same: use the entire model as one megakernel. That's very interesting.
8. Chip Startups
Jordan, I have a question for you. You don't have to name names, but if you're a chip startup whose whole thesis is fast tokens—like the megakernels you're talking about, such as Jalapeño—are you worried about interactivity in this area changing faster than general-purpose GPUs overall? What do you think?
I'm not worried, because demand is completely ahead of supply. It is possible to build chips, and anyone can find a customer. I have faith in that. People need all kinds of chips.
In the long term, these chip startups will be able to deploy what they're working on, and that will make or break their businesses. If you have a large business, you should be able to handle the supply chain and build a group of chips. You should be able to get good yield, build servers, build a data center, and train people to operate all of it.
In some ways, getting something to work well is much more difficult than building it. If the interactivity is high enough, and you have both high interactivity and high throughput, I'm here for it. I think that's the case.
If you don't need a lot of dynamic batching, the ratio between throughput and interactivity can still work. If you reach 1,000 in interactivity, you'll get 1,000 in throughput. If you reach 2,000, your throughput is 2,000. Therefore, I will favor custom silicon that provides as much high interactivity as possible.
But that will be very expensive. For example, if you use a batch size of 1 and get 1,000 tokens per second, each user is paying for the entire system. You can do one thing per batch, but if batch size 1 is really fast, you can pipeline everything. You can achieve high effective throughput even without dynamic batching. So, by that definition, fast tokens do not necessarily mean inefficient tokens. It doesn't seem that way to me.
I think it depends on the hardware design. As they say, it is expensive: HBM is caviar, while SRAM is even more expensive. I don't understand the energy-consumption implications. That will be interesting.
Another concern is that many alternative chip startups, without naming names, rely on general-purpose GPUs for prefill. That is fine, but your argument is that chips are limited and people need more of them. If a chip cannot run efficiently, though, it still seems that you need an NVIDIA GPU. I don't know.
Yes, but I think the meaningful setup is to have prefill and decode as a pool. As the pool size increases, the traffic can be divided between prefill and decode in a disaggregated setup.
If you know the prefill-to-decode ratio for the entire data center in advance, then that is absolutely possible. But that is a problem, because the workload is going to change, and inefficient silicon will go to waste. If the workload pattern changes and you need more decode or less decode, you have to account for that.
Chips are monolithic. If you can change the setup, that makes more sense.
In all circumstances, if you rely on NVIDIA, one of your partners has to be a GPU provider. If you buy NVIDIA GPUs and more of them, business decisions have more influence on your ability to win than technical decisions.
NVIDIA is clearly taking over. All of these chip startups need some technical justification for what they are doing. Otherwise, why would someone own $20 billion worth of chips that are not operating? It is impossible to build something that works for less than the price of an NVIDIA GPU. Anyone else can do the same thing.
The question is whether this is technically beneficial to someone, or whether it creates demand signals for NVIDIA. This is a business issue. Do these companies have enough capacity to build and deploy enough chips to run the workloads? NVIDIA's GPU exports are controlled, and data centers want to buy those GPUs and build around them. Stopping that is a very large challenge.
Yes, I think so too. At this stage, chip startups generally say they can tape out a chip. It has been built. It is proven. CPU, GPU, and custom silicon have been constructed by everyone. MTIA and Maia—and even the hyperscalers—have proven that they can build general-purpose silicon.
You don't need to build this specific silicon for decode. If necessary, build a second chip specifically for prefill. Build a prefill-and-decode chip, right? Habana Labs has done it, I think—or at least claims to have done it.
Yes, the design is not free. RTL is much cheaper than it used to be, though. The software moat—NVIDIA's CUDA software moat—takes time, effort, and money to develop. If you spend that time, effort, and money on alternatives, you may still not reach NVIDIA's performance standard.
I hope chip startups will do the pre-silicon work and then build silicon for decode. Whether they work with NVIDIA, Broadcom, Marvell, or another vendor, they will have to figure out the trade-offs involved in working with multiple vendors and removing some margin.
I hope they can deploy enough of these systems and produce enough chips, but that will be very difficult because distribution, networking, and data center operations are extremely important.
But there is excitement, man. During the conversation that somehow got buried, I know, TPU inference, for me, is the most exciting, as a matter of fact. Because it is a one-way street. It’s like, did you know? They are TPUs. They are now starting to sell externally, and then they’ll say, “We stopped everything for us to keep going.”