Why AI is Unbundling Faster Than You Think with Tarun Chitra | Ep 164
Tarun Chitra sees crypto’s frontier narrowing from world-changing infrastructure to “TradFi plus,” but not necessarily its addressable market. If stablecoins rise from roughly $300 billion to $1 trillion, he expects lending and trading to grow more than 3x through financial reflexivity. Logan Jastremski’s sharper formulation is that crypto can remain narrowly about finance while on-chain volume still grows 1,000-fold.
The investable crypto exposure is increasingly order flow and execution, not blanket ownership of protocol tokens. Chitra argues DeFi value capture migrated from the “thick protocol” toward wallets such as Phantom and toward MEV and execution, leaving fee switches “stuck in the middle” without durable control of flow. Hyperliquid offers a comparatively legible volume-times-basis-points model; Solana and ETH have much less defined token-accrual stories.
AI agents could change market microstructure by replacing part of passive investing rather than magically beating professional traders. Instead of buying a uranium ETF, a user might state a thesis and let an agent construct and continuously rebalance a personalized basket, dispersing trades across a 24/7 market. Chitra estimates that if agent-directed retail flow reached 20-30% of volume, it could materially weaken today’s opening-and-closing concentration and the predictable arbitrage surrounding ETFs.
Chitra’s strongest crypto-AI thesis is sovereign computation and cryptography, not decentralized training for its own sake. He concedes, “I was wrong” that decentralized learning could not work, but still sees InfiniBand and network optimization as structural advantages for centralized data centers. The higher-value opening may be selective FHE, TEEs, ZK proofs, and GPU integrity attestations—privacy as a targeted feature for expensive operations, not a mass-market product people willingly pay double to use.
Open-source AI is unbundling into the same functional layers as DeFi: interfaces resemble wallets, routers resemble DEX aggregators, models resemble protocols, and inference providers resemble liquidity providers. Beneath an apparently simple interface lies an order-flow market deciding “who processes your token, who picks you a GPU, who guarantees the price.” If DeFi’s history rhymes, value could concentrate at the interface and execution edges rather than automatically accruing to the model itself.
Compute is becoming a financial commodity with spot tokens, dated futures, GPU capacity curves, and routing economics. OpenRouter reportedly takes about 5%, while inference providers compete on price, speed, latency, uptime, and hardware; some discount standard pricing by 30-40%, while others charge more for faster Cerebras-based output. Chitra expects nonlinear premiums for networked 4x, 8x, and 16x GPU clusters and sees on-chain markets as a natural venue once hardware can prove what it computed.
Third-party wrappers retain strategic value even if models become capable of generating tools and learning inside enormous contexts. Enterprises do not want to hand their workflows and proprietary context to two model companies, which is why Ramp, Cursor, Databricks, and Palantir are all building routers. Chitra’s bet is that active learning will not eliminate this market because its compute demands are “excessive,” while wrappers can preserve context, sovereignty, and model interchangeability.
1. Crypto has matured into finance before exhausting its volume opportunity
Chitra’s top-level diagnosis is that crypto now resembles “TradFi plus.” The gap between on-chain and centralized trading has narrowed, while the grand questions around L1s, L2s, ZK, scaling, and reliable exchange infrastructure have yielded to incremental work bringing off-chain assets on-chain.
He describes technological progress as an S-curve whose slope now appears to be declining: crypto may be approaching a plateau, although “I don’t think we know for sure.” Gauntlet’s move toward RWA and institutional finance reflects that judgment, as does his admission that he has not felt inspired to write research papers because fewer open technical questions remain.
Jastremski’s pushback—worth keeping—is that a narrower product can still grow parabolically. Hyperliquid, Robinhood’s chain, Base, and perhaps Solana may constitute another “design maze,” while on-chain trading could expand 1,000-fold as real assets arrive.
Chitra’s monetary mechanism is explicit: moving stablecoins from about $300 billion to $1 trillion should generate more than 3x growth in lending and trading. Stable money supports nonlinear financial activity, so DeFi volumes should carry beta to stablecoins, “not to Bitcoin.”
2. Token value accrual is losing to wallets, order flow, and execution
Chitra argues that AI is “destroying” Bitcoin’s economics because the opportunity cost—or risk-free rate—for the average data center has changed. He is equally blunt on ETH: he thinks even ardent supporters in 2026 will be forced to concede that its value-accrual story “simply doesn’t exist,” despite applications burning some ETH.
Solana sits awkwardly between general infrastructure and vertically integrated exchanges such as Hyperliquid. Tokenized shares may be valuable for Solana or Ethereum, but Chitra does not see that usefulness translating automatically into meaningful accumulation for the underlying tokens.
His original attraction to DeFi was its unbundling of investment banking: Maker or Uniswap could function like individual bank divisions, letting users compose only what they needed. DeFi then fragmented again into interfaces, routers, liquidity protocols, LPs, MEV, and validators.
That second unbundling overturned the “thick protocol” thesis. Chitra sees front ends such as Phantom and MEV capturing much of the economics, while protocol fee switches occupy a shrinking middle without control of order flow; Jastremski therefore prefers execution exposure as volumes rise from today’s roughly $5-10 billion toward global equities’ approximately $800 billion.
3. Agent portfolios could dissolve the clock that organizes modern markets
Chitra expects a bot-dominated market to differ fundamentally from conventional HFT. Traditional low-latency competition exists partly because human and institutional flows concentrate near the open and around 3:30-4:00, when ETF and mutual-fund rebalancing forces predictable transactions.
His uranium example carries the argument: an ETF bundles miners, the commodity, transportation, storage, and disposal, but must trade transparently at prescribed times. Investors accept being “eaten by the wolves” through creation-redemption arbitrage in exchange for effortless thematic exposure.
An agent could instead translate “I want uranium exposure” into a personalized portfolio, favoring recycling for one user and extraction for another. Rebalancing would occur when individual signals arrive, scattering liquidity across a 24/7 market and making continuous responsiveness more important than extreme speed during fixed windows.
Chitra connects this to David Easley’s “volume clock,” under which market makers measure time through traded volume rather than wall-clock minutes. If bots operate continuously and personalized flows stop clustering, that normalization may disappear; the resulting market structure may need a name other than HFT.
4. Agents are a successor to passive products, not free alpha machines
Jastremski recalls the late-2024 pitch that autonomous agents would manage portfolios; he says the market had a giant pump, but only the meme coins really pumped, and he was unsure those investments made a profit. His objection was simple: if an agent could reliably outperform, why would firms competing with Citadel or Jump distribute it freely?
Chitra’s answer is that agents replace passive investing, not active management. That framing is commercially unpopular because passive fees are thinner, but it still represents a large market: ETFs went from below roughly 5% of market capitalization around the financial crisis to more than 50%, with more ETF tickers than individual stocks.
The conditional prediction is specific: agent-directed retail trading at 20-30% of volume could change ETF and fund dynamics and disperse order flow throughout the day. Chitra does not expect the extreme endpoint—many regulated, concentrated, or disclosure-bound pools of capital cannot delegate everything to autonomous agents.
5. The winning interface may begin with a spoken thesis rather than a ticker
Chitra expects a new platform for younger users, extending what Robinhood and Coinbase proved with millennials. It may manage a wallet through permissions and safeguards resembling Privy or Turnkey, but he does not know whether its primary interface will be text, visuals, or audio.
His preferred example begins at 1 a.m., with a user in their underwear describing a nuclear-power documentary. The system clarifies that the actual thesis is uranium exposure, proposes the relevant assets, constructs the portfolio, and executes—compressing hypothesis formation, analysis, and trading into one conversation.
Jastremski’s broader question is whether distribution incumbents such as Coinbase, Robinhood, or Stripe inevitably own that interface. Chitra leaves it open: the winner may be a fintech-DeFi hybrid that uses crypto for 24/7 margin and lower costs without caring “what an AMM curve is.”
In action rather than rhetoric, Chitra bets on DeFi practitioners learning to integrate traditional assets. Yet he allows that a less rigid consumer entrant could challenge Robinhood, Interactive Brokers, and Coinbase by combining friendly AI with crypto rails almost invisibly.
6. Crypto’s remaining bosses are sovereign AI and verifiable identity
Chitra identifies sovereign and private AI as one unfinished problem; Jastremski adds verifiable credentials and identity as the second. Neither necessarily requires tokens or blockchains, but both may require cryptographic tools—ZK proofs, FHE, attestations, and related techniques—that crypto helped turn into tested, formally verified products.
Chitra explicitly revises his prior view of decentralized learning: “I was wrong. They managed to do it.” His remaining skepticism is economic and physical—consumer NVIDIA 3000-, 4000-, and 5000-series GPUs outside data centers might total 5-10 gigawatts, while leading centralized labs were already approaching two to 2.5 gigawatts and could keep scaling.
InfiniBand and specialized networking remain the centralized advantage because data movement, encoding, and network topology can matter more than arithmetic optimization. A top-ten OpenRouter model may require about two H100s merely for weights, while the GLM-5.2 example requires perhaps 25 H100 equivalents before substantial context.
Logan calls privacy “a feature,” while Chitra says it is not a product. Users may say they want it but resist paying double; ZK resembles insurance whose value appears only when integrity is challenged, making monetization difficult outside high-value, low-frequency operations such as rotating root credentials or unlocking $500 million of stake.
7. Open-source AI is recreating DeFi’s unbundled market structure
Chitra’s mapping is the episode’s central frame: interfaces or wrappers are wallets, routers are DEX aggregators, models are protocols, and inference providers are liquidity providers. Closed labs package those layers together, even when a session internally routes among specialized models for compression, memory, or generation.
Open-source interfaces such as Hermes expose only the front end; underneath, routers select models and providers such as Together, B10, or Modal according to price, speed, latency, throughput, and uptime. “You don’t see all of this stack,” but its hidden order flow decides who serves each token and allocates each GPU.
Provider competition already resembles prop-AMMs or MEV. A model creator may establish a reference price—Chitra gives the hedged example of Zhipu pricing GLM at “$140 per million input tokens or something like that”—while smaller data centers discount by 30-40% or charge premiums for faster Cerebras-based output.
Under Chitra’s analogy, DeFi’s value-capture pattern could repeat: economics migrate upward to interfaces and downward to inference execution, while routers behave like constrained brokers and open-source models may act as loss leaders for inference demand. Hyperliquid remains an exception because its threat model differs.
8. Compute markets will need wrappers, futures, and cryptographic settlement
Jastremski asks whether increasingly capable models will absorb their own wrappers. Chitra sees durable enterprise demand for independence: businesses are reluctant to surrender automation and proprietary context to two vendors, driving Ramp, Cursor, Databricks, and Palantir to develop their own routing layers.
The technical counterforce is active learning. Chitra sketches a progression from 2023’s pre-training scale, through 2024’s reinforcement learning and task-specific harnesses, toward systems that generate harnesses in real time and learn recursively inside contexts that could expand from one million to 100 million tokens.
He nevertheless bets that third-party wrappers survive because real-time active learning is computationally expensive. They can accumulate organizational knowledge, connect workflows, switch backend models, and preserve sovereignty; over time, wrappers may internalize routing just as wallets internalized aggregation rather than surrendering its convenience fee.
Compute itself could split into spot token prices, dated token futures, GPU capacity indices, and cluster premiums. Direct providers already show nonlinear pricing from 4x to 8x to 16x clusters depending on shared InfiniBand topology, creating an unstandardized “yield curve.”
Chitra’s endpoint is on-chain settlement for compute: a GPU or cluster proves its integrity, cycles, and completed work, then acts as its own oracle. That could replace today’s “handmade” legal contracts and support commodity-style trades between token output and GPU input costs, including divergences caused by failed physical delivery.
OpenRouter is an early specimen rather than the final form: Chitra says it takes about 5%, while Jastremski estimates its earnings at “$40-50 million now.” Private enterprise workloads operate like a dark pool. Jastremski says current task routing remains crude—benchmark matching plus a price ceiling—but the speakers expect better evaluation, futures, and wrapper-router consolidation.
Full transcript
AI is simply killing the Bitcoin economy because the opportunity cost—the risk-free rate for the average data center—has completely changed. So, if the Bitcoin component disappears, ETH will have no real accumulation of value. I think even the most ardent ETH supporters in 2026 will be forced to admit that the concept of accumulating ETH value simply does not exist.
The reason it's similar to DeFi is that the interface for the end user is the same. Whether you use cloud code and your own interface, or Hermes and something custom, this is all you see. You don't see the whole stack. But underneath the stack, there is an ecosystem of order flow: who processes your token, who picks you a GPU, and who guarantees the price.
Open-source models are disaggregating just like the cryptosphere. But what's interesting is that instead of people competing as network nodes, we see entire data centers competing with each other. Because there is a huge amount of capital invested in data centers, hundreds of players are competing to create the next token, but the market structure is very similar to MEV or prop trading. This is just the beginning.
Thank you, Tarun, for coming to the podcast. We haven't seen each other for a long time. I'd like to talk about the state of the market. I know you discuss this every week on The Chopping Block, but maybe at both a higher and lower level, because I feel like the crypto industry has changed a lot. You're a veteran, so I'd like to get your thoughts not so much on Gauntlet, but on the venture capital market in general.
1. Speculative Narratives Fade, Trading & Payments Remain
If you describe the crypto market now, it's like TradFi plus. It seems like we no longer have the big dreams or aspirations that we had maybe 5 years ago, right?
When we were thinking about how to scale blockchains, or even 3 years ago, when the question was how to create a reliable trading infrastructure on the network, there used to be a huge gap between online and centralized trading infrastructure. This gap has narrowed, of course, with certain compromises.
But then the question arose: beyond improving trading and capital efficiency, is there any next big thing? It's not like there are L1s or L2s, all these big innovations, or ZK technologies anymore. It seems like there is nothing like that anymore.
Most things look like gradual but useful optimizations to make it easier to move off-chain assets onto the blockchain, without a significant surge in speculative hype. I don't count all these memecoins. It seems to me like it's just hot capital moving in circles, and I know there's a whole group of people who will ask, "How is that different? Isn't that what the crypto market is all about?" Everything is fascinating, but for me it's not like that.
I think that's a great starting point, because I agree. It was similar to Web3. It wasn't something like internet capital markets, so to speak. It was the metaverse. It was NFTs. I think there is much more restraint now.
The general arc was Web2, and then it became Web3. It definitely narrowed down to Web2.5 or Web2.7, or something in between. The feeling that crypto is now about trading and payments is like the next stage in the evolution of finance.
I was at the Out East conference earlier this week, hosted by TAI, and it was great. But everyone said, "Crypto is finance," and then they discussed artificial intelligence and agents. It looks like this is the barbell approach that is starting to take shape: long on trading and finance on the crypto side, and then potentially long on AI, with perhaps an intersection between the two.
I mean the finance sector, RWAs, and the institutional direction. I think there are a lot of interesting things there. Gauntlet is definitely moving in that direction.
It's just that I haven't felt inspired to write research papers for a long time, because there aren't that many open questions in crypto. That's not necessarily a bad thing, right? This means that the technology is mature, but if I were to think of technological innovation as a sigmoid, an S-curve, are we at an inflection point where we're passing a plateau, or is there more to come? I don't think we know for sure, but it feels like the slope is decreasing. At least, that's how I perceive it.
I think we've kind of explored this design space, so to speak, in some sense. But it was very much like Web2 or Web3, and then we did the same thing with scaling.
We had Ethereum, which had low throughput. Then we had NEAR and sharding. We had app chains, Cosmos, L2s, and L3s with high throughput. But now I think we're doing the same thing in the trading landscape.
You have, for example, Hyperliquid, for better or worse, like AWS Tokyo. Then you have L2s like Robinhood and Base, and maybe Solana, which is globally distributed. If cryptofinance explores this design space, I'm interested in your thoughts, but I feel like if the focus on cryptofinance or trading narrows, maybe there won't be 100 or 1,000 new primitives, but you can still increase the volume of trading on the network 1,000-fold. From that perspective, even though the focus is on trading, you can still have parabolic growth.
2. AI’s Impact on Bitcoin Economics & Data Center Opportunity Cost
If there is $1 trillion in stablecoins, there will be a money-multiplier effect, where the volume of lending should grow nonlinearly. If we go from $300 billion to $1 trillion, that's a 3-fold increase. You expect more than a 3-fold increase in lending activity. You expect more than a 3-fold increase in trading volume because usually, when the money supply expands, these reflexive things grow a little faster, assuming they're stable. Stablecoins are inherently a very stable expansion tool, unlike pure crypto assets.
You would expect the DeFi sector and trading volumes to have a beta relative to stablecoins, not to Bitcoin. That's how I see it.
Another point about Bitcoin: I think AI is just destroying its economics because the opportunity cost—the risk-free rate for the average data center—has completely changed. So if Bitcoin disappears, ETH will have no real accumulation of value. I think even the most ardent ETH supporters in 2026 would agree that the story of ETH accumulating value simply doesn't exist, right?
Robinhood's blockchain is great at burning ETH, isn't it?
There are uses for ETH, but the amount of value that ETH itself gains from this is negligible. As for Solana, it seems to be stuck somewhere in the middle. It doesn't seem like it will be as good as Hyperliquid. People are making incredible efforts to get closer to this, but I'm just not sure whether they'll be able to achieve the goal.
Tokenized shares have definitely become a good thing for them, and potentially for ETH. However, I don't think there is any history of value accumulation for these tokens. Likewise, I believe that DeFi tokens—these fee-switch mechanisms—are not viable in the long term. Some of them may work because they have truly verticalized their order flow and may charge additional commissions, like Phantom.
3. ETH Value Accrual, Solana Positioning & DeFi Token Sustainability
But the interesting aspect of DeFi is that, ironically, I got interested in it in 2018 and 2019 because of Maker, and then Uniswap around November 2018. What interested me about DeFi was that it was breaking up investment banking. Each protocol was like a division of an investment bank, and instead of paying for a subscription to all the divisions of the bank, you could choose the ones you needed at any given time and combine them.
It was interesting because, from the perspective of capitalism, it's just disaggregation and recombination. This disaggregation of finance made it easier to program and algorithmize, which was a very new thing.
But then DeFi itself fragmented even more: into the interface layer, the routing and solver layer, the protocols that held liquidity, the liquidity-provider layer, MEV, and validators. Honestly, I don't think anyone cares much about the validator economy these days.
If we look at it, DeFi has now fragmented into these stacks. If you look at where the value capture is happening, there was this thick-protocol thesis at the beginning, right? The protocol takes away all the value; the periphery receives nothing.
But if you look at it in practice, the front end and MEV side actually take up most of the cost. So this fee switch is kind of stuck in the middle. It tries to capture an ever-smaller share of the balance sheet.
My problem with DeFi tokens is that they have the ability to capture real value for real risk, but they're in a part of the market where they don't have enough control over order flow. I don't know if they can do it sustainably for a long time the way it's going now.
I'm very optimistic about order flow. On any given day, trading volumes are around $5 billion to $10 billion on the network. I think the global stock markets are about $800 billion. If we could get to, say, $1 trillion—it's a big number—I think order flow would become more and more valuable.
I think it's worth betting on order flow and execution. The tricky part, which you mentioned, is the different blockchains and where they are on the trading curve. For example, with Solana, what is the accumulation of value? I think it's obviously less defined than with Hyperliquid, where it's a little easier to say: volume times basis points, minus certain expenses.
Yes. I don't know. It's very interesting. I think we've explored this maze of solutions to some extent. With the scaling and performance of blockchains, our thinking now is that everything is increasingly moving toward the end state, which I think is closer to HFT.
We discussed earlier that I am a big believer in quants coming in and doing more quantitative operations on the network. Because if you go from low-bandwidth networks to much higher-bandwidth ones, all the way to HFT, it will be about 10 GB of traffic. You analyze it, perform algorithmic actions, and trade based on this information.
Yes, I think it will be a little different from HFT in one particular aspect, and that's okay. The last time I was in HFT was in 2017. It's been a long time, but back then, from a market maker's perspective, everyone knew that everyone was using low-latency FPGAs or, in some cases, ASIC algorithms against each other. But there were also slower algorithms, such as statistical arbitrage funds, TWAP engines, or others, which were willing to accept worse execution or longer position-holding times.
If you go up a level higher, there were already people there. In such a market structure, the reason for HFT's existence is different from a market structure where there are no people at all. I think we're moving toward a market where everyone is represented by a bot strategy, and there's not as much speculative human action anymore. If we get to that point, the concept of HFT will change, because HFT has always been based on low latency, because there are traders who, to some extent, have high throughput, right?
They trade in high volumes, but only once a day, around 3:45. If you look at stock volume, most of it is concentrated at the beginning and end of the day, and in the middle, the volume, relatively speaking, isn't that big. I believe this whole paradigm is changing, which means that HFT will become much more like “time trading” than constant quoting. And I think there will be another name for this market structure, at least that's how I see it.
When the model assumes that everyone is a bot 24/7, it changes everything. From 3:30 to 4:00 every day, that's the time when you need to have the lowest latency, because that's when everyone who has to rebalance ETFs, mutual funds, and so on—the ones who actually do it manually and have a fixed portfolio to rebalance—they're all trading during that time. If you're not fast enough, you won't be able to close those deals quickly.
4. Trading Design Space & On-Chain Volume Upside
But in a world where there are no time constraints, where markets are present 24/7, it will significantly change the entire landscape of trade-offs. I don't quite understand how to imagine it, but it seems to me that everything changes without the concept of time. There is such a concept in HFT, and indeed in the economy as a whole.
There is a very classic economic article about high-frequency trading called “volumetric clocking.” The volume clock is more about how people measure time. Market makers measure time differently from takers. Takers measure time in terms of how quickly I can physically fill this large order, which may be larger than what's in the order book.
But market makers do everything from the perspective of, say, a volume clock, meaning they calculate time through volume. So it can take a long time with a huge trading volume. That's kind of the same thing as less volume in a short amount of time, and that equivalence—how you weigh them—is the concept of how much of a market maker you are and how much of a taker you are.
This article by David Easley is from 2007, the eighth year; this was during the financial crisis. But when I worked in HFT, people always referred to the concept of a volume clock—that the speed of trading adjusts to it, and that's how you normalize data across markets. I think that concept has disappeared in this world where everyone is a bot 24/7. So I think it will have a new name.
I like to think in extreme categories. Maybe HFT should be rethought, but to me, the trajectory looks like an increase in blockchain trading if crypto succeeds. Obviously, we started with monkey pictures, then we moved on to memecoins, and now we have stocks. If we continue on this path, hopefully it will be like a black hole where we pull more real assets into the network. I hope that, for me, being long HFT is simply being long on-chain trading volumes.
But I'm curious. I like the train of thought regarding the time horizon, but why do you explain it with the phrase “everyone is a bot”? After all, there are many people in the stock market who are structurally forced to trade at certain times under certain restrictions, like me as an ETF investor.
If I think about the capital flow of someone buying an ETF, let's say I have a desire to get a lot of exposure to uranium. I did some research. I read a bunch about it. I think there are a bunch of new nuclear power plants being built or something, so I want to bet on uranium rising.
The problem is that the uranium trade involves a lot of different things. This could be a play on the commodity itself. This could be a play on mining companies. This could be a play on storage and transportation. This could be a play on—what's the word I'm looking for? Well, when you have spent uranium rods and you have to dispose of them—the disposal side. But that's it: it's all uranium beta, right?
You can say, “Okay, I want all of that, but I don't want to figure out how to build a portfolio, so I'll just buy this uranium ETF that says it's going to buy them all in a certain proportion.” But the thing is, this fund has to trade every day at the end and the beginning of the day, and that bleeds you, you know, arbitrage volume for all the HFT traders who are doing the creation-redemption arbitrage: “I go buy the stocks, I create the ETF,” or “I take the ETF and redeem it for the underlying stocks.”
That means that there are participants in the market who are structurally forced to behave in a certain way, which is predictable, publicly known, and anyone can play to get ahead or rebalance. It has become a kind of Faustian bargain in the US markets: we'll give you this transparency and all that in these passive products, but you're always going to be eaten by the wolves to some extent.
The question is, how worried are you about being eaten by wolves versus what I think is a 10-year trend, versus losing 20 basis points a day or something, right?
In a world where everyone expresses their preferences directly—for example, they have an agent who says, “I want exposure to uranium. Build a portfolio, buy and trade on my behalf”—instead of having an ETF do it, and in a world where all these assets are traded 24/7 on the blockchain, the rebalancing will happen on a personal level, based on each user's preferences, not because they're all forced to buy one ETF and they'll trade at different times.
Suddenly, you lose this notion of concentrated trading volume because people will act in a more timely manner. The agent gets a signal, and the signal says, “Okay, I need a trade.” Your uranium agent may be different from my uranium agent. So earlier, let's imagine that before we both became ETF buyers, our liquidity was pooled and it was a single transaction.
But now everything is getting scattered. Your trade could happen in the morning and mine could happen in the evening, because I like recycling companies and you like extractives, right? The more personalization, the more fragmentation of trading volumes occurs. This changes priorities: speed at certain moments is less important, and the ability to react quickly is more important.
If you imagine that people's individual preferences are coded in this way and a large part of the market becomes like this, then it's a very different idea of the structure of trade than what people are used to now. Logically.
Interestingly, yes, it is on the agents' side. It's hard. It seems to me, again, using the example of the East, everyone says, “We are very optimistic about on-chain trading through agents.” And I wasn't so optimistic about it because it seems to me like it's just a rebranding of AI, when in reality it's more like algorithmic trading, similar to what HFT does.
However, regarding your argument, I haven't really thought about the individual level and how that might change the flows, as well as the transition from a 9:30-to-4:00 regime to a 24-hour regime.
Yes, yes, yes. I mean, the most extreme option is when everyone has their own agent and no one buys ETFs anymore. Nobody invests in liquid funds anymore. They just say what they want, and the agent goes and builds a portfolio, right?
But everything won't be quite like that. There will be a certain portion of the flows that will remain really passive, where they just don't want the agent to take on all that risk—whether for regulatory reasons or because it's a family office that is the majority shareholder and can't trade through the agent without filing a disclosure. So there will definitely be this segment that will not be able to switch to the agents' side.
But it's clear that if all the retail flow suddenly shifts, people will stop buying ETFs and start switching. It's similar to when I was in college, just before and during the financial crisis. The conventional wisdom was: passive trading and ETFs are a bad idea; nobody will buy them.
And that's before ETFs. If you look at the total market capitalization of ETFs, it was still below 5% at that time or so. And today? More than 50%. There are more ETF tickers than actual stocks. This is madness.
These are huge numbers. I think people at the time were saying, “No one will buy these passive products. This is so ridiculous. They will be overtaken by other market participants,” and so on, whatever. But then there was this huge explosion of retail access to trading. There was a demographic that wanted to be active retail traders, and another group that succumbed to FOMO but didn't want to monitor their assets every day.
They wanted to treat it like their retirement account (401(k)), and the ETF became the perfect product that combined both approaches. That is, I can express my opinion and my preferences. I want uranium, but I don’t need to think about it more than once. Isn’t that right? So, I think there is a natural tendency toward a certain amount of passivity and activity, and it is constantly changing due to technology and changes in market structure.
I really think that if retail trading through agents exceeds 20% or 30% of volume, it will completely change the dynamics of ETFs and funds, and then the order flow throughout the day will start to disperse and not be as concentrated. This is the type of thing that I think could happen; it’s more of a microstructure issue. But I don’t think we’ll ever see an extreme version where absolutely every participant is an agent. I just think there are enough pools of capital that are forced not to do it.
I agree. I think it was funny at the end of 2024, when everyone was like, “Okay, agents are going to manage our portfolios.” XBT. Yes. They all went up—a giant pump—and all the hedge funds, but only the meme coins really pumped. I don’t know if their investments made a profit.
5. AI Unbundling Thesis: Open-Source Models vs Data Centers
Yes. Well, actually, that was my thought, because in my imagination it all looked like this: “Okay, this agent is incredibly smart. It will manage your capital and make you money.” And, to me, this is different from what you expressed, namely: “I have my own view of the world. Let the agent express this view or help me find a way to benefit here.”
But the way it was presented, at least in 2024, was that the agent was superintelligent and would make you money. I thought to myself then: Citadel, Jump, and all these high-frequency trading firms spend billions building great models, and if this agent can really consistently make you money, why are they giving it to you for free?
Yes, yes, yes. I think it’s more of a replacement for passive investing rather than active investing. I think that’s the whole difference. And the reason people don’t want to admit it is that the fee structure in passive investing is much weaker than in active investing. So many can’t sell the idea that the total market size of this product is that large if they say it will only replace ETFs. However, in my opinion, it’s still a significant improvement.
So, if this were to evolve into the kind of autonomous-agent worldview that you’re talking about, would that significantly change the current order-flow structure from MetaMask, Phantom, or similar products in favor of agents?
I think this changes the order-flow structure for crypto wallets, but also for traditional financial assets. To some extent, Robinhood and, to a lesser extent, Coinbase are a kind of derivative of the post-crisis mood: “I don’t trust my financial advisor; I want to do everything myself.” That was in the spirit of millennials, right? That era of retail trading was built on this, and everyone in the HFT sector scoffed at it until it started generating so many orders that it became clear: “Okay, now we really have to deal with this.”
WallStreetBets, GameStop.
Yes, yes. Well, that was even later. I’m referring more to 2016, when the Chinese stock market was booming, because that’s when emerging markets were really growing rapidly. Retail trading on Robinhood really took off back then, and of course, the ICO boom of 2017—that’s what really gave them a big boost.
But there is a certain new modality, so to speak, for people under 25 years of age. It can evolve as they grow older, which is essentially a thesis that both Robinhood and Coinbase have proven for millennials. It’s not a given that they will be able to support this for those who are under 25 now.
So I think there will be an interface—a kind of new interface—aimed at very young people who want to get their experience first, and then those who win at it will scale it up and create a network effect for the platform.
I don’t know if it’s the same thing. My suspicion, based on past experience when we moved from super-active fund management to passive funds, is that we’re going through a transformation again where we go back to active investing, but instead of you doing it, the agent does it, rather than you buying the ETF.
In that kind of universe, I think there will be a new platform, a new version of your agent—something that holds and manages your wallet, maybe something like Privy or Turnkey, where there are certain safeguards and security mechanisms. You grant certain permissions up front, and then it acts on its own.
I think there will be a new interface, and I don’t know what it will look like. Will it be text-based? Perhaps. Will it be that you draw what you want, or visually or audibly represent your request? Maybe. I can totally imagine an audio version, right? People talk about what they think their thesis is, and the AI tells them, “Actually, this is what your thesis is.”
You are fascinated by uranium. It could be something like, “I was reading about nuclear technology today, and it seems like this industry is growing a lot. I want to access it. What is the best way to do this?” Your analyst says, “Hey, you should actually just buy a bunch of uranium assets because they’re doing well right now.” And you say, “Okay, put together a uranium portfolio for me and buy it, okay?”
This is a completely different modality from what it was 10 years ago, when it was fashionable to say, “Oh, I followed those steps and then I just bought a uranium ETF, right?” So I think that this process, from hypothesis generation to execution—this pipeline—is going to change, and it seems like the current UX across all of these services is not right.
It will look different somehow. I think it will be something like an interface plus an agent that accepts multimodal data, perhaps audio.
Regarding audio, I used to be very skeptical about artificial intelligence working with sound. I thought, who wants to talk anyway? But now I see that we have devices where we talk and whisper into our microphones. I can imagine you pitching your idea or hypothesis and then saying, “Okay, figure out how to trade this because I want to get a share.”
It’s like you’re talking to your work colleagues. Do you understand what I mean? Like that.
Yes. This is interesting. I think it’s quite interesting right now, especially with Coinbase, and to some extent with Robinhood, where at least there was a thought that traditional new fintech players like Coinbase or Robinhood, through their distribution, would be able to enter the market—or even Stripe—and create things like Tempo simply by bringing in all the users. But somehow that didn’t happen.
It is still unknown what Robinhood could potentially do. But for the most part, there are still crypto enthusiasts who seem to live in their own world. We have our own crypto applications. And then there are traditional players—instead of traditional banks, it’s probably fintech.
It seems like we’re moving toward that same “mullet” model that everyone’s been talking about all along. It’s like KYC on the front end, or some pools with rules or segmented pools, a nice interface, and DeFi on the back end.
Do you think the new players emerging in the venture capital market will be more likely to be crypto enthusiasts, or will they be people from traditional finance who understand crypto and are trying to bridge both worlds?
Well, I think if you analyze my actions, not my words, then I’m somewhere in the middle. I believe that it is DeFi people who will learn and understand how to integrate traditional financial instruments.
I think that’s what Gauntlet is focusing on in the context of TradFi custody and asset acquisition, and giving TradFi people the benefits of crypto that we don’t even think about, like 24/7 margin. This doesn’t really exist in TradFi. Even in your Robinhood app, it’s fake because the calculations only happen the next day, right?
But once you get these TradFi assets and show them, “Hey, it’s easier for your agent to build your uranium portfolio at 1 a.m. when you’re in your underwear watching a documentary about nuclear power and thinking, ‘I want nuclear power,’” I think it works the same way that retail trading apps have worked for the last decade: They’ve given people the freedom to do it anytime. I believe the purpose of crypto now is to show this.
Regarding the confrontation between new players and old ones, in my opinion, the question is still open. I really think there could be a new platform emerging—a competitor to Robinhood, Interactive Brokers, Coinbase, and so on—that can leverage AI in a way that big players, who are too slow or clumsy, can’t because of their rigidity.
But I’m not sure if this company will be crypto-native. This company may use crypto simply because it’s more convenient, but I’m not sure if they even know what an AMM curve is. Do you understand what I mean? There will be some company at the intersection of fintech and DeFi. I don’t know exactly what it is, but some consumer company will be able to attract users through friendly AI and lower costs.
Overall, wrapping up the crypto topic before moving on to AI, do you think crypto will remain primarily about finance and trading, or do you think we’ll see something else?
We are either in too much of a hurry or simply haven’t found the product-market fit for this yet. So, I think this is a great transition to the topic of AI.
6. The DeFi Mapping (Harnesses, Routers, Models, Inference)
I think there are really only 2 “bosses” that the crypto world has yet to truly conquer, where it has either failed or, like DeFi, achieved some success but is still trying to figure out how to evolve further. For me, these 2 bosses are sovereign and private artificial intelligence, which may not need blockchains or cryptocurrencies, but probably needs cryptography, like zero-knowledge proofs or other advanced cryptographic techniques.
So, that’s the cryptographic part of the crypto industry. Another, so to speak, second final boss is the area of verifiable credentials and identity. Again, this can be seen as pure cryptography. WLD is a token, and Worldcoin as a product is now just a memecoin for another business. But I really think that the current discussion between China and the U.S., and the fact that people suddenly realized they were giving all their data to 2 companies—and that businesses were giving their data away as well—is important.
That’s a little different from social media, right? There’s a lot of buzz on social media about you giving away all your data. But why did people use social media?
They used social media to show others who they wanted to be, right? Instagram is the social network with the highest revenue per user in terms of advertising spend. And why is that? Because it’s about aspiration. The user shows you who they want to be, not who they really are. That means you can sell them a product that fills that gap and sell them this illusion.
Here, you tell people a lot more. You reveal your entire chain of thought, and countless ways emerge to take advantage of you. I think businesses will realize this over time. Of course, local models and routers are the first wave now, but there will be another one where decentralized and sovereign solutions become relevant again.
It’s like we started with the stupidest way to decentralize. Decentralized learning was simply impossible from a feasibility perspective. I didn’t believe in decentralized learning because I thought it wouldn’t work. I was wrong. They managed to do it.
My conclusion was this: all we saw from the training side was that it takes an incredible amount of computing power to build a state-of-the-art model, and data—and data for monetization—because learning is a loss-making product for the sake of inference.
Yes.
Therefore, if you can create a high-quality enough model, you can monetize it through inference. I kept thinking, “Okay, if the crypto aspect of coordination really works, how many GPUs are there outside of data centers that we could leverage?” I think I asked Rock, “How many potential gigawatts do we get if we add up all of NVIDIA’s 3000, 4000, and 5000 series consumer GPUs?” He said somewhere between 5 and 10.
If you look at the leading laboratories, like xAI and Colossus, they’re approaching 2, maybe 2.5 gigawatts. I thought, “Okay, let’s optimistically assume that there are 10 gigawatts.” If you get 50% of that, it will be 5 gigawatts. Elon will probably reach 5 gigawatts in a few years, and then how do you go beyond that? I thought it was complicated, but I’m pretty skeptical about their combination.
At the same time, I believe in cryptography, because I think that’s why I doubt that distributed systems are the main value here. The value of the cryptographic side seems much higher, including things like FHE, where you can potentially provide encrypted data for the model, so to speak.
Oh, you mean FHE?
Totally, yes—sorry, FHE. You can pass models and perform certain actions on encrypted data. I think another thing is that inference workloads are extremely regular, and you can optimize where to apply expensive cryptography. It’s possible to configure the pipeline and use cryptography only for certain parts, providing certain statistical guarantees.
I think this is the part where, of course, there are people working on it, but it seems like it’s still early days. If you look at the performance and usage curves, they’re not great yet, but this is a case where 1 algorithmic improvement is enough to make everything very fast.
And what about NVIDIA TEE? I’m not saying TEE is the best, right? They get hacked all the time, and there are many reasons why they are structurally built to be hacked all the time, but that’s another story for another day. There is a certain sense in which you can run many models—not advanced ones, but Kimmy 2 or, say, GLM 5, not 5.2—quite easily on Nvidia T with a performance reduction of about 30%. It’s not even that much.
So I think there’s a lot going on in that direction. And yet I believe that InfiniBand, high-speed specialized networks—all of these are what are really killing decentralized systems. When I worked with ASICs from 2011 to 2016, for example, placing ASICs on the network and encoding and decoding data gave a much higher performance gain than trying to optimize the computation.
Network optimization and computational optimization are completely different things, and decentralized systems will never be able to tune the network correctly. You keep trying to do it over the WAN. That’s why I was skeptical at first. Regarding decentralized learning, I thought the network overhead would be crazy.
I started studying the operation of data centers. I felt like I had ignored AI for too long and was mainly trying to understand the basic building blocks of how it all worked from a hardware perspective. I opened OpenRouter and thought, “Okay, let’s take the 10 most popular models. How much hardware do we actually need?”
The general argument against AI was that you’d just run it all locally on devices, there wouldn’t be any capital expenditure on AI, and it would all just disappear. The least intelligent model in the top 10 on OpenRouter seems to require about 2 H100s just for the model weights, and that’s without even considering the longer context. For something more advanced like GLM-5.2, you need probably 25 H100 equivalents just to run a model with minimal context.
So I’m optimistic about AI and equally optimistic about crypto. To me, crypto is long volume, and AI is obviously long intelligence, but I feel like the crypto-AI intersection is cryptography because we’re putting so much money into cryptography.
Yes. It’s not monetized at all in decentralized systems, because the value of a ZK proof is only realized when I need to challenge it. It doesn’t create a constant flow. If you think about a DEX transaction or a loan issuance on the network, they create a transfer of value in every block in which they exist.
The problem with a ZK proof is that it’s like insurance: it has zero value and zero revenue unless there is a successful appeal, and then it’s, “Okay, the integrity was breached; the evidence was useful.” So it’s difficult to set a price for it.
What premium do you charge for something like this? It’s like evaluating a rare event that never happened, right? This isn’t like insurance for something that happens with a certain known frequency. So I think historically this has been a real problem for advanced cryptography in terms of value capture.
I think people certainly want privacy online, but then I just don’t believe people actually want it.
It’s a bit like privacy. I think it’s a feature.
Yes. In my opinion, it’s not a product.
Yes, exactly. It’s like privacy on social media. Revealed preferences and stated preferences are extremely different.
Yes. I find it hard to imagine that this is important enough to pay for. Are people willing to pay double the transfer fee or double the cost of the loan?
No.
Right. So for me, it’s inherently related to the fact that it doesn’t scale. This will only be useful for the 5 “whales” who make 1 transaction per month. Then there is no network; it’s just, like you said about the feature, that’s what the feature is, right?
That was my difficult part. I support privacy, I support network privacy and encryption. I know everyone is generally excited about Zcash, but the cryptographic stuff that they’re doing there is hard for me to be excited about. Maybe if you think of it as the new Bitcoin with privacy, then I understand that, but otherwise I find it difficult to get excited about it.
I believe that the cryptocurrency industry has become a kind of DARPA for advanced cryptography. It has helped transform academic developments into ready-to-implement, tested, and formally verified products. Just as people found commercial applications for GPS that were radically different from its original purpose, even though it took 20 years, I expect others to start using these cryptographic tools in other areas.
For example, Cloudflare uses a bunch of ZK credentials in its work, which were borrowed from open-source projects in the crypto sphere. So examples of use already exist. I think we just rushed a little bit, trying to find ways to monetize them right away.
I came across a short excerpt on Twitter where one of the founders of Google was talking about Zcash. He said that zero-knowledge proofs are an extremely interesting technology, and that this cryptocurrency was the first to implement them. I found it interesting that the founders of Google followed the development of ZK.
Yes. No, I mean that many large companies use ZK inside their systems to synchronize data between data centers or keys. They exchange keys and get proof that they both used entropy correctly and that key rotation was successful. It’s almost like a blockchain, only between their 3 nodes.
This is different from a fully decentralized environment because they require integrity checks, not just regular hashes. They need to know the full history of the origin of the computations to ensure that no attacker has seized control of the data center credentials.
These technologies exist, but they are mostly high-value, low-frequency use cases. This is where, in my opinion, the problem with monetizing most privacy-protection technologies lies. They are ideal for very wealthy individuals or organizations that need to perform an extremely expensive and important operation, such as changing the private key to a data center’s root credentials.
This is as valuable as, say, unlocking my $500 million in staking that everyone is laughing at right now. But how often will you do this? It’s never going to be something high-frequency.
So, the question for me is this: What part of the AI stack will have high privacy demand, and where? How do you limit that computationally so that it's appropriate? I still haven't seen a good answer to that, but I think there will be something there.
So, it sounds like you're more optimistic about crypto-native primitives that have been designed for AI use cases than perhaps about intersections beyond the agent side that are unpacking the AI stack.
7. Agents, Preference Expression & Unbundling Traditional Products
Yes. One thing I will say is that the best digital asset is computing power, right? Bitcoin was the first way, in a sense, to directly monetize computing. Now, was it effective? No. But it worked at doing that, didn't it? It's still working at doing that.
A lot of the market-structure activity that's emerging around compute trading, building all these derivatives on compute, warehousing, and reconciliation—I think we're actually doing this podcast because I wrote that tweet that was a little bit cryptic. That's why I opened with the one about the DeFi guide to OpenAI market structure: bindings are wallets, routers are DEX aggregators, models are protocols, and inference providers are liquidity providers.
Yes. And it was a great tweet. So, historically, thank you.
Closed models sell you all of this as one unit, right? You go into Claude, and it's a binding. Every session you enter is routed. In fact, internally, at Anthropic, OpenAI, and others, they already do some routing between multiple models rather than just using one.
For example, compression—when you go beyond the context window, it compresses and gives you a memory hint about the previous thing—is often handled by a separate model that's smaller and can run faster than the main one, because it's really more about generalization. This is a specific task, so you can train a smaller model. Internally, in many of these advanced labs, when you talk to a model, you're actually talking to a set of models, and some of them are specialized for certain tasks to make things faster.
In a way, that's what routing is. Of course, they find the computation for you, and that's how you get this verticalized product.
The open-source stack, just like in crypto, is all about unbundling. Like Linux and the cryptosphere, the point of open-source software is that you share it, allowing anyone to run the same code, and then you have a way of verifying it. The strange thing is that verification is missing here, and this is probably one of the areas where I think cryptographic proofs will be very useful in the long run.
What happens is this: You have an interface, say Hermes, that works like cloud code, but it routes requests to the heap better than OpenRouter, by the way, or OpenClaude. I tried to get OpenClaude to work, got very frustrated, switched to Hermes, and it worked right after installation. It was wonderful.
Yes. Well, I think OpenRouter's metrics also indicate that everyone just moved there.
You have an interface, and some people have specialized interfaces, like for robotics or something like that. Some use a universal one, like Hermes. There is Klein, and there is OpenCode. There are a huge number of them.
These interfaces, in my opinion, were created completely independently by the AI open-source community. Well, to a certain extent. I mean, the Hermes team was in the cryptosphere. OpenRouter was also in the cryptosphere. They actually draw a lot of inspiration, and in the case of OpenRouter, for example, they have a lot of people from the crypto industry on their team. So they really understand the DEX-aggregation model very well, in a way that no one else in Silicon Valley could, because they all hate crypto.
What's interesting is that you have a third party that routes your request to a model that meets certain criteria. Maybe it's the cheapest option, or maybe it's the fastest. For example, I want this to run on Cerebras rather than NVIDIA because I need a faster token-generation speed.
Then the router forwards the request to the hardware vendor—the inference provider, like Together, B10, Modal, etc.—who executes the model for you. The reason it's similar to DeFi is that the interface for the end user remains the same. Whether you use cloud code and a proprietary interface, or Hermes and something custom, this is all you see. You don't see all of this stack.
But underneath the stack, there is this ecosystem of order flows: who processes your tokens, who picks up GPUs for you, and who guarantees you a price. Open-source models are unbundled just like crypto is unbundled, and this market is growing extremely quickly.
What's interesting about this is that I actually see the on-chain world as the right way to value these derivatives and support their trading 24/7, because it's the perfect digital asset. For example, a GPU with proper integrity checks—which may require modern cryptography—can tell you exactly how much computation was spent creating that token. Once you have that as an oracle, you can estimate the cost aggregated across all GPUs and all data centers.
The interesting thing that's very similar to crypto in the OpenRouter structure is that if you look at the list of providers—for example, if you select a particular model—you'll see a bunch of providers. These providers are actually something like prop AMMs. You have an order book where they tell you the price for the input token and the price for the output token, and then OpenRouter gives you some quality metrics.
They give you latency, bandwidth, transfer rate, and some idea of uptime. Their routing algorithm is, of course, a bit simplistic, but it takes into account price and also their quality. It directs you to them depending on how well they work.
There's an interesting thing you can notice: Some data centers, usually newer, smaller ones, actually compete on price, offering 30%-40% less than the standard price. The standard price is, at least for Chinese models, how they make money: They sell inference for their own model. So they're the ones who set the first price.
I mean, I'm Zhipu, I created GLM, and I say it's $140 per million input tokens or something like that. Then you see all these third-party inference providers joining in and setting different prices. Some will be even more expensive because they say, “We optimized the kernels. We work on Cerebras. We work a little faster than the creator of the model.”
Some will be cheaper because they're simply trying to get a stream of orders from the router and beat the routing algorithm. So this part is the crypto part.
What's interesting is that instead of people competing as nodes in the network, it's data centers competing. Because of the huge capital investment in them, hundreds of these players are competing to create the next token.
I'm simplifying a lot, because this isn't exactly crypto, where the asset is always fungible, right? Here, access to an H100 can vary, power consumption can vary, and there are many factors of non-interchangeability. But much of this market structure closely resembles MEV and prop trading.
This is just the beginning. The only research I'm working on right now is trying to understand, formally and based on the data that I can get publicly, what parts of this market are actually supposed to replicate what we're seeing in crypto.
We talked about value capture in DeFi earlier, remember? Protocols in the center, wallets here, and then MEV solvers and validators below. You've seen that value capture has moved over time from protocols, where it began, to wallets, and down to execution. Hyperliquid is an exception, but there is a different threat model, so let's put it aside.
If we take the same analogy and ask, “Hey, is there enough mathematical similarity to the crypto market structure to expect the same thing?”—if that's the case, then all the cost goes to the interfaces, such as Hermes, or to the inference providers, along with the model and the router.
The routers in DeFi do capture some of the cost, perhaps even more than the protocols themselves, but they're more like brokers. Their pricing can't go up, in a sense. This is the part with OpenRouter: The protocol is the model developer, the model itself.
Your idea of a model that works as a loss leader for the public good to get demand for inference is what the open-source model is a bit like. So the question is: If you think the structure of the crypto market will move here, then that will tell you where the value will be concentrated.
I like this breakdown. That's very interesting. In terms of hardware, I really liked the Hermes. My main question goes back to your previous point about computation.
I think we've seen that you just keep scaling the computation. Models are getting smarter. Of course, researchers and engineers are driving things forward, but to a large extent, it looks like this: “Okay, let's go from 100,000 GPUs to 1,000,000, and then to 10,000,000.” Models are generally getting bigger and more intelligent.
I'm curious: My main question about the wrapper is whether models are getting so smart that they can subsume the wrapper itself, which I don't know, or whether people are demanding a third-party wrapper because they don't want to be tied to a vendor and want to be able to change models on the backend.
The latest thing is, you know, Satya Nadella, CEO of Microsoft, Alex Karp, CEO of Palantir—all these people who were on TV. It was all about the latest developments, which I think is very relevant for business.
This brings us back to the previous thought: Businesses are realizing that they're handing over all the automation of their business processes to 2 companies that can then play for a lead, and the dispute between Figma and Anthropic is the best example of this.
So I think there is a real demand for a third-party shell and router that allows you to remain independent. That's why you see all these companies—Ramp introducing a router, Cursor introducing a router, Databricks introducing a router, Palantir introducing a router—everyone has their own now. So I think this is the inevitable path: confidentiality without advanced cryptography will be ensured precisely through a third-party shell.
8. Real-Time Harness Generation & Active Learning
On the other hand, there are rumors about what the next evolution will be after reasoning models. So let's have a quick history lesson, of course, through the lens of my vision and knowledge—not someone who was closer to the events. In 2023, with Stable Diffusion plus GPT-3, it all came down to the fact that more is better: pre-training. In 2024, it was, “Okay, reinforcement learning and directed reinforcement learning—how exactly do you feed the data after training?”
This in-context learning is actually useful for achieving another order-of-magnitude increase in performance across many different assessments and practical cases, so it makes sense. Then there was the idea of the “harness”: “Hey, actually, if you customize reinforcement learning for different tasks, you can get much better performance than from a generic RL environment.” Then there’s the current meta-level—I don't even know the real word—this meta-harness, something universal that can answer everything, like a cloud-based Codex type code, which arguably killed Cursor. That's why, I would say, they eventually went for the takeover.
This thing could determine which environment to use right at the time of the request, right? But the problem with these environments from a data-consumption perspective is that you now have to pay a lot more for the labeled data to set up different environments. That's why people talk about the cost of data becoming equal to the cost of computing for learning.
To do these specialized RL things, where you have many RL harnesses and choose which one to use, you need a lot more labeled data—a lot more for each specific task domain—to understand how to search in that space. So there's a belief that once one of these models gets good enough, instead of getting data like this and then generating synthetic data through RL, because the whole point of RL is that you're generating synthetic data from the data you have, the model can generate harnesses in real time, right?
It's as if it creates this harness itself. This is something I have also studied. So pre-training is just scaling, and then you do reinforcement learning for specific tasks. You have “on-the-fly calculations” during output, like, “What do you choose: low, medium, high, or ultra?” High.
Then I saw Dario on a podcast where he said, “Eventually we'll have context windows, both during training and just longer ones, that will increase from 1 million to 100 million.” He said, “As context windows get bigger, we'll also be able to do recursive learning directly in context based on your information.”
These active-learning things—I was thinking about this more from a theoretical perspective. There's a famous reinforcement-learning theory, and the question is, “Is this on-the-fly generation of tools equivalent to active learning?” If this is true, then perhaps the concept of a separate toolkit becomes meaningless.
But I believe that for use cases where you want to control your own data and context so it doesn't leak, you will probably have to use third-party solutions. So I think there will be a market for this. There is no doubt about it. The question is, will the growth of this kind of active learning be so dominant that it destroys that market? I'm betting not, because the computational cost of doing so seems excessive.
I don't know if we even have enough capacity for this right now.
Yes. I think the advantage of third-party tools is that you can gradually accumulate a knowledge base and connect to different platforms where everything works much more smoothly. At least, I think so. Apparently, it can be done through Claude and its own tools, but again, you become more dependent.
So the ability to have context and different workflows with your model or your toolkit, while still being able to change the model, is very valuable. And I think the sovereignty aspect is very important here. I find it ironic that American companies operate like a planned socialist economy, while Chinese models are pure capitalism with competition.
But I think both sides will come to active learning. I think if the financing of computing resources becomes efficient—not only routing, but also how I reserve capacity—right now it's mostly a spot market, although everyone is trying to create derivatives exchanges so that you can reserve GPUs for the future.
9. Cryptography, Verifiable Compute & On-Chain GPU Markets
Actually, I think it will be somewhat similar to basis Bitcoin trading. The market will be divided. There will be a market with a spot price per token, a futures price per token, and a market for GPU bands, GPU indices, or some kind of homogenization of these resources.
There will be a certain static premium for tokens, since tokens are like a final product. This is exactly what I need to generate. GPUs are more of an input cost, and they are obviously highly correlated, but there will be times when they will deviate very significantly. There will be a strategic trading strategy, as in commodity trading, aimed at profiting from these divergences.
I think this part can be implemented in DeFi. That's where, again, the beauty of computing comes in: with cryptography—the bare minimum—you can think of the GPU as its own oracle, able to cryptographically prove that it did everything correctly.
Returning to the topic of networks, even in my example with OpenRouter, when you need to combine a network of many GPUs, I'm buying a single GPU, but to me it looks more like a long position, like my MacBook Pro, which potentially has enough RAM to run local models. But if they are not networked, how do you evaluate a standalone GPU, even if there is cryptography? Wouldn't it be easier to evaluate a GPU if it worked separately? But when they are online, how do you evaluate networked GPUs compared to non-networked ones?
Yes, that's a great question. I mean, there's already a big gap visible. If you go to Prime Intellect or direct GPU providers like Lambda, you will see a nonlinear price gap between a 4x, 8x, and 16x cluster; you get something like a yield curve.
How many of them are in one InfiniBand subnet, and how many are not? The more you have, the more you pay. I think there is some yield curve that has not been standardized yet. So I believe that when this becomes standardized, the crypto industry will be a perfect place to trade.
If the cryptography is good enough that I can run a TEE or a ZK proof, and I, as a GPU, can provide proof of integrity, statistics of the calculations used, the number of cycles performed, and exactly what was processed, then the GPU group can act as its own oracle, publish it to the blockchain, and people can make calculations instantly.
Instead, we now have a kind of handmade legal contract, where people trust each other because prices are rising, but it's quite possible that some of these loans will simply explode. In my opinion, this is where the real crypto-AI lies: in moving the trading of computing power to the blockchain.
So it still looks like this: long on crypto trading, long on AI volumes, long on cryptography, and also long on hardware.
Well, I feel like I don't know about you, but I personally have been completely immersed in the crypto market. Then you look up and say, “Oh, there's a lot of interesting things happening in AI.” Of course, the companies developing the models are now valued exactly as they are valued. The question arises: if you missed this trend, how do you continue to bet on the growth of AI?
That's why I'm diving much deeper into data centers and hardware. On the inference side, especially when moving from pre-training to inference, the workload is significantly different. We usually do prefill, which still requires more GPU power compared to decoding.
I would say that OpenRouter is essentially a decentralized output, where the decentralization occurs not at the node level but at the data-center level, with each provider owning part of a data center or a certain number of megawatts. Does OpenRouter make money? Are they like reselling withdrawal services?
They have certain corporate decisions, and yes, they have a commission. Of course. They take about 5%.
Okay, so they have a pretty significant commission. They earn, I think, $40–50 million now. They charge 5% for $100 or $1 of withdrawal. So the commission already exists.
But they also handle enterprise workloads and private workloads that are not shared publicly, so it's effectively a dark pool. I don't know if they want to call it that, but it's something like that.
10. Closing Thoughts: Where Value Accrues Next
Especially now that there are multiple routers and each company makes its own version of one, I think we'll see something very similar to crypto. It will be more like intent solving—that is, combining an operator with an order execution. But there has to be some concept of futures, like in crypto, where the asset remains fungible forever, whereas here you don't have that guarantee.
It's like a corn harvest. If part of the crop fails, it's like, “The refrigeration broke. I can't deliver the physical commodity to you.” Because of that physical-settlement aspect, you need some concept of dated futures, but I think they will be traded. I'm very optimistic about routing in the long term.
We can wrap up if we need to. Yes, sorry. I understand. Let me just write to this person and apologize.
Yeah, let me know how much more you can talk, or we can wrap up. We can finish soon.
So, briefly about routing: I'm optimistic about routing and essentially “downgrading,” so to speak, to less intelligent models for the sake of cheaper tokens. But my main question is: how do you conduct an assessment for specific tasks? Because otherwise, how do you understand which model to choose for a particular task?
In my opinion, the most interesting thing is that routers don't do much in this regard. They take the benchmark results, match your query to a group of tests, and simply choose the model that is best, provided it costs less than X. They operate very much by the rules of the models. It's not that great, as far as I can tell. I think it's definitely going to improve; I think that part of the stack is going to change, but that's going to happen when futures come out.
Does it live in a wrapper or in OpenRouter?
I think in the long run it probably lives under wraps. I think the analogy with DeFi is Uniswap and Phantom. Sure, they could go pay Matcha or 1inch or—sorry—whatever the Solana DEX aggregator is called: Jupiter, or whatever. Jupiter, Titan, whatever. Yes, all these guys.
Or they just create their own solution internally, and in ETH wallets, many already do their own routing. That's just the wrapper analogy, right? After all, they don't want to lose that value, especially if their entire advantage is charging a convenience fee.
Mhm. So I think they are merging to some extent, and that's why we see enterprise routers, right? Because this is the fusion of these 2 things. Interesting.
Well, Tarun, I'm very grateful. We had a wide-ranging conversation—from the evolution of crypto to trading infrastructure and all the cool things happening with artificial intelligence. So I am very grateful. Very interesting conversation. Thank you for inviting me.