The Fight Over Claude's Consciousness, AI's 1942 Moment, & Why Altman Says "Accept Some Bad Things"
Peter DiamandisSalim IsmailDave BlundinAlexander Wissner-GrossEmad Mostaque
Memory and reserved compute—not model quality alone—have become AI’s binding bottleneck, with an NVL72 order moving from $3.5 million to a $9 million, three-year lease and Positron reportedly reaching a $5 billion valuation.Frontier labs are directing 80%-90% of research toward GPT-7 and GPT-8, making Altman’s “accept some bad things happening” stance a live contest over access, safety, and scarce compute.
Ep. 035 - Tech DD’s, Performance Projections, Supply Chain, Investment Thesis (Consulting)
SemiAnalysis Consulting is turning bespoke AI-infrastructure diligence into scalable products, including a token-supply simulator linking chip deployment, workload allocation, throughput, and token pricing.NeoCloud underwriting centers on contractual survivability: 99.9% availability expectations, sub-90% performance, and mismatched 15–20-year leases versus five-year GPU contracts can threaten financing and trigger remedies, while capacity timelines determine whether clients build or lease.
Compute Is A Trillion-Dollar Market Trading In Group Chats | Brett Harrison & Andrawes Bahou
An approximately $1 trillion compute market still trades through calls, texts, Slack, and WhatsApp, while roughly 500 GPU-service providers and opaque brokerage chains create heterogeneous supply and “phantom power” without a standard forward curve.Architect is pursuing dated futures and options, with billing-system benchmarks, as CFTC reviews index reliability, but 20–25% financing costs expose unresolved credit risk.
Crusoe CEO: Why Everyone Gets GPU Depreciation & AI Energy Costs Wrong
Harry StebbingsChase Lochmiller
Crusoe says AI infrastructure will follow cheap, available power, targeting roughly 200 MW in Abilene within one year versus 2.5 years elsewhere.Vertical integration cut medium-voltage delivery to 28 weeks from a 100-week quote, while five-year GPU contracts and managed services diversify cash flow.Rising Hoppers usage prices challenge simple depreciation assumptions, but labor, community disruption, and uncertain IPO timing remain risks.
Why Memory Is AI's Biggest Bottleneck with Vikram Sekar | EP 170
Memory—not raw compute—is AI’s deepest systems bottleneck, making lower-memory architectures a potential lever across bandwidth and networking.As 144-GPU systems push power and cooling higher and copper fails at 400 gigabits, optics, DRAM-on-logic and specialized inference architectures may gain relevance, while Anthropic’s S1 points to more than $500 billion in obligated capital and compute and possible over-ordering.
Ep. 034 - The Fight for Fast Tokens, TPU v7, Vera Rubin, and Engrams (AI Supply Chain, InferenceX)
Jordan NanosCam QuiliciBryan ShanAlec Ibarra
Engrams co-design model architecture and memory hierarchy, retrieving 2-gram and 3-gram embeddings from DRAM or SSD.Hybrid Engram-MoE layers outperform either endpoint, while AgentX traces show 100,000 input tokens at P50 and cache-hit rates above 98%.DRAM offload matters for hundreds of agents at 40–60 tokens/s; TPU V7's early Pareto position still leaves software, utilization, pricing, and demand unresolved.
Why AI Leaders Have Changed Their Minds About AI Safety, Elon on UHI, Anthropic’s IPO
Peter DiamandisSalim IsmailDave BlundinDr. Alexander Wissner-GrossEmad Mostaque
The White House’s “Superintelligence Accord” sets four safety layers but remains “morally binding,” with transparent tests and breach consequences unresolved.Anthropic’s proposed IPO combines $4.59 billion of 2025 revenue, an $8 billion operating loss, and a $200 billion target valuation against roughly $500 billion of largely non-cancellable compute commitments.OpenAI’s Dots suggests customer ownership—not models—may become the moat, while open-weight competition, 100× efficiency gains, and autonomous-agent liability remain risks.
The Future of Frontier Model Architectures with Walter Goodwin, Founder & CEO of Fractile
Fractile is betting frontier inference will be constrained by memory bandwidth and cost, not compute, shifting from SRAM to a DRAM architecture expected to be fully operational in the second half of next year.Goodwin claims 25 times more bandwidth per chip than an HBM-based design, potentially making sparse MoE models economical, while its 150-person team targets a three-to-six-month lead; foundry cycles, ramps, and three-to-five-year amortization remain constraints.
Recursive Language Models — Alex Zhang, MIT PhD
Alex Zhang sees openings for smaller AI labs in overlooked research bets and output architectures, with JEPA potentially cutting binary-classification inference costs by 400 times while expert verification remains scarce.RLM and Prime Agent replace trajectory-as-a-prompt loops with persistent state and code-native delegation, but reported swarm economics—10,000 agents, 88 hours, 130 billion output tokens, and about $40 million—make orchestration and reliability the key watchpoints.
We’re Reaching the Physical Limits of Chips
AI’s investable bottlenecks are shifting toward power, memory, cooling, and data transfer as Moore’s Law nears a physical plateau.Micron, SK hynix, and Samsung control 95% of memory, whose price rose 700% this year as hyperscalers absorbed supply.Photonics could relieve several constraints, with co-packaged optics expected beside GPUs within five years as memory pressure pushes specialized hardware beyond ever-larger models.









