[BidClub_]
307 episodes2 active
Language
MoonshotsEN · 167 min

The Fight Over Claude's Consciousness, AI's 1942 Moment, & Why Altman Says "Accept Some Bad Things"

Peter DiamandisSalim IsmailDave BlundinAlexander Wissner-GrossEmad Mostaque

Memory and reserved compute—not model quality alone—have become AI’s binding bottleneck, with an NVL72 order moving from $3.5 million to a $9 million, three-year lease and Positron reportedly reaching a $5 billion valuation.Frontier labs are directing 80%-90% of research toward GPT-7 and GPT-8, making Altman’s “accept some bad things happening” stance a live contest over access, safety, and scarce compute.

SemiAnalysisEN · 35 min

Ep. 035 - Tech DD’s, Performance Projections, Supply Chain, Investment Thesis (Consulting)

Jordan NanosAbhilash Jain

SemiAnalysis Consulting is turning bespoke AI-infrastructure diligence into scalable products, including a token-supply simulator linking chip deployment, workload allocation, throughput, and token pricing.NeoCloud underwriting centers on contractual survivability: 99.9% availability expectations, sub-90% performance, and mismatched 15–20-year leases versus five-year GPU contracts can threaten financing and trigger remedies, while capacity timelines determine whether clients build or lease.

20VCEN · 61 min

Crusoe CEO: Why Everyone Gets GPU Depreciation & AI Energy Costs Wrong

Harry StebbingsChase Lochmiller

Crusoe says AI infrastructure will follow cheap, available power, targeting roughly 200 MW in Abilene within one year versus 2.5 years elsewhere.Vertical integration cut medium-voltage delivery to 28 weeks from a 100-week quote, while five-year GPU contracts and managed services diversify cash flow.Rising Hoppers usage prices challenge simple depreciation assumptions, but labor, community disruption, and uncertain IPO timing remain risks.

FrictionlessEN · 73 min

Why Memory Is AI's Biggest Bottleneck with Vikram Sekar | EP 170

Logan JastremskiVikram Sekar

Memory—not raw compute—is AI’s deepest systems bottleneck, making lower-memory architectures a potential lever across bandwidth and networking.As 144-GPU systems push power and cooling higher and copper fails at 400 gigabits, optics, DRAM-on-logic and specialized inference architectures may gain relevance, while Anthropic’s S1 points to more than $500 billion in obligated capital and compute and possible over-ordering.

SemiAnalysisEN · 63 min

Ep. 034 - The Fight for Fast Tokens, TPU v7, Vera Rubin, and Engrams (AI Supply Chain, InferenceX)

Jordan NanosCam QuiliciBryan ShanAlec Ibarra

Engrams co-design model architecture and memory hierarchy, retrieving 2-gram and 3-gram embeddings from DRAM or SSD.Hybrid Engram-MoE layers outperform either endpoint, while AgentX traces show 100,000 input tokens at P50 and cache-hit rates above 98%.DRAM offload matters for hundreds of agents at 40–60 tokens/s; TPU V7's early Pareto position still leaves software, utilization, pricing, and demand unresolved.

MoonshotsEN · 145 min

Why AI Leaders Have Changed Their Minds About AI Safety, Elon on UHI, Anthropic’s IPO

Peter DiamandisSalim IsmailDave BlundinDr. Alexander Wissner-GrossEmad Mostaque

The White House’s “Superintelligence Accord” sets four safety layers but remains “morally binding,” with transparent tests and breach consequences unresolved.Anthropic’s proposed IPO combines $4.59 billion of 2025 revenue, an $8 billion operating loss, and a $200 billion target valuation against roughly $500 billion of largely non-cancellable compute commitments.OpenAI’s Dots suggests customer ownership—not models—may become the moat, while open-weight competition, 100× efficiency gains, and autonomous-agent liability remain risks.

No PriorsEN · 36 min

The Future of Frontier Model Architectures with Walter Goodwin, Founder & CEO of Fractile

Walter GoodwinSarah Guo

Fractile is betting frontier inference will be constrained by memory bandwidth and cost, not compute, shifting from SRAM to a DRAM architecture expected to be fully operational in the second half of next year.Goodwin claims 25 times more bandwidth per chip than an HBM-based design, potentially making sparse MoE models economical, while its 150-person team targets a three-to-six-month lead; foundry cycles, ramps, and three-to-five-year amortization remain constraints.

Latent SpaceEN · 101 min

Recursive Language Models — Alex Zhang, MIT PhD

swyxVibhuAlex Zhang

Alex Zhang sees openings for smaller AI labs in overlooked research bets and output architectures, with JEPA potentially cutting binary-classification inference costs by 400 times while expert verification remains scarce.RLM and Prime Agent replace trajectory-as-a-prompt loops with persistent state and code-native delegation, but reported swarm economics—10,000 agents, 88 hours, 130 billion output tokens, and about $40 million—make orchestration and reliability the key watchpoints.

SourceryEN · 62 min

We’re Reaching the Physical Limits of Chips

Molly O'SheaAnnie Lamont

AI’s investable bottlenecks are shifting toward power, memory, cooling, and data transfer as Moore’s Law nears a physical plateau.Micron, SK hynix, and Samsung control 95% of memory, whose price rose 700% this year as hyperscalers absorbed supply.Photonics could relieve several constraints, with co-packaged optics expected beside GPUs within five years as memory pressure pushes specialized hardware beyond ever-larger models.

SourceryEN · 33 min

The $10T AI Buildout Has a Photonics Problem

Molly O'SheaHerwig Van HoveYannick De Koninck

AI’s next bottleneck is the interconnect fabric: models no longer fit on one GPU, while agentic calls make latency critical, putting photonics alongside compute as core infrastructure.Optical bandwidth addresses copper and power constraints, but lasers, tools, substrates, throughput, and yield remain scarce as NVIDIA’s demand shock runs ahead of supply; Themaa targets 2027 production and 2028 full ramp.