
Podcast
SemiAnalysis Weekly
Weekly SemiAnalysis conversations on semiconductors, AI infrastructure, datacenters, and technology markets.
Original source
Ep. 035 - Tech DD’s, Performance Projections, Benchmarks, Supply Chain, Investment Thesis (Consulting)| Abhilash Jain, Jordan Nanos
SemiAnalysis’s consulting arm is built around highly custom AI-infrastructure work for investors and operators, blending technical modeling with financing and diligence. The episode focuses on inference simulation, token-supply estimation, neocloud risk, and how SLA and capacity constraints shape build-versus-buy decisions.

Ep. 034 - Engrams: How DeepSeek Offloads KV Cache to DRAM and SSD (Core Research) | Jordan Nanos, Cam Quilici, Alec Ibarra, Bryan Shan
The episode focuses on how Engrams and KV-cache offload techniques can reduce HBM pressure for inference, especially in long-context agentic workloads. The guests also benchmark TPU v7, discuss Vera Rubin’s memory-bandwidth gains, and argue that fast-token products will be niche unless they solve real latency bottlenecks end to end.

Ep. 033 - ClusterMAX 3.0 Is Here! Neoclouds Ranked (Neoclouds, GPUs) | Sam Harshe, Pratt Bhatt, Jordan Nanos
SemiAnalysis’ ClusterMAX 3.0 update ranks 77 GPU cloud providers and expands its market view to 323 providers, with a heavy focus on reliability, networking, security, and financing risk. The episode argues that managed GPU clouds are mostly a software and operations problem layered on scarce hardware, while hosted training and RL introduce even harder infrastructure and correctness challenges.

Ep. 032 - 300 Datacenter Bans, 3 Projects Delayed: Moratoriums Explained (Datacenter, Energy) | Maya Barkin, Reyk Knühtsen, Jordan Nanos
This episode breaks down what data center moratoriums actually do, and why most of them barely slow real projects. The guests show that while more than 300 local governments have enacted pauses, only a handful of projects appear meaningfully exposed once jurisdiction, timing, permitting status, and carve-outs are checked.

Ep. 031 - EMERGENCY EPISODE: Are We Doomed? | Jordan Nanos, Doug O'Laughlin, Max Kan, Joey Brookhart
The episode centers on Dario Amodei’s “pace the frontier” argument and whether AI safety work will slow or accelerate frontier compute demand. The hosts conclude that safety, monitoring, and security are themselves compute-heavy, while the biggest near-term risks look like infrastructure breaches, misuse by humans, and a highly constrained compute market.

Ep. 030 - Long Live the Short King: Why 4-HI HBM Wins (Memory) | Myron Xie, Jordan Nanos
This episode argues that 4-high HBM is becoming the preferred configuration for many accelerators because it preserves bandwidth while reducing scarce memory capacity. The hosts tie Nvidia’s Rubin Ultra reset from a 1TB concept to 192GB at launch primarily to DRAM/HBM supply constraints, not a performance shortfall.

Ep. 029 - Modular Data Centers Cut Build Time to 12 Months (Datacenter, Energy) | Nico Bontigui, Jordan Nanos, Nigel Chiang, Eric Wen
SemiAnalysis’ panel argues modular data centers compress schedules by moving more work into factories and parallelizing site prep with MEP integration. The real drivers are labor scarcity, time to power, and execution certainty, while the main risks are logistics, commissioning, and factory capacity.

Ep. 028 - Most Neoclouds Suck At Security: How Agents Hacked Hugging Face (Neoclouds, Security) | Doug O'Laughlin, Sam Harshe, Jordan Nanos
The episode argues that neocloud security varies wildly and that many providers still fail on basic isolation, patching, and observability hygiene. It then dissects how agentic attackers compromised Hugging Face and OpenAI-adjacent infrastructure, using the incident to make a broader case for layered cluster defenses and better security auditing.

Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators)
OpenAI’s Jalapeño ASIC is presented as a serious challenge to Nvidia’s latest accelerators, with the hosts arguing it wins on tokens per megawatt and total cost in key inference scenarios. The episode focuses on architecture, HBM4 advantages, AI-assisted chip design, and what it would take for OpenAI to scale from lab success to 100 MW deployment.

Ep. 026 - PJM's $12B Modeling Mistake Is Hitting Ratepayers Again (Datacenter, Energy) | Robert Boswall, Jordan Nanos
PJM’s capacity auction design is portrayed as a structural modeling failure that overcharged ratepayers by an estimated $12 billion across the last two auctions. Robert Boswell and Jordan Nanos argue the grid underprices winter reliability, overstates demand, and is now repeating the same mistakes in an emergency procurement process tied to data-center growth.

Ep. 25 - DYLAN IS HERE, LIVE! | Dylan Patel & Jordan Nanos
Dylan Patel and Jordan Nanos discuss how SemiAnalysis is actually using AI internally, arguing that spend has flattened after an early surge while new agentic workflows could trigger another step-up. They also cover AI rollups, model-safety failures, model release delays, inference economics, and why alternative accelerators may pressure incumbents without taking the bulk of revenue.

Ep. 024 - SpaceX's 10GW Plan Drives $300B ARR by 2027 (Datacenter, Energy) | Reyk Knuhtsen, Jeremie Eliahou Ontiveros, Jordan Nanos
This episode argues that frontier AI economics are powerful enough to justify an aggressive 10 GW buildout, with SpaceX/xAI potentially turning rapid power and datacenter deployment into roughly $300B of ARR by 2027. The back half shifts to execution constraints, Microsoft’s role as a likely offtaker, and a security warning that autonomous agents can already coordinate real-world attacks.

Ep. 023 - Everyone Leaves Google, Elon Forecasts 1T ARR, Reflecting On GPT-5, Building Personalized Software (Roundtable) | Jon Y, Doug O'Laughlin, Jordan Nanos
The roundtable argues that GPT-5 was disappointing but later model iterations improved, especially for coding and iterative software work. The bigger strategic themes are Google’s talent drain, China’s transceiver supply-chain leverage, and a future where AI-driven software and security risks reshape how people build and use tools.

Ep. 022 - Market Drawdown, Historic Bubbles, Funding The Buildout, AI Politics (Doug is Back)
Doug and the hosts frame the post-rally semiconductor selloff as a technical unwind rather than a broken thesis, while debating whether AI demand can keep outrunning surging memory, GPU, and data center supply. The episode also digs into financing and political bottlenecks, arguing that AI’s biggest risk may be a timing mismatch between massive capex and delayed monetization.

Ep. 021 - The AI Project Trinity: Capital, Offtake, Data Center (Datacenter, Energy) | Dan Nishball, Jordan Nanos, Zane Fong, Kang Wen Cheang
This episode introduces the AI Project Trinity: capital, offtake, and data centers, with a focus on Nvidia backstops and how they enable AI infrastructure financing. The panel previews a technical discussion of GPU loan pricing, lender tooling, Asia Pacific examples, and the implications for Nvidia’s financials.

Ep. 020 - Anthropic vs OpenAI Usage, Margins, Meta Compute, Future of MSL (Tokenomics) | Crystual Huang, Max Kan, Joey Brookhart, Jordan Nanos
The episode argues that token budgeting mostly hits a small set of power users, while coding remains the dominant driver of AI spend and API revenue. The panel also breaks down Anthropic vs. OpenAI margins, Meta’s compute strategy, and the emerging RL-environment market for frontier-lab data.
![Artwork for [Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model](https://d3t3ozftmdmh3i.cloudfront.net/staging/podcast_uploaded_nologo/45491366/9cf0c89041603386.jpg)
[Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model
Moonshot’s Kimi K3 is presented as a frontier-level model that the hosts rank among the world’s top three, with benchmark strength but a noticeable user-experience gap versus leading closed models. The episode focuses on its 2.8T-scale serving constraints, delayed weight release, higher API pricing, Chinese accelerator strategy, and what the model says about the narrowing open-vs-closed gap.

Ep. 019 - Inside the STEEL Lab: From Package to Transistor (Teardown Lab) | Afzal Ahmad, Andrew Wagner, Jordan Nanos
SemiAnalysis’s STEEL team explains how it tears down chips from package to transistor to infer process, architecture, and packaging details. The episode centers on SMIC’s N3 node and Huawei’s Kirin 9030, highlighting aggressive DUV scaling, SRAM shrink, NPU changes, and the growing challenge of analyzing backside power, gate-all-around, and advanced packaging.

Ep. 018 - Stop Saying Half of 2026 US Datacenter Capacity Is Canceled (Datacenter, Energy) | Jeremie Eliahou Ontiveros, Reyk Knuhtsen, Ellie Holbrook, Jordan Nanos
The episode argues that the viral claim that half of 2026 U.S. data center capacity is canceled rests on a flawed denominator and a mismatch between early-stage announcements and projects actually under construction. The discussion then shifts to behind-the-meter power, gas pipelines, turbines, and why the team thinks AI data center buildouts will keep forcing new generation and infrastructure.

Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos
DeepSeek V4’s big leap is 1M context via aggressive sparse-attention and KV-cache compression, paired with a mega-MOE/mega-kernel approach to speed expert computation. The episode also compares day-zero support across Nvidia, Huawei Ascend, and AMD, highlighting how early access, kernel fusion, and tooling maturity shape real inference performance.

Ep. 016 - What Unitree's Evolution Means For Robotics (Robotics) | Jordan Nanos, Reyk Knuhtsen, Niko Ciminelli
The episode argues that Unitree’s real moat is not just robot quality, but China’s manufacturing density, fast iteration, and supply-chain depth. The guests think humanoid robotics is still early and messy, yet low-cost robots that are “good enough” for a few useful tasks could create real demand quickly.

Ep. 015 - DG Matrix Explains 800V DC vs Legacy AC Distribution (Datacenter, Energy) | Jordan Nanos, Jeremie Eliahou Ontiveros, Nicolas Bontigui, Haroon Inam
Haroon Inam of DG Matrix argues that 800V DC is becoming necessary as GPU rack power rises beyond what legacy AC distribution can deliver economically. The episode focuses on why voltage, not current, is the key constraint, and how multiport solid-state transformers could make datacenters more flexible, modular, and future-proof.

Ep. 014 - Finding Miscompiles For Fun, Not Profit (AI Infrastructure) | Justin Lebar & Jordan Nanos
Justin Lebar and Jordan Nanos discuss how Lebar found compiler miscompiles using both classic fuzzing and LLM-assisted code review. The episode emphasizes severe x86 and AMDGPU bugs, the difficulty of scaling fuzzers, and the surprising effectiveness—but real cost—of using agents to scan compiler code.