
Summary
This episode reviews ClusterMAX 3.0, SemiAnalysis’s framework for ranking neocloud GPU providers based on hands-on testing, reliability, networking, and operations. The new rankings move Nebius into platinum alongside CoreWeave, put Oracle and Google Cloud in gold, and push Azure down while several others land in silver and bronze. The hosts argue that current GPU pricing is deeply backwardated, so near-term availability dominates software quality and lets providers earn strong margins even with weak cluster tooling. A major technical focus is health checks and failure remediation, including active vs passive checks, DCGM-based failure injection, node quarantine workflows, and architecture-specific recovery for HGX and NVL72 systems. The episode also previews endpoint testing, explains why NCCL and networking are decisive differentiators, and expands into neocloud financing, insurance, and the infrastructure consequences of rising power density and older GPUs staying economically useful.