SemiAnalysis

The GPU Power-Performance Curve Most Clusters Ignore | Researcher Conversations at GTC

Original source
Artwork for The GPU Power-Performance Curve Most Clusters Ignore | Researcher Conversations at GTC

Summary

This episode focuses on Pebble’s approach to data center energy optimization for AI workloads, with an emphasis on maximizing tokens per watt rather than simply increasing power. The core technical claim is that GPU power/performance is non-linear: beyond a point, extra power produces diminishing or saturated token output. Pebble says inference is often memory-bound, with energy spent moving weights into HBM while SMs are underutilized, especially when prefill waits on decode. Its system ingests telemetry from inference servers and NVIDIA GPUs, deploys via Helm on Kubernetes, and dynamically tunes per-GPU power caps and clock frequencies after learning workload behavior for a few days. The company is also exploring grid-responsive data centers and flexible power arrangements, claiming there is 100 gigawatts of flexible power in the U.S. that AI clusters could access through curtailment tradeoffs.

Notes