Google TPU v8: training vs inference hardware split, Virgo networking
- —First TPU generation with two distinct SKUs: v8t for training, v8i for inference — signals that training and inference have fundamentally different hardware requirements
- —TPU 8i carries 3x more on-chip SRAM than v8t, enabling fast decode without specialized inference chips like Groq LPU; Vik argues SRAM is the key, not exotic architectures
- —TPU 8i HBM capacity is higher than training chip — memory bandwidth is critical as more users and agents hit the same inference endpoint
- —Virgo scale-out architecture replaces Jupiter network using Optical Circuit Switches (OCS) with high switch radix, reducing networking layers and latency
- —Boardfly scale-up: 4 TPU 8i per board (copper) → 8 boards per rack (AEC) → 36 racks per pod (OCS) = 1,152 TPUs per pod; OCS is the structural backbone throughout
TPU v8 dual-SKU design shows architectural maturity; Virgo + Boardfly networking innovations position Google's own infrastructure ahead of relying on merchant silicon
Training vs inference hardware bifurcation is becoming structural — inference chips need more SRAM and HBM, training chips can trade memory for more compute tiles #催化劑
OCS (Optical Circuit Switches) becoming the structural networking substrate for hyperscale AI pods; Google's Virgo architecture makes this a long-term shift #催化劑