top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

Paper Analysis | One Architecture, Five Generations, 3600x Growth: Google Dissects Its Own TPUs — and the Hidden Hero, Optical Circuit Switching

2 days ago
7 min read

Those who once claimed "ASICs are too rigid to keep up with DNN evolution" probably don't want anyone digging up their old quotes now.

Four heavyweight Google authors — Norman Jouppi, Sridhar Lakshmanamurthy, Cliff Young, and David Patterson (of RISC, TPU and ML benchmarking fame) — lay out five TPU generations, from TPU v2 to Ironwood (TPU 7), in a paper for the July/August 2026 issue of IEEE Micro. The conclusion fits in one sentence: the architecture barely changed, yet the system grew 3600x.

For STT readers, the real value isn't the PR line "Google is fast and efficient" — it's three things:

1. How an architecture defined in 2017 survived the full workload turnover to Transformers and Diffusion without being retired; 2. OCS (optical circuit switching) is what keeps a machine of nearly 10,000 chips alive, schedulable, and deployable in stages — and it is still unique to TPU; 3. Power has officially replaced cost as the new ceiling. perf/TCO gives way to perf/Watt, and Google introduces a new metric along the way: CCI.

1. Background: Not a Spec Sheet, but a Retrospective on a Bet That Paid Off

TPU v1 (2016) was only an inference chip, yet it ignited an entire industry. The paper even quotes the quip that TPU v1 "launched a thousand chips" (a nod to Helen of Troy's face that "launched a thousand ships"). Within four months Intel spent billions of dollars acquiring Nervana and Movidius, and within two years VCs poured about $3 billion into more than a hundred DNN chip startups.

But the real protagonist is TPU v2 — Google's first training supercomputer, deployed in 2017. The paper argues something counterintuitive: the biggest risk of a domain-specific architecture (DSA) is that design-to-production takes 2–3 years, and by the time it ships DNNs have already changed — yet the TPU v2 microarchitecture has carried all the way to Ironwood without being replaced.


2. The Core Question: The Workload Changed Three Times — Why Didn't the Architecture?

In 2016, MLPs were 61% of the workload and RNNs 29%; by 2026 Transformers dominate at 74%, RNNs are almost gone (<0.5%), and Diffusion has overtaken CNNs. The Transformer paper appeared in December 2017, and within 15 months it accounted for 21% of Google's production workload. Source: Jouppi et al., IEEE Micro 2026, Figure 1
In 2016, MLPs were 61% of the workload and RNNs 29%; by 2026 Transformers dominate at 74%, RNNs are almost gone (<0.5%), and Diffusion has overtaken CNNs. The Transformer paper appeared in December 2017, and within 15 months it accounted for 21% of Google's production workload. Source: Jouppi et al., IEEE Micro 2026, Figure 1

This chart is the foundation of the whole argument: the workload underneath turned over three times, while the hardware architecture above it stayed put. It means the bet TPU v2 made was placed on the "greatest common divisor" of DNNs, not on whatever model was hottest at the time.


3. The Core of a Stable Architecture: Two Big Cores, Systolic Arrays, BF16


Blue is the compute datapath, green is HBM, purple is the host connection, and yellow is the interconnect router. This v2 diagram still holds for every training TPU through Ironwood. Source: Jouppi et al., IEEE Micro 2026, Figure 2 (originally Norrie et al., 2021)
Blue is the compute datapath, green is HBM, purple is the host connection, and yellow is the interconnect router. This v2 diagram still holds for every training TPU through Ironwood. Source: Jouppi et al., IEEE Micro 2026, Figure 2 (originally Norrie et al., 2021)

The paper identifies several key decisions TPU v2 got right:

  • Just two big cores (TensorCores): a balance between the long-wire latency of one giant core and the software burden of stitching together many small ones. The paper closes with Seymour Cray's famous line: "If you were plowing a field, which would you rather use: two strong oxen or 1024 chickens?"

  • A systolic array for the MXU: v2 used a 128×128 multiply-accumulate array delivering 32,768 operations per cycle — a design later copied by almost every accelerator.

  • BF16, the first DNN accelerator to depart from IEEE floating point: the judgment was that for DNNs "dynamic range matters more than precision," so for the first time the exponent (8 bits) was larger than the mantissa (7 bits). FP8 and FP4 all followed suit.

  • A compiler-managed memory hierarchy: DMA plus scratchpad SRAM instead of CPU-style caches, making behavior predictable and reproducible.

These four, plus HBM as main memory, custom ICI links to build a supercomputer, and vector units for non-matrix operations, make up the six key decisions listed at the end of the paper.

4. Scale: In 8 Years, Nearly Every Dimension Grew by One to Three Orders of Magnitude

From v2 to Ironwood, almost every dimension jumped by one to three orders of magnitude. Source: Jouppi et al., IEEE Micro 2026, Table 1
From v2 to Ironwood, almost every dimension jumped by one to three orders of magnitude. Source: Jouppi et al., IEEE Micro 2026, Table 1

Using v2 as the baseline, here is what the table really says:

  • Per-chip HBM capacity/bandwidth ~10x: 16 GiB / 700 GB/s → 192 GiB / 7300 GB/s

  • Per-chip peak compute ~100x: 46 BF16 TFLOPS → 2307 BF16 (or 4614 FP8) TFLOPS

  • Supercomputer chip count 36x: 256 → 9216 chips

  • System-wide shared HBM ~400x: 4 TB → 1.77 PB (a new record for AI supercomputers)

  • perf/Watt ~30x; system-wide peak compute ~3600x

That 3600x works out to a compound annual growth rate of nearly 100% — achieved in an era when Dennard scaling is long dead and Moore's Law has stalled. The paper puts it bluntly: TPU has crossed the so-called "Accelerator Wall."

5. The Truth About Evolution: Always Four TPUs per Board — Only the Scale Changes

HBM stacks sit on either side of the compute die: 4 on v4, 6 on v5p, and 8 on Ironwood, which also has two compute dies. Four TPUs per board has never changed. Source: Jouppi et al., IEEE Micro 2026, Figure 3
HBM stacks sit on either side of the compute die: 4 on v4, 6 on v5p, and 8 on Ironwood, which also has two compute dies. Four TPUs per board has never changed. Source: Jouppi et al., IEEE Micro 2026, Figure 3

What you can't see is the jump from air cooling (v2) to liquid cooling (from v3 onward). The paper stresses that since TPU v3 in 2018, Google's training has been fully liquid-cooled and has never gone back.

Microarchitectural evolution is almost entirely "the same thing, bigger and more of it": the MXU grew from two 128×128 arrays to four 256×256 (BF16) plus four 512×512 (FP8); Ironwood even adds redundant rows inside the MXU to recover yield.


6. STT Focus: OCS Optical Switching — the Machine's Circulatory System

The paper states that so far only two innovations remain unique to TPU: OCS and SparseCore. For an industry that talks about CPO and Optical I/O every day, this is an easily overlooked reminder: optics isn't only winning at the chip edge — it has been deployed at the system interconnect layer for years.


A 4×4×4 cube (64 chips) and its OCS optical interconnect. Each cube has 6 faces with 16 optical links per face, 96 links in total, all going to OCSes; the 3D Torus wraparound requires opposite faces to land on the same OCS, so each cube connects to 48 OCSes. Source: Jouppi et al., IEEE Micro 2026, Figure 4 (originally Jouppi et al., 2023)
A 4×4×4 cube (64 chips) and its OCS optical interconnect. Each cube has 6 faces with 16 optical links per face, 96 links in total, all going to OCSes; the 3D Torus wraparound requires opposite faces to land on the same OCS, so each cube connects to 48 OCSes. Source: Jouppi et al., IEEE Micro 2026, Figure 4 (originally Jouppi et al., 2023)

How it works:

  • TPU v4 was the first supercomputer to use OCS: based on 3D MEMS micromirrors with millisecond-scale switching, it essentially adds a physical-layer optical switching layer beneath ICI that can route around failures.

  • Why a 4×4×4 = 64-chip cube? A cube gives the best bisection bandwidth in a 3D Torus, and 64 chips plus 16 CPU hosts fit exactly into one rack.

What OCS delivers isn't "faster" — it's operability of the entire machine:

  1. Availability: failures are routed around at the physical layer and the 3D Torus is rebuilt. Ironwood has 2304 CPU hosts — without OCS, host availability would need to exceed 99.9% to sustain system-wide goodput.

  2. Scheduling: with OCS, any two cubes can be picked from anywhere in the machine. That's why Ironwood has 9216 (9K) chips rather than a power of two — it can run four 2K-slice jobs simultaneously and still keep 16 spare cubes.

  3. Staged deployment: v3 couldn't be used until all 1024 chips were installed and tested; from v4 on, each rack is independent and goes into production as soon as its 64 chips are in.

  4. Single-tier network: the whole pod stays on one flat network, and failed parts are simply swapped out.


Without OCS, the alternative would require spare TPUs and standalone crossbar switches in every rack, potentially consuming nearly half the rack space.

STT view: this section reframes optical switching from a nice-to-have into the circulatory system of the training fabric. While the market's eyes are fixed on chip-edge CPO, Google is reminding everyone that system-level optical switching has been quietly paying dividends for five years.


7. Hardening Resilience: Ironwood Builds "Catching Bad Chips" into Silicon

What training supercomputers fear most isn't crashes — it's silent data corruption (SDC). It used to be handled by software health checks; Ironwood pushes it into hardware:

  • FBIST (Functional Built-In Self-Test): integrated into the MXU, it runs high-coverage tests during manufacturing, burn-in and operation, catching chips that pass structural tests but are actually defective.

  • VPU hardware replay unit: using idle VLIW slots, it re-runs odd-lane operations on even lanes and compares results — with zero performance loss and negligible power. When a bad unit is caught, OCS takes it offline immediately for repair.

For a sense of scale: Google trained Gemini 2.5 with synchronous data parallelism across multiple 8960-chip v5p pods spanning several data centers, achieving 93% goodput (Gemini 1.0 on v4 reached 97%).


8. The New Ceiling: Power, Not Cost

Normalized to v2 = 1: v3 = 1.8, v4 = 4.9, v5p = 5.2, and Ironwood jumps to 29.3. Higher is better. Source: Jouppi et al., IEEE Micro 2026, Figure 5
Normalized to v2 = 1: v3 = 1.8, v4 = 4.9, v5p = 5.2, and Ironwood jumps to 29.3. Higher is better. Source: Jouppi et al., IEEE Micro 2026, Figure 5

The paper cites Vahdat et al.: new data centers are finding it increasingly hard to secure enough power, so performance per watt now matters more than performance per TCO. Of Ironwood's 29.3x, the single v5p→Ironwood generation contributes about 6x.


9. Counting Carbon Too: The New CCI Metric


v4 totals 339 (operational 249 / embodied 90), v5p 292, and Ironwood only 79. Ironwood's operational CCI is about 3.7x better than v5p's, and its embodied carbon about 3.8x better. Embodied carbon is amortized over a 6-year lifetime. Lower is better. Source: Jouppi et al., IEEE Micro 2026, Figure 6
v4 totals 339 (operational 249 / embodied 90), v5p 292, and Ironwood only 79. Ironwood's operational CCI is about 3.7x better than v5p's, and its embodied carbon about 3.8x better. Embodied carbon is amortized over a 6-year lifetime. Lower is better. Source: Jouppi et al., IEEE Micro 2026, Figure 6

perf/Watt only captures operational emissions, not embodied carbon. Google therefore proposes CCI (Compute Carbon Intensity): the CO2e emitted per useful floating-point operation, where total CCI = operational CCI + embodied CCI. The counterintuitive part: newer TPUs draw more power and carry more embodied carbon per chip, but because they are so much faster and run for fewer seconds, the carbon per floating-point operation actually goes down. Across all three generations, operational carbon is about 75% of the total (the opposite of smartphones, where 87% is embodied).


10. Verdict: Where This Paper Sits in Technology History

The paper closes by ranking its six winning bets in order of importance:

  1. Systolic arrays for matrix multiplication

  2. Wide-range, narrow-width floating point (BF16/FP8/FP4) instead of high-precision, wide IEEE floating point

  3. HBM as main memory

  4. Custom high-speed links (ICI) that assemble accelerators into a supercomputer

  5. A software-managed memory hierarchy using DMA plus scratchpad SRAM instead of CPU-style caches

  6. Vector units for non-matrix operations

The authors draw a bold analogy: in the 1960s CPUs converged on microcode, backward compatibility and caches (pioneered by the IBM 360), and only in the 1990s shifted to superscalar out-of-order execution (Pentium Pro). Their bet is that the six points above will be the defining features of training accelerators in the 2020s.


STT's take on this paper: this isn't a spec showcase — it's a textbook on architectural discipline. The core lesson fits in one sentence: bet on the greatest common divisor, not on whichever model is hottest right now. And for the optical communications industry, it adds an official stamp: optical switching in AI training fabrics is deployed, hard to replace, and still unique infrastructure.


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page