Paper Analysis | One Architecture, Five Generations, 3600x Growth: Google Dissects Its Own TPUs — and the Hidden Hero, Optical Circuit Switching
Those who once claimed "ASICs are too rigid to keep up with DNN evolution" probably don't want anyone digging up their old quotes now.
Four heavyweight Google authors — Norman Jouppi, Sridhar Lakshmanamurthy, Cliff Young, and David Patterson (of RISC, TPU and ML benchmarking fame) — lay out five TPU generations, from TPU v2 to Ironwood (TPU 7), in a paper for the July/August 2026 issue of IEEE Micro. The conclusion fits in one sentence: the architecture barely changed, yet the system grew 3600x.
For STT readers, the real value isn't the PR line "Google is fast and efficient" — it's three things:
1. How an architecture defined in 2017 survived the full workload turnover to Transformers and Diffusion without being retired; 2. OCS (optical circuit switching) is what keeps a machine of nearly 10,000 chips alive, schedulable, and deployable in stages — and it is still unique to TPU; 3. Power has officially replaced cost as the new ceiling. perf/TCO gives way to perf/Watt, and Google introduces a new metric along the way: CCI.
1. Background: Not a Spec Sheet, but a Retrospective on a Bet That Paid Off
TPU v1 (2016) was only an inference chip, yet it ignited an entire industry. The paper even quotes the quip that TPU v1 "launched a thousand chips" (a nod to Helen of Troy's face that "launched a thousand ships"). Within four months Intel spent billions of dollars acquiring Nervana and Movidius, and within two years VCs poured about $3 billion into more than a hundred DNN chip startups.
But the real protagonist is TPU v2 — Google's first training supercomputer, deployed in 2017. The paper argues something counterintuitive: the biggest risk of a domain-specific architecture (DSA) is that design-to-production takes 2–3 years, and by the time it ships DNNs have already changed — yet the TPU v2 microarchitecture has carried all the way to Ironwood without being replaced.
2. The Core Question: The Workload Changed Three Times — Why Didn't the Architecture?

This chart is the foundation of the whole argument: the workload underneath turned over three times, while the hardware architecture above it stayed put. It means the bet TPU v2 made was placed on the "greatest common divisor" of DNNs, not on whatever model was hottest at the time.
3. The Core of a Stable Architecture: Two Big Cores, Systolic Arrays, BF16

The paper identifies several key decisions TPU v2 got right:
Just two big cores (TensorCores): a balance between the long-wire latency of one giant core and the software burden of stitching together many small ones. The paper closes with Seymour Cray's famous line: "If you were plowing a field, which would you rather use: two strong oxen or 1024 chickens?"
A systolic array for the MXU: v2 used a 128×128 multiply-accumulate array delivering 32,768 operations per cycle — a design later copied by almost every accelerator.
BF16, the first DNN accelerator to depart from IEEE floating point: the judgment was that for DNNs "dynamic range matters more than precision," so for the first time the exponent (8 bits) was larger than the mantissa (7 bits). FP8 and FP4 all followed suit.
A compiler-managed memory hierarchy: DMA plus scratchpad SRAM instead of CPU-style caches, making behavior predictable and reproducible.
These four, plus HBM as main memory, custom ICI links to build a supercomputer, and vector units for non-matrix operations, make up the six key decisions listed at the end of the paper.
4. Scale: In 8 Years, Nearly Every Dimension Grew by One to Three Orders of Magnitude

Using v2 as the baseline, here is what the table really says:
Per-chip HBM capacity/bandwidth ~10x: 16 GiB / 700 GB/s → 192 GiB / 7300 GB/s
Per-chip peak compute ~100x: 46 BF16 TFLOPS → 2307 BF16 (or 4614 FP8) TFLOPS
Supercomputer chip count 36x: 256 → 9216 chips
System-wide shared HBM ~400x: 4 TB → 1.77 PB (a new record for AI supercomputers)
perf/Watt ~30x; system-wide peak compute ~3600x
That 3600x works out to a compound annual growth rate of nearly 100% — achieved in an era when Dennard scaling is long dead and Moore's Law has stalled. The paper puts it bluntly: TPU has crossed the so-called "Accelerator Wall."
5. The Truth About Evolution: Always Four TPUs per Board — Only the Scale Changes

What you can't see is the jump from air cooling (v2) to liquid cooling (from v3 onward). The paper stresses that since TPU v3 in 2018, Google's training has been fully liquid-cooled and has never gone back.
Microarchitectural evolution is almost entirely "the same thing, bigger and more of it": the MXU grew from two 128×128 arrays to four 256×256 (BF16) plus four 512×512 (FP8); Ironwood even adds redundant rows inside the MXU to recover yield.
6. STT Focus: OCS Optical Switching — the Machine's Circulatory System
The paper states that so far only two innovations remain unique to TPU: OCS and SparseCore. For an industry that talks about CPO and Optical I/O every day, this is an easily overlooked reminder: optics isn't only winning at the chip edge — it has been deployed at the system interconnect layer for years.

How it works:
TPU v4 was the first supercomputer to use OCS: based on 3D MEMS micromirrors with millisecond-scale switching, it essentially adds a physical-layer optical switching layer beneath ICI that can route around failures.
Why a 4×4×4 = 64-chip cube? A cube gives the best bisection bandwidth in a 3D Torus, and 64 chips plus 16 CPU hosts fit exactly into one rack.
What OCS delivers isn't "faster" — it's operability of the entire machine:
Availability: failures are routed around at the physical layer and the 3D Torus is rebuilt. Ironwood has 2304 CPU hosts — without OCS, host availability would need to exceed 99.9% to sustain system-wide goodput.
Scheduling: with OCS, any two cubes can be picked from anywhere in the machine. That's why Ironwood has 9216 (9K) chips rather than a power of two — it can run four 2K-slice jobs simultaneously and still keep 16 spare cubes.
Staged deployment: v3 couldn't be used until all 1024 chips were installed and tested; from v4 on, each rack is independent and goes into production as soon as its 64 chips are in.
Single-tier network: the whole pod stays on one flat network, and failed parts are simply swapped out.
Without OCS, the alternative would require spare TPUs and standalone crossbar switches in every rack, potentially consuming nearly half the rack space.
STT view: this section reframes optical switching from a nice-to-have into the circulatory system of the training fabric. While the market's eyes are fixed on chip-edge CPO, Google is reminding everyone that system-level optical switching has been quietly paying dividends for five years.
7. Hardening Resilience: Ironwood Builds "Catching Bad Chips" into Silicon
What training supercomputers fear most isn't crashes — it's silent data corruption (SDC). It used to be handled by software health checks; Ironwood pushes it into hardware:
FBIST (Functional Built-In Self-Test): integrated into the MXU, it runs high-coverage tests during manufacturing, burn-in and operation, catching chips that pass structural tests but are actually defective.
VPU hardware replay unit: using idle VLIW slots, it re-runs odd-lane operations on even lanes and compares results — with zero performance loss and negligible power. When a bad unit is caught, OCS takes it offline immediately for repair.
For a sense of scale: Google trained Gemini 2.5 with synchronous data parallelism across multiple 8960-chip v5p pods spanning several data centers, achieving 93% goodput (Gemini 1.0 on v4 reached 97%).
8. The New Ceiling: Power, Not Cost

The paper cites Vahdat et al.: new data centers are finding it increasingly hard to secure enough power, so performance per watt now matters more than performance per TCO. Of Ironwood's 29.3x, the single v5p→Ironwood generation contributes about 6x.
9. Counting Carbon Too: The New CCI Metric

perf/Watt only captures operational emissions, not embodied carbon. Google therefore proposes CCI (Compute Carbon Intensity): the CO2e emitted per useful floating-point operation, where total CCI = operational CCI + embodied CCI. The counterintuitive part: newer TPUs draw more power and carry more embodied carbon per chip, but because they are so much faster and run for fewer seconds, the carbon per floating-point operation actually goes down. Across all three generations, operational carbon is about 75% of the total (the opposite of smartphones, where 87% is embodied).
10. Verdict: Where This Paper Sits in Technology History
The paper closes by ranking its six winning bets in order of importance:
Systolic arrays for matrix multiplication
Wide-range, narrow-width floating point (BF16/FP8/FP4) instead of high-precision, wide IEEE floating point
HBM as main memory
Custom high-speed links (ICI) that assemble accelerators into a supercomputer
A software-managed memory hierarchy using DMA plus scratchpad SRAM instead of CPU-style caches
Vector units for non-matrix operations
The authors draw a bold analogy: in the 1960s CPUs converged on microcode, backward compatibility and caches (pioneered by the IBM 360), and only in the 1990s shifted to superscalar out-of-order execution (Pentium Pro). Their bet is that the six points above will be the defining features of training accelerators in the 2020s.
STT's take on this paper: this isn't a spec showcase — it's a textbook on architectural discipline. The core lesson fits in one sentence: bet on the greatest common divisor, not on whichever model is hottest right now. And for the optical communications industry, it adds an official stamp: optical switching in AI training fabrics is deployed, hard to replace, and still unique infrastructure.




Comments