The Next Decade of AI Supercomputing Interconnects: The Architectural Shift from GPU-Centric to Fabric-Centric
Introduction: Compute Is No Longer a GPU Race — It's an Interconnect Race
The AI boom has pushed the world's data centers into one unmistakable trend:
Bigger models → exploding token consumption → training and inference costs become a fundamental infrastructure problem.
From 2023 to 2025, the leading global CSPs (Google / AWS / Meta / Microsoft / Alibaba / Tencent / ByteDance) have raised capex one after another:
Google: $85 billion
Meta: $66–72 billion
AWS: over $100 billion
Microsoft: over $30 billion in a single quarter
The four major U.S. CSPs alone will spend more than $361 billion in combined capex in 2025.
The most important point here — the structure of the spending is changing:
From "buying more GPUs" → "building a dedicated AI fabric."
This article therefore focuses on:
How is AI cloud compute architecture moving from the GPU era to ASIC + fabric?
What new phases are scale-up / scale-out / scale-across interconnect technologies entering?
Why are CPO, OCS, backplanes and optical I/O the biggest AI supply chain opportunities of the next 10 years?
The Global CSP Compute Map: The Main Battlefield of AI Fabrics
The reasons global clouds are moving to in-house ASICs are remarkably consistent:
Lower TCO, better energy efficiency, control of the supply chain, and avoiding dependence on NVIDIA alone.
Below, we break down the compute and interconnect architectures of the major players.
Google: TPU + OCS, the World's Most Forward-Looking Optical Switching Architecture
Seven Generations of TPU (v1 → v7 Ironwood)
Starting with TPU v4, Google clearly shifted its focus to:
3D torus interconnect
All-optical switching (OCS)
1.6T optical modules
Rack (64 chips) → superpod (4,096 chips)
TPU v7 (Ironwood) goes further:
192GB HBM (a big jump over v5e/v5p)
ICI (inter-chip interconnect) bandwidth of 9,600 Gbps
Peak compute ≥ 4,614 TFLOPS (BF16)
This isn't a linear improvement — it treats the interconnect as a core performance metric.
3D Torus: A Topology Tailor-Made for AI Training
Typical Ethernet uses Clos/fat-tree; TPU chose a torus because it:
Offers high scalability
Is more stable for large batches and ultra-large models
Delivers predictable transmission latency under a fixed topology
64 TPUs → a 4×4×4 3D network is Google's optimized sweet spot.
OCS: Google's "Next-Generation Technology That Skips CPO"
In its 4,096-chip superpod, Google deploys:
48 OCS units (Optical Circuit Switches)
Each supporting 320×320+ optical cross-connects
Zero buffering, zero serialization, zero retiming
This means:
An architecture with truly all-optical paths is the end-state AI network.
Three major advantages of OCS:
Latency an order of magnitude lower than Ethernet
Far lower power than Ethernet switching
Well suited to AI training workloads with "heavy traffic but rarely changing topology"
Optical Module Demand (TPU v4/v5/v6/v7)
TPUs require large numbers of optical modules:
Architecture | TPU : optical modules |
In-rack DAC | 1 : 4 |
4,096-chip superpod (OCS) | 1 : 1.5 |
Large-scale fat-tree network | 1 : 4.5 |
Google's optical module procurement strategy is crystal clear:
Massive purchasing + full control of the architecture + aggressive in-house development.
AWS: Trainium → Teton, from Copper to Backplane
AWS's strength isn't raw compute but engineering and cost optimization.
It doesn't chase the "highest TOPS" but the best performance/cost/power ratio.
Trainium 2: In-Rack Interconnect Dominated by AEC + DAC
NeuronLink v3: 32DP × 32Gbps per chip
In-rack DAC: 1:9
Rack-to-rack AEC: 1:1
AWS's focus is:
Delivering 80% of large-model training performance without using ultra-expensive GPUs.
Trainium 3 (Teton): Going All-In on Backplanes
Teton PDS / Teton Max introduce:
A backplane replacing copper-packed cable trays
More PCIe switches (32–40)
Liquid cooling + high-density topology (64–72 chips)
AWS knows full well:
Copper density is the ceiling; the backplane is the long-term solution.
This is exactly the same direction as NVIDIA Rubin.
Meta: MTIA and Minerva, a Heavily Customized Network Architecture
Meta has three world-class capabilities:
It understands AI models and workloads
It understands data center design (Meta has driven it since the Clos architecture)
It can co-design ASICs, NICs, switches and racks as one system
MTIA-T: 800G Interconnect × Heavy DAC Use
Each MTIA-T chip connects via 4×800G → scale-up
Connected to TH5/TH6 → 8×800G
Very high in-rack DAC ratio: 1:12
Meta's strategy is:
Beat high-priced GPUs with "moderate bandwidth + heavy design optimization."
Scale-out (two-tier fat-tree) optical module demand:
MTIA : 800G optical modules = 1 : 8
Minerva: Meta's In-House ASIC + Broadcom J3 Switch
16 MTIA-T compute blades
6 network blades (scale-up + scale-out)
The difference from Google/NVIDIA:
Meta relies more heavily on Ethernet switching
Very high demand for optical modules and DACs
NVIDIA: NVLink + Optical Modules + Backplane on Three Fronts
NVIDIA doesn't just sell GPUs; it sells:
A complete AI fabric (GPU + NVSwitch + NIC + Spectrum).
That makes NVIDIA the standard-setter for the world's AI networks.
GB200: The Peak of the 800G Era
576-GPU rack
GPU : 800G optical modules ≈ 1 : 1.5–2.5
NVLink 5.0 at 1.8TB/s bidirectional
C2C runs entirely on copper (large numbers of differential pairs)
GB200's biggest problem isn't performance but:
Extremely high cooling density (6 pairs of UQDs)
Extremely high cable density (thousands of copper cables)
Hence NVIDIA launched GB300.
GB300: Entering the 1.6T / CPO Era
NVLink 6.0 + CX8 NIC (800G):
Liquid-cooling lines increase to 14 pairs of UQDs
NICs move to the 1.6T generation (CX9)
CPO switches (Spectrum) begin to be adopted
Rubin / Feynman: NVIDIA's Most Critical Architectural Turning Point
Rubin = orthogonal backplane + CPO + full liquid cooling
The Rubin architecture brings:
Elimination of cable trays
A fully backplane-based rack (same direction as AWS Teton)
Lower latency, better stability, simpler cabling
Scale-up that looks more like a mainframe
Feynman then takes:
NVLink to 7,200GB/s
Spectrum to 204T (the CPO era)
NICs to CX10 (the 3.2T era)
NVIDIA is taking the AI fabric toward:
PCB → CPO → all-optical architecture → light-speed fabric
Three Interconnect Tiers: Scale-Up / Scale-Out / Scale-Across
1) Scale-Up (In-Rack Interconnect): Determines Training Throughput
Mainstream technologies:
DAC / AEC (short reach)
Orthogonal backplane (medium reach)
NVLink / PCIe switch
OIO (future)
The next generation will be:
Optical backplanes replacing copper backplanes.
2) Scale-Out (Rack-to-Rack Interconnect): Determines Cluster Size
Mainstream technologies:
CPO (800G → 1.6T → 3.2T)
OCS (led by Google)
High-density fiber (MPO → MMC)
Ethernet (Broadcom Tomahawk)
InfiniBand (NVIDIA Quantum)
The main battlefield of scale-out is optics.
3) Scale-Across (Data-Center-to-Data-Center Interconnect): The DCI Era
Technology directions:
Coherent optics
Hollow-core fiber (low latency)
400G ZR / ZR+
Spectrum-X (AI-specific L2/L3 fabric)
Big Trends in AI Interconnect Over the Next Decade
Trend 1: AI's Bottleneck Is Shifting from Compute to Bandwidth
GPU TFLOPS are no longer the point; what really drives training speed is:
NVLink
PCIe switch
NIC
Switch ASIC
Optical modules
Topology
The AI fabric will be the essence of cloud competitiveness.
Trend 2: Copper Will Gradually Give Way to Optics — Possibly Over 5–8 Years
The migration path:
Copper → backplane → optical backplane → CPO → OCS → OIO
Silicon photonics will play three roles along the way:
CPO optical engines (the core of 1.6T/3.2T)
OIO (the end state)
Optical backplanes (fiberizing the backplane)
Trend 3: AI Data Centers Will Become a Three-Layer Fabric
No longer "server + switch," but:
Compute Fabric (GPU/ASIC/NVLink/PCIe switch)
Optical Fabric (CPO, OCS, optical backplane)
Cooling Fabric (full liquid cooling, UQD, manifold)
This is an architecture completely different from the traditional data center.
AI Data Centers Are Entering a Golden Decade of Optoelectronic Convergence
The world is entering a massive turning point:
Compute in the AI era isn't about stacking GPUs — it's about stacking interconnect.
NVIDIA, Google, AWS, Meta and Huawei are all doing the same thing:
Upgrading electrical interconnect to optical interconnect
Turning the copper density problem into an optical solution
Turning the network from Ethernet into a fabric
Moving thermal management from air cooling to full liquid cooling
Over the next 10 years, what will truly dominate the AI data center value chain isn't the GPU, but:




Comments