top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

Paper Analysis | Winzer Nails the I/O Bottleneck of AI Clusters: Why Copper Can't Keep Up, WDM Can't Win, and CPO Is the Only Way Out

2 days ago
11 min read

In this SPIE invited paper, Nubis' Peter Winzer pins down the whole AI-cluster interconnect dilemma with a single "scissors gap" chart: AI parameter counts are exploding at roughly 2000% a year, while the I/O bandwidth that feeds them grows only 20% a year. He then walks through the Roofline model, the three-tier I/O ladder of NVL72, fan-out topologies and energy curves to arrive at a pointed conclusion: the distance wall of copper in the scale-up network is the real ceiling of this generation of AI clusters; the answer is not WDM, nor retimed optical modules, but linear optics that "pretend to be copper" (LPO / CPO). This is not a new-technology announcement but an architecture-level reading of "why now, and why this path" - worth going through figure by figure.

1. Background

The author, Peter J. Winzer, is a heavyweight in optical communications - years of high-speed transmission work at Bell Labs, a long list of OFC / ECOC best papers, and now a core figure at the startup Nubis Communications. "High-density High-Speed Linear I/O For AI Clusters" is an invited paper at SPIE Photonics West 2025 (Next-Generation Optical Communication XIV, Vol. 13374), co-authored with D. J. Harding of Nubis.

We picked this paper not for any stunning new device data, but because Winzer rebuilds, from first principles, the case for why AI clusters must move to optical interconnect. Plenty of articles talk about CPO, but most stop at slogans like "copper is hitting its limit"; Winzer, unusually, fills in the physical and system-level argument behind every "why". For anyone trying to understand the industry logic of Co-Packaged Optics (CPO), it is a rare skeleton-level textbook.


2. What Problem Is This Paper Solving?

In one sentence: in an AI cluster built from hundreds of thousands of accelerators (xPUs), how much I/O should each get, over how many meters, on what medium and at what power, so that expensive compute isn't left idle and "underfed"?

Winzer breaks it into five questions: How much I/O bandwidth does an xPU need? Fat pipes or fan-out? What reach should links support? How much I/O density is required? How much power and cost can each I/O socket be allotted? These five questions map neatly onto the six figures that follow.


3. Figure 1: AI Parameters Grow ~2000%/Year, I/O Only 20%

This figure shows the driving force of the entire paper - long-term growth curves for a set of information-technology metrics (log y-axis, so slope equals growth rate).

Line them up and the gap is absurd: per-lane interconnect speed - Ethernet, PCIe, HBM and the like - has long grown at only about 20% per year; Ethernet switching capacity at about 40% per year; microprocessor floating-point performance (Flops/s) at about 75% per year. And AI model parameter counts grow at roughly 2000% per year (400x every two years).

Compute grows 75% a year; the pipes feeding it grow 20%. That scissors gap is the source of all the pain in this generation of AI infrastructure.

Winzer honestly adds a caveat: this explosive growth will saturate sooner or later (Netflix streaming and wireless networks went through the same), but even after saturation, the gap accumulated in the early years is enough to keep the technology supply chain catching up for many years. In other words, this is not a short-term theme but a structural, long-term gap. I/O is the "supporting technology" being called on to catch up.


Figure 1: Long-term growth curves of information and communication technologies - AI parameter count (~2000%/year) far outpaces per-lane I/O (~20%/year), switching capacity, Flops/s and every other hardware metric. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 1
Figure 1: Long-term growth curves of information and communication technologies - AI parameter count (~2000%/year) far outpaces per-lane I/O (~20%/year), switching capacity, Flops/s and every other hardware metric. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 1

4. Figure 2: A GPU's I/O Is a Three-Step Ladder, Dropping 10x at Each Step Out

The left side of this figure is the Roofline model, the right side is NVIDIA's NVL72 cluster architecture; they need to be read together.

Left (Roofline) first answers "is there enough I/O?" System performance depends on three things: processor compute CP, I/O bandwidth B, and the algorithm's arithmetic intensity I (operations per byte of I/O, Flops/Byte). When a system is I/O-bound, actual compute is only I x B, far below the processor's nominal CP - a pricey GPU sits idle because data can't get in. The orange line in the figure is an Nvidia B200, clearly sitting on the "I/O-bound" slope, as is Google TPUv1. This confirms with data that many AI platforms really are shackled by I/O.

Right (NVL72) reveals a key structure: a GPU's I/O is a tiered ladder. Taking the B200 as an example -

  • To its own 192 GB of in-package memory: 8 TB/s, i.e. 64 Tbps

  • To the scale-up domain (72 GPUs and 18 NVSwitches in the same rack): 7.2 Tbps per GPU (over passive copper)

  • To the scale-out domain (across racks, a 5,184-GPU cluster): only 800 Gbps per GPU (over single-mode fiber)

Each step outward cuts bandwidth by about 10x. This 64T -> 7.2T -> 0.8T ladder is the foundation for understanding every trade-off that follows. Incidentally, total scale-up domain bandwidth reaches 518 Tbps, and the two-tier scale-out network exceeds 8 Pbps in total - at that scale it's no longer a "network", it's a breathing data center.

We break down this scale-up / scale-out layering in detail in The Great Optical Packaging Transition (Part 2): CPO's Three-Stage Evolution - worth reading alongside.

Figure 2: Left, the Roofline model - the orange B200 line sits clearly on the I/O-bound slope; right, NVL72's three I/O domains - on-package 64 Tbps, scale-up 7.2 Tbps, scale-out 800 Gbps, roughly 10x apart at each step. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 2
Figure 2: Left, the Roofline model - the orange B200 line sits clearly on the I/O-bound slope; right, NVL72's three I/O domains - on-package 64 Tbps, scale-up 7.2 Tbps, scale-out 800 Gbps, roughly 10x apart at each step. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 2


5. The Distance Wall: Copper Locks Scale-Up into a Single Rack, at the Cost of 120 kW

The photo of passive copper cable (the inset) in Figure 2's right panel looks unremarkable, yet it kicks off the paper's most cutting argument.

Why does scale-up use passive copper and only scale-out use optics? Because scale-up needs 10x the bandwidth of scale-out, and copper is cheap and power-efficient - use it as long as it holds up. But copper's price is distance: at 200 Gbps/lane, passive copper reaches only about 1 meter. That 1 meter directly limits the scale-up domain to "however many GPUs fit in one rack".

So to squeeze 72 GPUs into one shared-memory domain, NVL72 crams 120 kW of hardware into a single rack - a power density extreme enough to cause cooling and power-delivery problems. And it only gets worse: once SerDes jumps from 200 Gbps/lane to 400 Gbps/lane, copper reach is cut at least in half again, and rack power density for the same 72 GPUs at least doubles.

Marvell spelled out this physical timetable even more plainly at Computex; I summarized it in Computex 2026 Keynote: The Copper Wall Is Moving Into the Rack. There are two ways through this wall:

  • Active Copper Cable (ACC): low power and low cost, it can grow the number of GPUs in a scale-up domain by up to 6x. For how the copper camp turns "reliability" into a business, see Earnings Call Highlights: Credo (CRDO).

  • CPO: erases the "distance" distinction between scale-up and scale-out altogether. Single-mode silicon photonics interfaces inherently reach several kilometers (typically specified at 2 km), enough to connect close to a million racks on a single floor. Once GPU interfaces move to CPO, scale-up domain size is no longer bound by distance but by latency, fan-out and switch radix - with 5 ns/m fiber latency and a radix of 1024, a single-hop shared-memory domain could in theory reach 1,024 GPUs.

That is CPO's real value to the system: it isn't "swapping copper for light", it's tearing down the spatial ceiling of scale-up entirely.


6. Figure 3: Full Fan-Out Networks - Why Single-Wavelength Beats WDM

This figure presents an argument that cuts deep into the optical-module camp - AI clusters need fan-out, not fat pipes.

In NVL72, each GPU isn't connected to its peers through one big pipe; it fans out to a large number of switches so any two GPUs can talk in a "single hop". The key number: each point-to-point link is only 2 x 200 Gbps = 400 Gbps full duplex, in both scale-up and scale-out. This has two decisive implications for technology choice -

  1. Using massive wavelength-division multiplexing (WDM) to aggregate a GPU's bandwidth into one fat pipe is fundamentally ill-suited to AI clusters.

  2. Massively parallel low-speed transmission, on the other hand, is inherently incompatible with standard network protocols (Ethernet, InfiniBand, NVLink, PCIe), and doesn't fit either.

The figure uses a fan-out of 512 xPUs across 16 switches as an example: with a single-wavelength, multi-fiber approach, one simple "fiber shuffle" achieves the desired topology; with a 16-wavelength WDM approach, every link must be demultiplexed, shuffled and remultiplexed - adding a pile of complexity and optical loss out of nowhere.

In AI clusters, aggregating bandwidth into fat pipes is the wrong move. Single-wavelength, spatially parallel is the right direction.

There's an intriguing industry tension here: Winzer clearly bets on single-wavelength, yet NVIDIA presented a DWDM laser-array approach at ECTC. I broke down that route battle in ECTC 2026 | NVIDIA | Light Shouldn't Be Forced into Electrons and Back; reading the two side by side makes both sides of this bet clearer.

Figure 3: Full fan-out topology of 512 xPUs x 16 switches - top, logical topology; middle, single-wavelength approach (fiber shuffle only); bottom, WDM approach (demux/shuffle/remux), with much greater complexity and optical loss. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 3
Figure 3: Full fan-out topology of 512 xPUs x 16 switches - top, logical topology; middle, single-wavelength approach (fiber shuffle only); bottom, WDM approach (demux/shuffle/remux), with much greater complexity and optical loss. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 3

7. Figure 4: The Real Trick of LPO and CPO - "Pretending to Be Copper"

The table on the left plus the chart on the right are the core of understanding "why LPO/CPO rather than conventional optical modules".

The left table lists the energy and reach of each interface: passive copper 0 pJ/bit (but only 1 m), active copper about 1.6 pJ/bit (3-4 m), retimed optics nearly 20 pJ/bit, Linear Pluggable Optics (LPO) about 10 pJ/bit, and CPO about 6 pJ/bit (including the laser, 2 km).

Why such a large gap? Winzer points to the key insight: make the optical interface "look like a passive copper cable" to the xPU's SerDes. Modern digital SerDes equalization is so powerful that it can correct not only PCB trace distortion but also the distortion of the optical link itself and of the electrical path between the far-end receiver and the other ASIC. If the SerDes can handle equalization on its own, the expensive, power-hungry retimer chip in the module is redundant - remove it and you get LPO; drop the package too and attach directly to the ASIC, and you get CPO.

Advanced CPO already reaches 6 pJ/bit today (including the laser, with a 3 dB optical link budget). If the laser is disaggregated to a better-cooled location, the "CPO engine" sitting next to the host ASIC can drop to about 4 pJ/bit - roughly the same as the SerDes driving it. For a B200: 4 pJ/bit extra on 7.2 Tbps of scale-up I/O works out to about 30 W, just 3% of the GPU's 1,000 W - practically negligible.

The right chart is a warning: optical-module energy has long improved at a pace of "10x every 10 years", but in the 8x100G and 8x200G generations it suddenly fell off the historical trajectory and stopped improving. To keep going down, the retimer has to go - that's LPO. This isn't optional; it's a necessary step to stay on the energy curve. Ultra-low-power CPO at 6 pJ/bit puts the curve right back where it belongs.

Figure 4: Left table compares passive copper (0 pJ/bit / 1 m), ACC (1.6 / 3-4 m), retimed optics (19 / 2 km), LPO (10 / 2 km) and CPO (6 / 2 km); right chart shows the 10x-per-decade energy trajectory stalling at the 8x100G / 8x200G generations, with LPO/CPO bringing it back on track. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 4
Figure 4: Left table compares passive copper (0 pJ/bit / 1 m), ACC (1.6 / 3-4 m), retimed optics (19 / 2 km), LPO (10 / 2 km) and CPO (6 / 2 km); right chart shows the 10x-per-decade energy trajectory stalling at the 8x100G / 8x200G generations, with LPO/CPO bringing it back on track. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 4

8. Figure 5: Nubis XT1600 - Pushing Beachfront Density to 230 Gbps/mm

This figure shows Nubis' own XT1600, a 16 x 100 Gbps (full-duplex) near-packaged optics (NPO) / CPO module, and the paper's only "physical deliverable".

Why is density the Achilles' heel? Because today's architectures use "edge escape": all I/O must squeeze out along the chip's edge, its "beachfront", so beachfront density (Gbps/mm) is the key metric. The numbers: pluggable optics top out at about 70 Gbps/mm (1.6T OSFP); advanced digital SerDes reach about 1 Tbps/mm; short-reach die-to-die interfaces like UCIe reach about 5 Tbps/mm. If optics stay at 70 Gbps/mm, they become the bottleneck of the whole system.

XT1600's answer is to use "area" to make up for "shoreline": its transceiver silicon photonics die is only 6.9 x 8.5 mm yet achieves 230 Gbps/mm of net beachfront density; with vertical fiber escape the modules can also be stacked in 2D (N rows), approaching 1 Tbps/mm at N=4 - right in the league of digital SerDes. The fully packaged pluggable module is 15 x 15 mm and still delivers over 100 Gbps/mm net. The figure also shows six modules tiled in 2D into an ultra-dense 10 Tbps cluster, and five modules on a PCIe card delivering 8 Tbps.

Figure 5: Nubis XT1600 high-density, low-power NPO/CPO module - left, a single electrically pluggable module (16 x 100 Gbps full duplex); center, six modules tiled in 2D for 10 Tbps; right, five modules on a PCIe card for 8 Tbps. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 5
Figure 5: Nubis XT1600 high-density, low-power NPO/CPO module - left, a single electrically pluggable module (16 x 100 Gbps full duplex); center, six modules tiled in 2D for 10 Tbps; right, five modules on a PCIe card for 8 Tbps. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 5

9. Figure 6: No All-Rounder - What's Missing Is a Standardized Socket

This technology trade-off table wraps up the paper - it scores every I/O option against seven metrics with checks and crosses, and the conclusion is harsh: no single technology meets them all.

Item by item: retimed pluggables (DSP, LRO) lose on power, retiming latency and density; LPO solves latency and power by removing the retimer but loses on density; passive copper (DAC) loses on reach and density, while fly-over copper fixes density at the cost of connector standardization; on the NPO/CPO side, fast I/O meets almost every criterion, with the only gap being "no standardized mechanical socket interface"; slow and ultra-slow CPO/NPO, meanwhile, are incompatible with scale-up protocols.

So what's really holding back volume production isn't that optics can't be built, but that there's still no common mechanical socket. The good news: the Optical Internetworking Forum (OIF) is developing a standard that lets fly-over copper and NPO/CPO share the same mechanical socket. Once it lands, the combination of the two will become a key piece of future AI cluster networks.

This "standardized socket battle" sits at the heart of the new scale-up standards war; for more, see What Is XPO? The Loudest New Scale-Up Standard at OFC 2026.


Figure 6: Trade-off table of electrical and optical I/O technologies - Pluggables (DSP/LRO/LPO), Copper (DAC/Fly-over) and NPO/CPO (Fast/Slow/Ultra-slow) against seven criteria: high density, low latency, low power, cross-rack reach, full radix, E/O and mechanical interoperability; the only gap for Fast NPO/CPO is mechanical socket standardization. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 6
Figure 6: Trade-off table of electrical and optical I/O technologies - Pluggables (DSP/LRO/LPO), Copper (DAC/Fly-over) and NPO/CPO (Fast/Slow/Ultra-slow) against seven criteria: high density, low latency, low power, cross-rack reach, full radix, E/O and mechanical interoperability; the only gap for Fast NPO/CPO is mechanical socket standardization. Source: Winzer & Harding, SPIE Vol. 13374, 133740F (2025) - Figure 6


10. Industry Links: How Far from Volume, and Who Benefits

String the six figures together and Winzer's industry roadmap is actually quite clear:

Near term (now to 1 year): ACC is the transitional solution for extending scale-up - low risk, low cost, growing the scale-up domain by up to 6x - so the copper and signal-conditioning (retimer/gearbox) camps still have meat on the bone. It also explains why market sentiment toward ACC/AEC has been warming recently.

Mid term: LPO goes first - because it's the necessary step to stay on the energy curve, and it doesn't require touching the package, so adoption friction is lowest. The real gating item is that common OIF socket standard.

End game: GPU interfaces move to CPO, the distance divide between scale-up and scale-out disappears, and shared-memory domains are defined by latency and radix. Who benefits? Those who own single-mode silicon photonics, high-density fiber coupling/escape, and external laser sources - precisely the three hardest and most valuable pieces of the CPO BOM.

A word on the downside risk: Winzer writes from Nubis' position (single-wavelength, linear optics), so his argument naturally leans toward "single-wavelength beats WDM". But NVIDIA's DWDM array route hasn't left the stage, and it holds the ecosystem's voice. The single-wavelength vs WDM bet isn't settled yet - when reading this paper, remember where the author is sitting.


Conclusion

The paper's biggest value isn't any single data point but the way it completes the full causal chain of "why AI clusters must go to linear optics": from the parameter scissors gap (Fig 1) -> the three-tier I/O ladder (Fig 2) -> copper's distance wall (the 120 kW price) -> fan-out knocking WDM out (Fig 3) -> "pretending to be copper" to cut power (Fig 4) -> trading area for beachfront density (Fig 5) -> and finally stalling on a standardized socket that doesn't exist yet (Fig 6).

If you take away just one line: the ceiling of this generation of AI clusters isn't compute or memory, but that 1 meter of scale-up copper; whoever first replaces that 1 meter with low-power optics redefines how big a shared-memory domain can be. And the key to that door is called linear optics.


References

  • P. J. Winzer and D. J. Harding, "High-density High-Speed Linear I/O For AI Clusters," Proc. of SPIE Vol. 13374, 133740F (2025). DOI: 10.1117/12.3045631

  • A. Gholami, "AI and Memory Wall," https://arxiv.org/pdf/2403.14123

  • P. J. Winzer, "The Future of Communications is Massively Parallel," J. Opt. Commun. Netw. 15, 783 (2023)

  • S. Williams, A. Waterman, D. Patterson, "Roofline: An Insightful Visual Performance Model," Communications of the ACM, 52(4), 65-76 (2009)

  • NVIDIA GB200 NVL72, https://www.nvidia.com/en-us/data-center/gb200-nvl72/

  • S. T. Le et al., "1.6-Tbps Low-Power Linear-Drive High-Density Optical Interface (HDI/O) for ML/AI," Proc. OFC, Th4C.4 (2024)

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page