[CPO Deep Dive 1/6] Have Pluggable Optics Hit the Wall? Before You Understand CPO, Understand This Wall
AI keeps stacking up more GPUs, but what runs out of breath first isn't compute — it's the network that "connects" those chips. Pluggable optical modules have carried data centers for two decades, and now they are hitting two walls at once: the power cost of moving each bit, and the bandwidth density a switch faceplate can hold. This wall isn't about "not fast enough" — it's that "go any faster and you can't afford the power." CPO (Co-Packaged Optics) was born to tear down this wall, but it isn't here to replace pluggables; it first gains a foothold in scale-up, a new battlefield where copper can no longer cope. This article breaks down the structure of that wall and lays the foundation for the whole series.
1. Compute Is Soaring, but "Connectivity" Gets Stuck First
For the past two years, all eyes have been on GPUs: how many, how many TFLOPS, how much HBM. But one fact is rarely spelled out — when you tie tens of thousands of GPUs into a single training job, overall efficiency is often determined not by per-chip compute, but by whether they can talk to each other fast enough and efficiently enough.
AI training is an "all-synchronous" workload: at the end of every step, thousands of accelerators must exchange gradients and align parameters. If the interconnect lags by even a beat, expensive GPUs sit idle. So the network has moved from backstage to center stage, becoming the real bottleneck of AI infrastructure.
And carrying this interconnect lifeline are optical modules. The problem: this pluggable-centric optical architecture is approaching its physical limits.
2. Everyone Thinks the Problem Is "Not Enough Bandwidth" — It's Really "Each Bit Costs Too Much"
First, let's dispel an intuitive misconception. Most people think the challenge for optical modules is simply "making them faster" — from 400G to 800G to 1.6T. Speeds are indeed advancing, but that's not the real reason for hitting the wall.
There are really two walls, and neither is speed itself:
The first wall is power. A signal leaving the switch ASIC must first travel an electrical trace (SerDes) before reaching the faceplate optical module for electrical-to-optical conversion. The higher the speed and the longer the distance, the more expensive that electrical transport becomes — and not linearly, but at an accelerating rate. As a single switch's total bandwidth climbs, just "getting the electrical signal to the module" eats an absurdly large share of the power budget.
The second wall is density. No matter how small a pluggable module gets, it still has to plug into the switch faceplate. Faceplate area is fixed, so there's a cap on how many modules you can fit. As the bandwidth each chip demands keeps doubling, the faceplate will eventually run out of room — a physical ceiling that engineering cannot postpone forever.
Translate these two walls into a more precise metric and you get "energy per bit" (usually measured in picojoules per bit). The problem with pluggables isn't that they can't reach high data rates; it's that the energy per bit can't come down when you have to sustain that rate while also driving the electrical signal across that SerDes distance. Data center power is budgeted, and the bigger the share interconnect consumes, the less is left for compute — that's what really keeps hyperscalers up at night: they're not short of bandwidth, they're short of "bandwidth they can afford."
The problem isn't that "light isn't fast enough" — it's that "the electrical signal is dragged too far and burns too much power."
Shorten the signal's electrical path and make the electrical-to-optical conversion happen as close to the ASIC as possible — that is the entire motivation behind CPO. It moves the optical engine from the faceplate onto the same package substrate as the ASIC, trading shorter distance for power and density. As for the sweeping changes in packaging this move brings, we break them down fully in the "Packaging Evolution" installment of this series.

3. Two Networks in One Picture: Scale-out and Scale-up Hurt in Different Ways
To understand where CPO will land first, you need to distinguish the data center's two networks, because they hit different walls.
Scale-out: connects an entire data center's racks and clusters horizontally so traffic flows between racks. This is the traditional Ethernet world, with reach from in-rack to across floors and even several kilometers. The pain here is mainly power and port density.
Scale-up: within a single package or node, binds tightly coupled GPUs, xPUs and memory into one ultra-high-bandwidth, ultra-low-latency compute unit. Reach here is extremely short (centimeters to meters), but the demands on bandwidth density and latency are extreme.
The key difference: scale-up has historically relied on copper. Copper is cheap, mature and reliable, and has long been the go-to for very short distances. But as bandwidth density climbs to the levels AI clusters demand, copper starts to buckle — it's a physics problem: the higher the signal rate, the worse copper's high-frequency losses (skin effect and dielectric loss), and the shorter the distance it can carry. Reaching further means adding retimers and equalizers, which brings the power right back; saving power means shortening the distance, yet GPU clusters keep getting bigger and spread further. Copper is stuck in a corner where reach, speed and power can't all be had at once, and thicker cables can't save it. So "replacing copper with optics" becomes the unavoidable option in scale-up — optical loss is nearly independent of data rate, which is precisely its crushing advantage over copper at high speed and long reach.
This explains an industry phenomenon that's easy to misread: why some are starting CPO in scale-up first. In scale-out, pluggables still have a stack of "life-extension tricks" available (covered in the next section); but in scale-up, copper has reached the end of the road and optics is the only way out. Here CPO isn't "a better option" — it's "almost the only option."

4. Pluggables Aren't Going Quietly: LPO, LRO and the Battle of Life-Extension Tricks
If you think pluggables will meekly yield to CPO, you're underestimating the survival instincts of a mature ecosystem. Facing the power wall, the pluggable camp has rolled out a whole lineup of "life-extension tricks":
LPO (Linear-drive Pluggable Optics): removes the power-hungry DSP in favor of linear drive, saving substantial power at the cost of stricter channel-quality requirements.
LRO (Linear Receive Optics): linearizes only the receive side — a compromise between removing the DSP entirely and keeping it all.
New modulation materials and SiPh: silicon photonics itself, plus new material platforms such as thin-film lithium niobate (TFLN), keep squeezing out room for pluggable energy efficiency.
What these technologies mean: pluggables in scale-out still have several good years ahead. They keep pushing the power wall back, forcing CPO adoption in scale-out to slow down — because as long as pluggables can hold cost and power in check with these tricks, data centers have no incentive to take on the risk and uncertainty of CPO's entirely new packaging.
Meanwhile, NPO (Near-Packaged Optics) and OBO (On-Board Optics) play a transitional role: moving optics one step closer from the faceplate to the ASIC without full co-packaging, giving the ecosystem a waypoint until CPO's packaging, yield and reliability mature.

5. The Unresolved Variable: The 200G/lane vs. 400G/lane Split
On the road to ever-higher speeds, there's another unconverged technical split worth keeping in view: whether a single lane should go 200G or 400G.
The higher the per-lane rate, the fewer lanes needed for the same total bandwidth — in theory more efficient and denser. But higher rates raise the difficulty bar on signal integrity, modulation materials, and lasers and modulators. The prevailing industry view: 200G/lane is relatively mature and highly compatible with current AI infrastructure, so it will be the near-term workhorse; 400G/lane faces tougher technical challenges and its adoption will lag. This split affects not only pluggables but directly drives the spec choices for CPO optical engines — which rate you bet on determines how you pair your lasers, modulators and packaging route.
What this means for readers: don't let the "faster is better" instinct carry you away. Every doubling of rate sets off a chain of material and packaging problems — modulators must be more linear, lasers more stable, packaging signal integrity tighter, and testing harder. The industry's real deployment order is usually "mature first, fastest later," not "fastest first." So when you see a company announce a higher per-lane rate, it's worth asking: is this "done in the lab," or "stable on the production line at a cost that works"? The two are often years apart.

6. Where Things Stand: CPO Opens an Additional Battlefield, Not a Replacement Race
Bringing this wall back to the current state of the industry, here's how to position it:
CPO isn't here to sweep pluggables into history. In fact, the pluggable market will keep growing for the foreseeable future — CPO should be understood as "an additional battlefield," not a zero-sum replacement race. It will first enter scale-up, where copper can't cope, and the high-end scale-out scenarios where power density is tightest, coexisting with pluggables for the long term.
And the threshold CPO really has to clear isn't only technical: to win data center buy-in, it must prove three things — that a multi-vendor business model works, that cost falls with scale, and that reliability and serviceability reach production grade. If any one of these falls short, adoption will keep slipping.
So once you understand this wall, you have the coordinates to read the entire series: the packaging, light sources, coupling, supply chain and vendor roadmaps that follow all answer the same question — to move electrical-to-optical conversion as close to the chip as possible, which puzzle pieces are we still missing?
Summary
AI's bottleneck is shifting from "can we compute fast enough" to "can we connect efficiently enough." Pluggable optical modules aren't hitting a speed wall but a power wall and a density wall — electrical transport per bit is too expensive, and the bandwidth a faceplate can hold has topped out. CPO's entire logic is to move electrical-to-optical conversion as close to the ASIC as possible, trading shorter distance to win back both. But it doesn't replace pluggables; it first gains a foothold in scale-up, where copper has reached its end, then seeps into high-end scale-out scenarios, all while being slowed by life-extension tricks like LPO/LRO and transitional technologies like OBO/NPO. Next time you see a headline saying "CPO is coming," the question to ask is: this time, which scenario's wall is giving way first? Understand the wall, and you'll understand the timing.




Comments