OFC 2026 - A Paradigm Shift in AI Cluster Architecture: Scale-Out and Scale-Up Optical Interconnects Explained - Broadcom
At OFC 2026, the talk by Anand Ramaswamy, solutions architect in Broadcom's Optical Systems Division, set the tone for the future of AI infrastructure. As million-GPU clusters move from blueprint to reality, traditional electrical (copper) interconnects are running head-on into physical limits. This article breaks down how Broadcom is using CPO (co-packaged optics) and the newly formed OCI (Optical Compute Interconnect) MSA to redefine the back-end network and scaling architecture of AI clusters.
Core Technology: CPO's Strength in Scale-Out Networks
Broadcom divides AI networking into three layers: the front-end network, scale-out (back-end network), and scale-up (accelerator interconnect). In scale-out, bandwidth requirements are jumping rapidly from 400G and 800G to 1.6T, making power management critical.
1. A Step-Change in Power: The Push to 6 pJ/bit
According to Broadcom's data, at 100G/lane the power gap between architectures is huge:
DSP pluggable (traditional pluggable module): about 15 pJ/bit.
LPO (linear pluggable optics): about 10 pJ/bit.
CPO (co-packaged optics): only 6 pJ/bit.
That means CPO cuts power by about 65% versus traditional optical modules, and saves 35% versus LPO. For AI data centers with millions of optical links, this translates directly into megawatt-scale power savings.
2. TH5-Bailly: From the Lab to 50 Million Hours of Stable Operation
Broadcom's second-generation CPO system, TH5-Bailly (a 51.2T switch), has demonstrated remarkable reliability.
Hardware: eight built-in 6.4T optical engines, each integrating more than 1,000 optical components.
Field data: trial results published by Meta in 2025 showed the Bailly CPO system achieved zero link flaps during its first 1 million hours of operation. In his talk, Anand updated the figure: cumulative operation has reached 50 million hours with no unrecoverable failures.
This is a strong rebuttal to market doubts about CPO serviceability and reliability. By moving the laser outside the package (remote laser source, ELSFP), Broadcom addresses the pain points of thermal management and field repair.
A New Landscape: OCI MSA Targets Copper's Last Stronghold
Scale-up (such as NVIDIA's NVLink architecture) is still dominated by copper (DAC/ACC), but as per-rack power soars and cabling complexity grows, copper has hit its limit. An NVIDIA GB200 NVL72 rack weighs up to 700 kg and contains 5,000 copper cable pairs - the space and cooling challenges are obvious.
OCI MSA: Specs and Ambitions
To break through copper's physical barrier, Broadcom teamed up with major players to found the OCI (Optical Compute Interconnect) MSA. Its core logic is finding the balance point between cost and performance:
Modulation: uses 53G NRZ rather than the more complex PAM4, to reduce latency and power.
Transmission: 4-wavelength WDM (wavelength-division multiplexing) plus Bi-Di (bidirectional transmission), halving the fiber count.
Target: keep total power below 10 pJ/bit (including SerDes), so it can compete head-on with copper on cost and power.
Market Impact: Generational Evolution of Switch Nodes
Broadcom has already published the roadmap for its next-generation switch, TH6-Davisson:
Throughput: 102.4 Tbps.
Interface speed: raised to 200G per lane.
High-radix advantage: 512 x 200G lanes, meaning more GPUs can be connected in a single-tier switching network, sharply reducing cluster latency and total cost of ownership (TCO).
Simple Tech Trend View: The "200G Tipping Point" for Optical Interconnects Has Arrived
2026 is the pivotal year in which the optical communications industry shifts from "component supply" to "system integration."
CPO is fully mature: with Meta's large-scale deployment data, Broadcom has shown that CPO is no longer an academic topic but a commercial technology in volume production. 12M equivalent hours of testing plus 50M hours in the field are the strongest ammunition for convincing hyperscalers to give up pluggables.
Scale-up is the next gold mine: the founding of the OCI MSA marks optics' formal entry into short-reach (<10m) GPU-to-GPU interconnects. While 53G NRZ may look "slow," under the demands of ultra-low power and high radix it is currently the optimal solution for volume production.
Supply chain impact: with Broadcom shipping more than 50 million laser chips per year, silicon photonics (SiPh) economies of scale will push costs down further. This will pressure traditional DSP vendors (such as Marvell) and push the industry toward lighter-weight, linear optical solutions (LPO/CPO).
We expect more AI chipmakers (especially ASIC vendors) to integrate the OCI specification into their SoC packages, and copper's role in AI clusters will formally shrink to ultra-short in-rack links only.
































Comments