OFC 2026 – Scaling AI Clusters: How Scale-Up and Scale-Out Optical Architectures Are Evolving – OpenAI / AMD / NVIDIA / Broadcom / TeraHop / Coherent
As AI workloads shifted structurally in 2025, moving beyond pure large-model training into a diverse mix of long-context, reasoning-heavy and agentic workloads, the underlying infrastructure is facing unprecedented pressure. At the Data Center Summit panel at OFC 2026, experts from OpenAI, AMD, NVIDIA and other industry leaders agreed: "I/O bandwidth density" and "power" have become the core bottlenecks to scaling AI clusters. This article breaks down the competing technical paths each major player is taking on scale-up and scale-out network architectures.
Core Technologies and Data: How Six Industry Giants Line Up
1. OpenAI: Defining AI Demand and Launching OCI-MSA
OpenAI's Binbin Guan noted that AI demand in 2025 is no longer just a story about model size: reasoning tokens grew 320-fold.
Technical position: Copper remains mainstream at 224G/lane, and 448G/lane is seen as "feasible but extremely challenging." Still, OpenAI is calling on the industry to deliver 2X bandwidth-density gains every generation and to push I/O power below <2 pJ/bit.
Key move: OpenAI co-founded OCI-MSA (Optical Compute Interface) to drive an open-standards-based "optical scale-up" architecture that extends scale-up networks from within a rack to multi-rack deployments.









2. AMD: CPO Is a System-Level Co-Design Challenge
AMD's Juthika Basak stressed that CPO (Co-Packaged Optics) is not just a component but a system enabler.
Technical data: Compared with traditional pluggable modules (>15 pJ/bit), CPO delivers roughly 3X better energy efficiency, bringing power down to ~5 pJ/bit.
Core view: For AMD, reliability, availability and serviceability (RAS) come first. To tackle thermal drift in micro-ring modulators (MRM), AMD argues that design must combine liquid cooling with multiphysics thermal modeling in a co-design approach.








3. NVIDIA: AI Factory Specs in the Blackwell Era
NVIDIA's Meer Sakib revealed the staggering specs of a 512K Blackwell GPU cluster.
Cluster data: A single AI factory draws 600MW; its internal L1 network runs over NVLink at 900GB/s unidirectional bandwidth, while scale-out requires roughly 1.8M optical transceivers.
Reliability advantage: NVIDIA's data show CPO delivers 3.5X better energy efficiency and 10X greater resiliency, significantly cutting the US$3 million daily loss caused by network failures (for a 512K cluster).







4. Broadcom: From Copper's Limits to 2.6M-Hour MTBF Reliability
Broadcom's Anand Ramaswamy gave a precise prediction of when optics will replace copper.
Data anchor: Copper (DAC/ACC/AEC) currently holds the line at <10 pJ/bit. Broadcom believes optical scale-up only becomes competitive once total optical interconnect power falls below 10 pJ/bit.
Proven reliability: In Meta's field testing, Broadcom's Bailly CPO system achieved an MTBF (mean time between failures) of 2.6M hours, keeping training efficiency on a 24K-GPU cluster above 90%.








5. TeraHop: A Volume-Production Breakthrough with the 12.8T XPO
TeraHop's Ryan Yu showcased the industry's first 12.8T XPO (Pluggable Optical Engine).
Specs: The module integrates 64 lanes of 200G, equivalent to eight 1.6T OSFP modules, supports module power of up to 400W, and integrates liquid cooling.
Supply-chain value: Its silicon photonics technology has accumulated more than 70 billion hours of device operation, with PIC FIT below 0.1, demonstrating very high yield and reliability.










6. Coherent: InP Capacity and Optical Circuit Switching (OCS)
Coherent's Steffen Koehler urged the industry to watch upstream materials.
Capacity warning: As demand for InP (indium phosphide) surges for EML and CW lasers, capacity will become extremely tight. Coherent has moved early to 6-inch InP wafers, targeting a doubling of output in 2026-2027.
New tool: OCS will no longer be limited to Google's front-end network; it will spread to six major use cases, including dynamic AI resource allocation and workflow optimization.






Consensus and Divergence: The Battle of Paths
Consensus: pluggables and CPO will coexist for the long term. CPO has the power advantage, but pluggable modules still win on serviceability flexibility and ecosystem breadth.
Divergence #1: transmission strategy. The industry is split into two camps: "Fast and Narrow" (400G+ per lane, relying on external lasers) and "Slow and Wide" (massively parallel lanes as promoted by OCI-MSA, with power as low as 2 pJ/bit).
Divergence #2: laser architecture. Whether to integrate the laser inside the module or push ELSFP external light sources to reduce failure risk and thermal load is still weighed differently by each player.
The Simple Tech Trend View:
Based on the data and timelines disclosed at this session:
Silicon photonics (SiPh) hits its breakout year: TeraHop's 12.8T XPO and Broadcom's CPO field data show that SiPh has resolved the reliability doubts of the past. Before the end of 2026, we expect a wave of high-volume 1.6T/3.2T SiPh solutions to enter pilot runs.
Liquid cooling becomes standard for optics: As optical engine power density climbs with 400G/lane, air cooling can no longer meet the needs of heat-sensitive devices such as MRMs. Over the next 18 months, pluggable optics with integrated liquid-cooled cold plates will become a key procurement criterion for high-end switches.
Packaging is the competitive edge: Leaders such as NVIDIA and Broadcom both emphasize fan-out wafer-level packaging and co-design. That is a long-term positive for advanced-packaging supply chains such as TSMC (CoWoS/SoIC), as optoelectronic integration shrinks from the PCB level all the way down to the package substrate.




Comments