top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

OFC 2026 - AI Interconnect Scale-Up and CPO Reliability Deep Dive - Meta Platforms

2 days ago
3 min read

At OFC 2026, Meta Platforms' Andrew Alduino delivered a heavyweight talk on the evolution of AI interconnect architecture. As generative AI's appetite for compute moves into the "Superintelligence" era, Meta revealed its latest experimental data on optical scale-up networks and co-packaged optics (CPO). The talk not only redefined the reliability benchmark for CPO, but also pointed the way for 1.6T and higher-speed interconnect solutions.


The 5GW Era: From Manhattan-Scale Sites to GB300 Clusters

Meta is building infrastructure to serve 3.4 billion daily active users (DAU). To pursue superintelligence, Meta has committed hundreds of billions of dollars to compute, with 2026 capital expenditure (CapEx) expected to exceed US$115 billion.


Data Centers the Size of Cities

Meta disclosed its data center project codenamed Hyperion, a single Louisiana site expected to reach 5GW of capacity. A comparison slide showed its footprint covering a large part of Manhattan. [💡 Editor's note: insert the data center footprint comparison from slide 2 here]


Rack Architecture Pushed to the Limit

On the interconnect hardware side, rack architecture is changing fast:

  • GB300 era: a single rack hosts 72 accelerators, using a copper backplane and liquid cooling.


  • 144-node double-wide rack: copper (with retimer assistance) can just about stretch to 144 nodes, but once scale reaches 256 nodes and beyond, copper's power, weight and physical limits become insurmountable bottlenecks.

  • Optical scale-up: this is the core driver behind Meta's push for CPO and OCI (Optical Compute Interconnect).


CPO vs. Pluggables: The Price of 65% Power Savings

Andrew Alduino noted that choosing an optical technology is fundamentally a trade-off across metrics.


According to Meta's test results, CPO shows an overwhelming advantage in performance and efficiency:

  • Power savings: a CPO link saves 65% of the power of a conventional retimed pluggable, and still saves 35% versus LPO (linear pluggable optics) at $100Gbps/lane$.


  • System-level benefit: in a 51.2T switch system, CPO saves more than 500W of power.


CPO still faces challenges in maturity, serviceability and ecosystem, but its low latency and high density make it the inevitable choice for scale-up domains of 256 nodes and beyond.


90 Million Hours of Field Data: Busting the "CPO Is Unreliable" Myth

The industry's biggest concern about CPO has long been its large "blast radius" when something fails. Meta published large-scale reliability data from its Bailly CPO system, with a sample of 90 million cumulative device-hours (400G-port equivalent).


Reliability Head-to-Head (MTBF)

Solution

Test condition

Cumulative device-hours

MTBF (million hours)

2x400G FR4 pluggable module

40degC stress test

~8M

0.71M

CPO Phase 1 (all failures)

40degC stress test

>40M

1.47M

CPO Phase 1 (excluding PLS-specific issue)*

40degC stress test

>40M


8.2M

CPO Phase 2 (non-serviceable components)

Room temperature

>50M

Too few failures to compute

*Note: the PLS issue was traced to degradation of an SMT component in the laser driver circuit, not a defect in the CPO architecture itself.

Failure Pareto

Meta found that the main failure source in Phase 1 CPO was the driver circuit of the ELSFP (external laser source). Excluding this known component issue, which can be fixed by redesign, CPO reliability (MTBF) is more than 10x that of pluggable modules.


Notably, for both pluggables and CPO, "dirty fibers and connectors" remain a shared challenge for stability.


Supply Chain View: From Closed to Open with the OCI MSA

To accelerate deployment of optical scale-up, Meta has teamed up with five partners to drive the OCI MSA (www.oci-msa.org).


  • NVIDIA and AMD: Meta has established deep collaboration with both GPU giants to ensure compatibility between compute architecture and optical interconnect.


  • Corning: provides capacity support for fiber infrastructure.


Meta's strategy is clear: run its in-house MTIA accelerators in parallel with partnerships with major vendors, and use CPO to extend AI cluster connectivity from rack scale to data center scale.


Andrew Alduino's OFC 2026 report is a landmark. CPO discussions used to stay at the academic level of "power savings", but Meta's 90 million hours of data make a strong case for the engineering rule that "higher integration means higher reliability".


Outlook:

  1. Copper-to-optics crossover arrives early: as ultra-dense racks like GB300 proliferate, copper's weight and thermal burden will push hyperscalers to start small-scale CPO/OCI cluster deployments before the end of 2026.

  2. ELSFP becomes the battleground: since laser source failures are the main pain point, vendors with strong external laser module capabilities (such as Broadcom and Lumentum) will gain a more critical position in the supply chain.

  3. Testing moves upstream: CPO requires system-level testing at the packaging stage, which will reshape profit distribution across the optical supply chain and give OSATs and switch ODMs a much bigger role.


"Small packages are more reliable than large packages" — this will become the motto of AI infrastructure designers for years to come.



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page