ECOC 2025 Tech Spotlight: Meta Evaluates Co-Packaged Optics (CPO) in Hyperscale Data Centers
Updated: 21 hours ago
Introduction
With AI model sizes and GPU counts growing rapidly, data center IO bandwidth demand is far outpacing gains in compute capability. As a result, IO power takes a rising share of total energy consumption, squeezing the power available for compute.
At ECOC 2025, Meta presented a comparative study of CPO and pluggable optical modules, examining from the angles of power, reliability and serviceability whether CPO can become the mainstream technology for the next generation of hyperscale fabric switches.
Content
1. Background: Network Requirements for Scale-Up and Scale-Out
Scale-Up (GPU interconnect within the rack): short reach (meter-class), mostly served by copper, but as data rates rise, channel loss and power problems worsen.
Scale-Out (across racks / data centers): reach of 10–100 meters or even kilometers, currently relying mainly on pluggable optical modules.
Challenges:
Power improvement is slowing in the 1.6T generation, and power is not expected to drop meaningfully in the 3.2T generation.
ecoc-2025-tech-spotlight-meta-evaluates-co-packaged-optics-cpo-in-hyperscale-data-centers
Rack design is constrained by the bandwidth × distance (bandwidth-reach product), sharply increasing pressure on power density and reliability.
2. Technology Comparison: Pluggable vs. LPO/LRO vs. CPO
Pluggable:
Mature, with a complete ecosystem and serviceability.
Drawback: power and density limits are becoming increasingly apparent.
LPO / LRO:
Advantages in power and cost while keeping the serviceability flexibility of pluggables.
Challenge: needs more mature interoperability and a broader ecosystem.
CPO:
Advantages: lower power, higher density.
Drawback: optical engines cannot be replaced in the field, so reliability and maintenance models need to be re-evaluated.
3. Meta's Experiments and Data
Meta built a large-scale hardware and software test platform and, under elevated-temperature stress, compared Pluggable (MiniPack3) with CPO (Broadcom Bailly 51.2T):
Power testing:
Compared with pluggables, CPO saves 65% of power.
Compared with LPO, CPO saves 35% of power.
Temperature has less impact because the lasers sit in replaceable modules.
Reliability testing:
More than 15M device hours accumulated (CPO), versus 2M device hours (pluggable).
CPO MTBF: 2.6M hours, about 5× that of pluggables.
No unserviceable failures occurred in the first 4M hours.
Failure case analysis: CPO failures were mostly laser modules or fiber-core contamination, with no need to replace the whole system.
Diagnostics and serviceability:
Example: contaminated fiber core → detected via the reflection metric (MPI) and restored by cleaning.
Example: abnormal laser bias → fixed by replacing the laser module.
Conclusion: CPO reduces human plugging errors and raises overall availability.
4. How CPO Relates to AI Training Efficiency
Cluster MTBF and efficiency:
LLM training requires periodic checkpoints; a mid-run failure forces a rollback, severely hurting efficiency.
Simulations show that when MTBF improves 5×, training efficiency of large clusters rises significantly.
Conclusion: CPO not only lowers power but, by improving reliability, also indirectly raises AI training efficiency.
5. Other Perspectives
Socket vs. Solder: Meta believes that whether socketed or soldered, there is almost no difference in CPO's in-field serviceability.
High-power laser risk: no failures observed so far, but as power rises, new failure modes may emerge and need continued monitoring.
Summary
Meta's research shows:
CPO's power advantage is clear: 65% lower than pluggables and 35% lower than LPO.
Reliability is greatly improved: MTBF rises to 5× that of pluggables, with no unserviceable failures.
A direct contribution to AI training performance: higher MTBF means higher cluster operating efficiency.
Challenges remain: the reliability of high-power lasers and progress on standardization need continued monitoring.
Overall, Meta's signal is clear: CPO is no longer just a lab concept but a viable option for the back-end network of AI data centers, delivering breakthroughs in both power and reliability.
Presentation Slides

















Comments