top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

ECOC 2025 Tech Focus: OIF on Energy-Efficiency Requirements and Multiple Challenges in AI Applications

3 days ago
2 min read

Updated: 23 hours ago

Introduction

The explosive growth of AI training and inference is driving rapid evolution in data center networks. Whether in-rack scale-up or cross-rack scale-out, optical and electrical interconnect design faces multiple requirements across reach, latency, reliability, and energy efficiency.

OIF, a cross-industry standards organization with more than 160 member companies, works to advance interoperability, energy-efficiency standards, and architectural coordination. At ECOC 2025, OIF shared its energy-efficiency requirements analysis for AI networks and discussed the trade-offs among CPO, CPC, and different link architectures.


Content

1. Link Types in the AI Compute Pod Architecture

OIF divides the links in an AI pod into five categories:

  1. Compute I/O (e.g., PCIe, CXL, UCI)

  2. Memory Interface (e.g., HBM, CXL over PCIe)

  3. Scale-Up (links within GPUs/accelerators; low latency required)

  4. Scale-Out (cross-pod links; longer reach required)

  5. Front-End Ethernet (connecting to networks outside the data center)

👉 This talk focused on 3. Scale-Up and 4. Scale-Out, because they are the most critical to AI compute and face the toughest power and latency challenges.


2. Reach Requirements

  • Scale-Up: today mostly within a single rack, primarily over copper. But next-generation AI clusters must span multiple racks, requiring ~20 m reach.

  • Scale-Out: cross-pod links need reach of ~100 m, which traditional copper struggles to meet — optics must step in.


3. Latency and AI Workload Characteristics

  • AI compute relies on matrix tiling: each accelerator completes part of the task before results are aggregated.

  • If some links have excessive latency, overall GPU efficiency drops.

  • Sources of latency:

    • Propagation delay

    • Error detection and retransmission delay

  • Solution: a very low bit error rate (BER) and an efficient FEC mechanism.


4. Reliability Requirements

  • Link reliability: low error rates must be guaranteed to avoid frequent retransmission.

  • Hardware reliability: if a GPU or link fails, training must restart from a checkpoint, sharply reducing training efficiency.

  • Pluggable vs. CPO:

    • Front-panel (pluggable): easy to maintain, but with higher reliability demands.

    • Co-packaged (CPO): short copper traces and good reliability, but harder to maintain.


5. Bandwidth Density and Energy-Efficiency Targets

  • Hyperscaler requirements:

    • >2 Tbps/mm shoreline density.

    • <4 pJ/bit energy consumption.

  • Design trade-offs:

    • CPO: shortest copper traces, best bandwidth density and energy efficiency.

    • CPC (Co-Packaged Copper): sits between pluggables and CPO.

    • Pluggable: flexible, but with long copper traces and high loss.

👉 OIF's energy-efficiency map shows: CPO has the most potential in the <4 pJ/bit, high-density region, but deployment must overcome manufacturing and serviceability challenges.


Conclusion

OIF's analysis at ECOC 2025 highlights:

  1. AI networks impose stricter requirements on reach, latency, and reliability, and traditional copper is gradually being replaced by optics.

  2. Scale-Up (20 m) and Scale-Out (100 m) are the key scenarios for future energy-efficient design.

  3. Latency is closely tied to BER: tail latency must be controlled through lower error rates and efficient FEC.

  4. Reliability is becoming key to AI efficiency: MTBF must improve dramatically to avoid checkpoint rollbacks.

  5. CPO is the most energy-efficient, but not the only answer: pluggables, CPC, and CPO will coexist depending on the application.

Overall, OIF's message is: AI factories need new energy-efficiency standards and a system-level design philosophy. It's not just about speed upgrades, but about finding the balance among energy efficiency, reliability, and scalability — which will determine the success or failure of future AI networks.


ECOC 2025 Tech Focus – Conclusion illustration
ECOC 2025 Tech Focus – Conclusion illustration 2
ECOC 2025 Tech Focus – Conclusion illustration 3
ECOC 2025 Tech Focus – Conclusion illustration 4
ECOC 2025 Tech Focus – Conclusion illustration 5
ECOC 2025 Tech Focus – Conclusion illustration 6
ECOC 2025 Tech Focus – Conclusion illustration 7
ECOC 2025 Tech Focus – Conclusion illustration 8
ECOC 2025 Tech Focus – Conclusion illustration 9
ECOC 2025 Tech Focus – Conclusion illustration 10
ECOC 2025 Tech Focus – Conclusion illustration 11
ECOC 2025 Tech Focus – Conclusion illustration 12
ECOC 2025 Tech Focus – Conclusion illustration 13

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page