top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

2025 OCP APAC Summit | Scale-up for AI: Balancing Compute, Memory and Networking | Broadcom

3 days ago
3 min read

Updated: 14 hours ago

Summary

At the 2025 OCP APAC Summit, Broadcom presented Ethernet-centric scale-up and scale-out solutions for building hyperscale AI training clusters. With the 100 Tbps Tomahawk Ultra switch chip and the AI Fabric Router, it supports distributed computing that scales from a single rack to across data centers, delivering <400 ns latency, 200,000+ GPU clusters and links spanning 100 km. This open-standards approach pushes next-generation AI infrastructure toward higher performance, lower power and simpler topologies.

2025 OCP APAC Summit – Summary illustration: At the 2025 OCP APAC Summit, Broadcom presented

Content

1. Background and Challenges

AI model sizes are growing exponentially, and a single rack can no longer hold all the compute. As demand extends to hundreds or thousands of GPUs/XPUs, systems must move beyond in-rack scaling (scale-up) into cross-rack and cross-data-center scaling (scale-out). This makes network requirements more demanding than ever:

  • High bandwidth: HPM-to-XPU bandwidth has reached 40–100 Tbps

  • Low latency: data exchange must complete within a few hundred nanoseconds

  • High reliability: fewer packet errors and retransmissions

  • Efficiency and scalability: lower optics and copper costs while supporting larger clusters

Background and Challenges illustration
Background and Challenges illustration 2
Background and Challenges illustration 3
Background and Challenges illustration 4
Background and Challenges illustration 5

2. Core Elements of a Scale-up Network

2.1 Current Bottlenecks

  • Copper backplanes have limited reach, capping the number of XPUs in a single domain (currently <100)

  • Existing switches lack sufficient radix and bandwidth

  • As GPU/HPM interfaces rise to 100 Tbps, current topologies struggle to keep up



2.2 Broadcom's Solution

  • Ethernet-based scale-up (SU) architecture:

    • Open standard, co-developed by the OCP community

    • Round-trip latency from XPU through the Ethernet switch <400 ns

    • Of that, the switch accounts for only 250 ns; the remaining 150 ns comes from up/down stack traversal

Core Elements of a Scale-up Network illustration

  • Tomahawk Ultra switch chip:

    • 50 Tbps bandwidth, 250 ns latency

    • Path to a future 100 Tbps version, reducing fiber count and simplifying network tiers

      Core Elements of a Scale-up Network illustration: Path to a future 100 Tbps version, reducing fiber count

3. From Scale-up to Scale-out

3.1 Internal Scale-out (Within the Data Center)

  • A single data center will host 128,000–200,000 GPUs

  • With 100 Tbps switches, a two-tier topology enables a simpler network:

    • 67% fewer optical modules

    • Lower latency

    • Higher reliability (fewer fiber links and hops)

From Scale-up to Scale-out illustration
From Scale-up to Scale-out illustration 2

3.2 Cross-Data-Center Scale-out

  • Reaching million-GPU clusters requires connecting multiple 50–60 MW data centers

  • Broadcom introduced the AI Fabric Router:

    • Supports links spanning 60–100 km

    • Deep-buffer design, multi-chip stacking, integrated HBM

    • Line-rate encryption to secure data across sites

From Scale-up to Scale-out illustration 3


4. Ethernet's Strategic Position

Broadcom emphasized Ethernet's ubiquity, openness and economics:

  • Multi-vendor interoperability

  • Free from proprietary or licensing restrictions

  • Lower cost than proprietary protocols (e.g., NVLink)

  • Continuously extended and optimized through the OCP community



5. Technology Trends and Outlook

  • Interface rates will move from 100G SerDes to 200G and 400G

  • Network topologies will shrink from three tiers to two, cutting latency and power

  • Optics will replace copper as the mainstream, driving hyperscale compute expansion

  • Cross-data-center integration creates a new form of AI supercluster



Conclusion

With its Ethernet-centric scale-up/scale-out solutions, Broadcom strikes a balance between low latency and high bandwidth to meet the needs of next-generation AI supercomputing clusters. Its open standards and collaboration with the OCP community will accelerate the industry's shift from in-rack compute to globally distributed, cross-data-center AI platforms.

Core value includes:

  • Network interconnect latency below 400 ns

  • 100 Tbps-class switch chips that simplify topology

  • Support for 200,000+ GPU clusters and links spanning 100 km

  • Long-term sustainability through open standards and industry collaboration

Over the next few years, with SerDes upgrades and wider adoption of optics, Ethernet will become the mainstream interconnect for hyperscale AI training and drive a transformation of the entire data center architecture.

Recent Posts

See All

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page