top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

OCP Global Summit 2025 | Broadcom & Arista | The Scale-Up Ethernet (SUE) Framework for AI/ML Accelerators

3 days ago
4 min read

Introduction

Network architecture in the AI era is moving from scale-out to scale-up.

Broadcom and Arista presented a new SUE (Scale-Up Ethernet) framework at OCP 2025,

aiming to extend Ethernet into high-bandwidth, low-latency interconnect between accelerators

and make it the foundational protocol for memory sharing and cooperative compute among GPUs/XPUs.

Broadcom senior architect Mohan Kalkunte said in the talk:

"AI networking today has three tiers: scale-up, scale-out and scale-across. Ethernet already dominates scale-out, and now it's ready to take over scale-up as well."

The talk showed how Ethernet, through architectural layering, protocol refinement and open standardization,

can break the closed GPU fabric model and let different accelerator vendors interoperate in a single ecosystem.


Content

1. From Scale-Out to Scale-Up: The Next Phase of AI Networking

In the past, the main challenge of AI training was connecting thousands of GPUs (scale-out).

Now, the number of GPUs within a single cluster has passed several hundred,

and high-speed intra-rack interconnect (scale-up) has become the system performance bottleneck.

Broadcom draws a clear distinction between these network tiers:

Network type

Typical use

Connections

Reach

Core technology

Scale-Up

Intra-rack GPU interconnect

10–100+

<10m

Copper / Short-Reach Optics

Scale-Out

Inter-rack cluster interconnect

1,000–100,000+

10–100m

Ethernet Fabric

Scale-Across

Data center interconnect

100K–1M GPUs

km-class distances

Coherent Ethernet / WAN Overlay

SUE (Scale-Up Ethernet) is designed for the first tier, the rack-scale fabric,

letting GPUs share memory and move tensor data over an Ethernet-based architecture

to achieve a compute model where a cluster acts like one giant GPU.


2. Why Use Ethernet for Scale-Up?

Traditional GPU fabrics (such as NVLink and Infinity Fabric) are proprietary architectures:

  • hard to integrate across vendors;

  • closed protocols with a packet layer that is hard to extend;

  • high physical-layer cost and a lack of standards.

Broadcom points to three major advantages of building on Ethernet:

  1. Scalable: reuses existing Ethernet PHY/MAC, switch silicon and test infrastructure.

  2. Open & Modular: lets each accelerator vendor run its own transport-layer protocol.

  3. Cost-Optimized: leverages the mature copper / AEC / optics ecosystem.

Broadcom has already deployed several SUE prototypes in internal testing,

some of which are integrated into Tomahawk 5 / 6 switches and CPO modules.


3. SUE Architecture: Separating Transport and Network

Broadcom stresses that SUE's key innovation is decoupling the transport layer from the network layer.

● Lower layer: Ethernet Networking (ESAN, Ethernet Scale-Up for Networking)

  • Defined at the link, MAC and PHY layers, led by Broadcom and Arista.

  • Core functions include:

    • Link-Level Retry (LLR)

    • Thread-Based Flow Control (CBFC)

    • Optimized Headers & Lightweight Framing

    • Low-Latency Deterministic Jitter Control

● Upper layer: SUE Transport

  • Left open for GPU/XPU vendors to implement freely.

  • Specific features can be chosen per application, for example:

    • Transaction Packing

    • Reliability Layer (hop-by-hop or end-to-end)

    • Memory Ordering Models

    • Congestion & Load Balancing Policies

    • Encryption / Security Options

    • Lightweight Retransmission (Go-Back-N)

This design makes SUE a "menu of choices" protocol,

so each vendor can pick and choose features based on its compute architecture, silicon resources and workloads.


4. High Bandwidth, High Efficiency, Low Power: Three Principles of AI Fabric

Broadcom notes that in a GPU cluster each GPU has 4–8 HBM stacks, for total bandwidth of up to 100 TB/s.

For multiple GPUs to share memory, interconnect bandwidth must reach at least 1/10 of HBM bandwidth.

SUE's design therefore follows three core principles:

  1. Bandwidth Density:

    • 8–12× higher than scale-out Ethernet;

    • target I/O of 3.2–6.4T per XPU.

  2. Power Efficiency:

    • must be embedded next to the GPU die, at under 5W per port;

    • traditional RDMA NICs can't be used, as their area and power are too high.

  3. Reliability & Determinism:

    • low latency and low jitter are design priorities;

    • occasional packet errors are handled with lightweight retransmission, without TCP-style overhead.


5. Implementation Examples and Topologies

Broadcom showed two 128-XPU cluster architectures:

Topology

Switch chip

GPU I/O

Port speed

Result

Option A

Tomahawk 5

3.2T

400G

Single-hop multi-plane architecture

Option B

Tomahawk 6

6.4T

800G

Double the bandwidth, low latency

SUE can already support single-tier clusters of a hundred-plus GPUs,

and as low-cost short-reach optics mature,

it is expected to extend to cross-rack multi-rack scale-up fabrics.


6. Open Standardization: The SUE Consortium and Working Groups

Broadcom has launched two main workstreams:

Workstream

Name

Function

Lead

SUE Transport WG

Scale-Up Ethernet Transport

Defines the XPU transport-layer protocol

Broadcom + Arista

ESAN WG

Ethernet Scale-Up for Networking

Defines MAC/PHY-layer functions

Led by Broadcom

EAN WG (upcoming)

Ethernet AI Networking

Integrating scale-out and scale-up

Multiple vendors participating

Broadcom called on more vendors to join the open collaboration

to ensure Ethernet becomes the universal AI fabric technology,

rather than just a server-to-server communication protocol.


Conclusion

The SUE Framework presented by Broadcom and Arista at OCP 2025

signals that Ethernet is moving from the data center backbone into the heart of the GPU.

"Ethernet is no longer just a network; it will become the unified medium for memory and compute."

Through an open, modular, low-power architecture, SUE lets accelerators interoperate in a standardized way

while preserving flexibility and differentiation. Broadcom expects future AI racks to be built around SUE + UALink,

delivering a truly open, scalable scale-up fabric.


Further Perspectives

  1. Technical impact

    • SUE is a key milestone in Ethernet's evolution into the compute interconnect layer.

    • Together with UALink (Layers 1–3) and UEC (Layers 4–5), it forms a complementary, complete AI fabric stack.

  2. Supply chain observations

    • Through Tomahawk 6 + CPO combined with Arista's software stack, Broadcom forms an integrated hardware-software ecosystem.

    • In the future, ODMs/OEMs (such as Supermicro and Dell) can build GPU racks directly on the open SUE architecture.

  3. Market trends

    • AI infrastructure is shifting from proprietary fabrics to open Ethernet.

    • Between 2026 and 2028, SUE + UALink + CXL may become the mainstream three-tier open architecture.

Recent Posts

See All

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page