OCP Global Summit 2025 | Broadcom & Arista | The Scale-Up Ethernet (SUE) Framework for AI/ML Accelerators
Introduction
Network architecture in the AI era is moving from scale-out to scale-up.
Broadcom and Arista presented a new SUE (Scale-Up Ethernet) framework at OCP 2025,
aiming to extend Ethernet into high-bandwidth, low-latency interconnect between accelerators
and make it the foundational protocol for memory sharing and cooperative compute among GPUs/XPUs.
Broadcom senior architect Mohan Kalkunte said in the talk:
"AI networking today has three tiers: scale-up, scale-out and scale-across. Ethernet already dominates scale-out, and now it's ready to take over scale-up as well."
The talk showed how Ethernet, through architectural layering, protocol refinement and open standardization,
can break the closed GPU fabric model and let different accelerator vendors interoperate in a single ecosystem.
Content
1. From Scale-Out to Scale-Up: The Next Phase of AI Networking
In the past, the main challenge of AI training was connecting thousands of GPUs (scale-out).
Now, the number of GPUs within a single cluster has passed several hundred,
and high-speed intra-rack interconnect (scale-up) has become the system performance bottleneck.
Broadcom draws a clear distinction between these network tiers:
Network type | Typical use | Connections | Reach | Core technology |
Scale-Up | Intra-rack GPU interconnect | 10–100+ | <10m | Copper / Short-Reach Optics |
Scale-Out | Inter-rack cluster interconnect | 1,000–100,000+ | 10–100m | Ethernet Fabric |
Scale-Across | Data center interconnect | 100K–1M GPUs | km-class distances | Coherent Ethernet / WAN Overlay |
SUE (Scale-Up Ethernet) is designed for the first tier, the rack-scale fabric,
letting GPUs share memory and move tensor data over an Ethernet-based architecture
to achieve a compute model where a cluster acts like one giant GPU.
2. Why Use Ethernet for Scale-Up?
Traditional GPU fabrics (such as NVLink and Infinity Fabric) are proprietary architectures:
hard to integrate across vendors;
closed protocols with a packet layer that is hard to extend;
high physical-layer cost and a lack of standards.
Broadcom points to three major advantages of building on Ethernet:
Scalable: reuses existing Ethernet PHY/MAC, switch silicon and test infrastructure.
Open & Modular: lets each accelerator vendor run its own transport-layer protocol.
Cost-Optimized: leverages the mature copper / AEC / optics ecosystem.
Broadcom has already deployed several SUE prototypes in internal testing,
some of which are integrated into Tomahawk 5 / 6 switches and CPO modules.
3. SUE Architecture: Separating Transport and Network
Broadcom stresses that SUE's key innovation is decoupling the transport layer from the network layer.
● Lower layer: Ethernet Networking (ESAN, Ethernet Scale-Up for Networking)
Defined at the link, MAC and PHY layers, led by Broadcom and Arista.
Core functions include:
Link-Level Retry (LLR)
Thread-Based Flow Control (CBFC)
Optimized Headers & Lightweight Framing
Low-Latency Deterministic Jitter Control
● Upper layer: SUE Transport
Left open for GPU/XPU vendors to implement freely.
Specific features can be chosen per application, for example:
Transaction Packing
Reliability Layer (hop-by-hop or end-to-end)
Memory Ordering Models
Congestion & Load Balancing Policies
Encryption / Security Options
Lightweight Retransmission (Go-Back-N)
This design makes SUE a "menu of choices" protocol,
so each vendor can pick and choose features based on its compute architecture, silicon resources and workloads.
4. High Bandwidth, High Efficiency, Low Power: Three Principles of AI Fabric
Broadcom notes that in a GPU cluster each GPU has 4–8 HBM stacks, for total bandwidth of up to 100 TB/s.
For multiple GPUs to share memory, interconnect bandwidth must reach at least 1/10 of HBM bandwidth.
SUE's design therefore follows three core principles:
Bandwidth Density:
8–12× higher than scale-out Ethernet;
target I/O of 3.2–6.4T per XPU.
Power Efficiency:
must be embedded next to the GPU die, at under 5W per port;
traditional RDMA NICs can't be used, as their area and power are too high.
Reliability & Determinism:
low latency and low jitter are design priorities;
occasional packet errors are handled with lightweight retransmission, without TCP-style overhead.
5. Implementation Examples and Topologies
Broadcom showed two 128-XPU cluster architectures:
Topology | Switch chip | GPU I/O | Port speed | Result |
Option A | Tomahawk 5 | 3.2T | 400G | Single-hop multi-plane architecture |
Option B | Tomahawk 6 | 6.4T | 800G | Double the bandwidth, low latency |
SUE can already support single-tier clusters of a hundred-plus GPUs,
and as low-cost short-reach optics mature,
it is expected to extend to cross-rack multi-rack scale-up fabrics.
6. Open Standardization: The SUE Consortium and Working Groups
Broadcom has launched two main workstreams:
Workstream | Name | Function | Lead |
SUE Transport WG | Scale-Up Ethernet Transport | Defines the XPU transport-layer protocol | Broadcom + Arista |
ESAN WG | Ethernet Scale-Up for Networking | Defines MAC/PHY-layer functions | Led by Broadcom |
EAN WG (upcoming) | Ethernet AI Networking | Integrating scale-out and scale-up | Multiple vendors participating |
Broadcom called on more vendors to join the open collaboration
to ensure Ethernet becomes the universal AI fabric technology,
rather than just a server-to-server communication protocol.
Conclusion
The SUE Framework presented by Broadcom and Arista at OCP 2025
signals that Ethernet is moving from the data center backbone into the heart of the GPU.
"Ethernet is no longer just a network; it will become the unified medium for memory and compute."
Through an open, modular, low-power architecture, SUE lets accelerators interoperate in a standardized way
while preserving flexibility and differentiation. Broadcom expects future AI racks to be built around SUE + UALink,
delivering a truly open, scalable scale-up fabric.
Further Perspectives
Technical impact
SUE is a key milestone in Ethernet's evolution into the compute interconnect layer.
Together with UALink (Layers 1–3) and UEC (Layers 4–5), it forms a complementary, complete AI fabric stack.
Supply chain observations
Through Tomahawk 6 + CPO combined with Arista's software stack, Broadcom forms an integrated hardware-software ecosystem.
In the future, ODMs/OEMs (such as Supermicro and Dell) can build GPU racks directly on the open SUE architecture.
Market trends
AI infrastructure is shifting from proprietary fabrics to open Ethernet.
Between 2026 and 2028, SUE + UALink + CXL may become the mainstream three-tier open architecture.

Comments