top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

SEMICON 2025 Silicon Photonics Summit | NVIDIA: The Network Is the Data Center

3 days ago
3 min read

Introduction

As artificial intelligence (AI) keeps scaling, the data center has evolved from a simple collection of servers into the "unit of compute" itself. In its talk at SEMICON 2025, NVIDIA stressed that the network is the core of the data center: whether it's a single server, a GPU cluster or a supercomputer spanning data centers, the key lies in how everything is connected and how data moves.

The talk took a deep look at three network architectures — scale-up (within the system), scale-out (across servers) and scale-across (across data centers) — and at the roles of InfiniBand, Spectrum-X Ethernet, DPUs and the photonic engine (CPO, co-packaged optics) in AI infrastructure.


Key Points

1. The Shift in the Data Center's Unit of Compute

Traditionally, the core of compute was the CPU, which then evolved into GPU servers. In the AI era, however, the entire data center is one computer.

  • Scale-up: interconnecting the GPU ASICs within a rack over an ultra-low-latency network so they operate like a single giant GPU.

  • Scale-out: connecting hundreds of thousands of GPUs across servers and racks to form an AI supercomputer.

  • Scale-across: connecting across data centers and even across cities, so distributed resources form a global computing network.

2. Foundational Technologies for Scale-Up and Scale-Out

  1. InfiniBand

    • Still the industry-recognized standard for low-latency, high-performance interconnect.

    • Well suited to large-scale AI training and inference.

  2. Spectrum-X Ethernet

    • NVIDIA's first Ethernet purpose-built for distributed AI computing.

    • Unlike traditional Ethernet, it offers low latency, zero jitter and high reliability.

    • Paired with smart NICs (SuperNIC) and distributed switches, it can simultaneously deliver:

      • Dynamic traffic control (avoiding congestion)

      • Packet ordering (preserving data consistency)

  3. DPU (Data Processing Unit)

    • Sits at the gateway for the data center's "north-south traffic."

    • Functions:

      • Security isolation: separates infrastructure management from application compute.

      • Storage integration: provides high-performance access and management capabilities.

3. Scale-Across: Interconnecting Data Centers

As AI models keep growing, a single data center is no longer enough, and data centers across cities and countries must be linked into one supercomputer.

  • Challenge: greater distance means more latency. Traditional deep-buffer switches can absorb congestion, but they add extra latency and degrade performance.

  • Solution: NVIDIA's distance-aware algorithms, which use distributed scheduling and real-time monitoring to avoid congestion and maintain performance.

  • Result: on long-distance links between data centers, measured results show nearly a 2x performance improvement.

4. The Role of Optical Engines and CPO

As GPU counts and bandwidth demand grow explosively, traditional electrical signaling is no longer power-efficient enough. NVIDIA's answer:

  • CPO (Co-Packaged Optics): integrating the optical engine directly into the switch package, shortening the electrical signal path and lowering power consumption.

  • Technical highlights:

    • Introduces micro-involved engines to support the bandwidth needs of future generations.

    • Innovative packaging approaches to ensure reliability and yield.

    • Measured results: at the same power, it can support GPU connectivity at a larger scale.

5. Looking Ahead

NVIDIA emphasized that next-generation data centers will feature:

  1. Massive GPU integration: a single workload can run across hundreds of thousands of GPUs.

  2. Optimized network layering: switches handle traffic distribution, while NICs handle ordering and congestion control.

  3. More energy-efficient optical solutions: CPO and advanced packaging will keep improving performance and energy efficiency.

  4. A truly global supercomputer: scale-across technology links data centers in different cities and countries into one.


Summary

In its SEMICON 2025 talk, NVIDIA clearly laid out its complete blueprint for AI data center networking infrastructure. From scale-up (GPU integration within the rack) and scale-out (expansion across servers) to scale-across (integration across data centers), NVIDIA's solutions cover every layer, using InfiniBand, Spectrum-X Ethernet, DPUs and CPO technology to build the next generation of AI supercomputers.

Going forward, as AI models and compute demand keep growing, network performance and energy efficiency will become core competitive advantages. NVIDIA's strategy isn't just to launch individual products but to build a complete distributed computing universe, truly fusing the world's data centers into one "giant AI computer."

Recent Posts

See All
SEMICON 2025 Silicon Photonics Summit: TSMC

At the SEMICON 2025 Silicon Photonics Summit, TSMC's 黃欣芬 presented the company's latest silicon photonics advances and its COUPE (Compact Universal Photonic Engine) platform, from hybrid-bonded EIC/PI

 
 
 

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page