top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

Is the GPU Era Ending? AWS Trainium 4, NVLink Fusion and the Next AI Interconnect Standards

2 days ago
8 min read

1. Introduction: a new compute war has begun in the AI era

For the past decade, we got used to equating "AI = GPU," as if training large models were NVIDIA's exclusive territory.

From GPT-3 and Stable Diffusion to recent multimodal models and agentic AI, NVIDIA has nearly monopolized the entire compute supply chain.

But starting in 2025, things look different.

AI models keep getting bigger and more multimodal, and inference volume is exploding with AI agents. "Compute" is no longer the only bottleneck — the real bottleneck has become interconnect.

GPU FLOPS are rising fast, but GPU-to-GPU interconnect bandwidth is not keeping pace, so efficiency gains in large-model training and inference are starting to slow.

AWS's Trainium 4 and NVLink Fusion, announced this year, mark a major turning point in the history of AI accelerators:

For the first time, a non-NVIDIA chip is allowed to connect directly to NVIDIA's NVLink fabric.

It signals that AI is no longer a single compute architecture, but an era where heterogeneous accelerators from many vendors coexist.

It also signals that the GPU's absolute dominance is being pried open, bit by bit, by other chipmakers and cloud companies.

From an engineering and data center architecture perspective, this article breaks down:

  • How performance evolves from Trainium 3 → 4

  • What NVLink Fusion really is

  • Where NeuronLink fits

  • The roles of UALink / PCIe / CXL

  • Why interconnect matters more than GPU compute

  • Why this will redefine the optical communications and silicon photonics industries

  • The far-reaching impact on Taiwan's supply chain

This article is not just about chips; it is about the overall architectural evolution of next-generation AI data centers.

2. GPUs got faster, but interconnect didn't keep up: AI's real bottleneck is "connection"


2.1 Exploding GPU FLOPS ≠ exploding training speed

Take NVIDIA GPUs as an example:

  • Hopper → Blackwell → Rubin: FLOPS grow 2–4× per generation

  • HBM bandwidth rises in step

  • Memory capacity keeps increasing

The compute looks formidable — so why hasn't training time dropped linearly?

Why do bigger models get harder to train?

The answer:

  • Interconnect *hasn't kept up.


2.2 Large-model training depends on GPUs constantly exchanging data

Large-model training isn't "one GPU" running; it's hundreds or thousands of GPUs exchanging data nonstop:

  • All-Reduce: gradient synchronization

  • All-Gather: MoE gating

  • Pipeline parallel: layer-by-layer transfer

  • Tensor parallel: matrix partitioning

  • FSDP / ZeRO: memory offloading

For these operations, 90% of the bottleneck is in the interconnect, not the GPU itself.


2.3 PCIe is growing far more slowly than GPUs

Interface

Bandwidth

Growth pace

PCIe Gen4 → Gen5

+2×

Every 3 years

GPU FLOPS (H100 → B100)

+4× to 8×

Every 2 years

GPU count (cluster)

Heading toward 512–4096 GPUs

Doubling rapidly

Model scale (GPT-3 → GPT-4)

175B → ∼1T

Exponential growth

If interconnect doesn't keep up, even the strongest GPU is useless.

That is the core reason AWS had to build Trainium and NeuronLink — and ultimately adopt NVLink Fusion.


3. AWS's Trainium strategy: the leap from Trainium 3 → Trainium 4

AWS began building its own AI accelerators in 2021: Inferentia (inference) and Trainium (training).

The 2025 launch of Trainium 3 and Trainium 4 is key to AWS's effort to break its sole dependence on GPUs.


3.1 Trainium 3: AWS's first 3nm AI chip

Trainium 3 is the version currently in commercial service. Highlights include:

  • Built on a 3nm process

  • Big jump in per-chip performance (3× the training throughput of the previous generation)

  • 40% better energy efficiency (performance per watt)

  • Paired with Trn3 UltraServer, integrating up to 144 chips

  • Suited to MoE, multimodal and long-context models

Trainium 3's goal is not to beat Blackwell, but:

to deliver "good enough" training performance at lower cost, so AWS data centers can fully control their supply chain.

3.2 Trainium 4: the key that truly unlocks next-gen AI architecture

Trainium 4's gains go far beyond "performance" in the traditional sense:

  • FP4 performance expected to reach 6× Trainium 3

  • FP8 performance up about 3×

  • Memory bandwidth up 4× over the previous generation

  • Major increases in interconnect bandwidth and parallelism

But the most important highlight is:


3.3 ✨ Trainium 4 supports "NVLink Fusion"

Something that hasn't happened in the AI accelerator industry for years:

For the first time, NVIDIA allows non-NVIDIA chips onto NVLink.

Trainium 4 becomes:

  • not a standalone cluster

  • not limited to NeuronLink

  • but able to join GPUs in a "Unified GPU-ASIC Fabric"

This completely changes data center design.


4. The NVLink Fusion revolution: heterogeneous accelerators enter the GPU fabric for the first time

To understand why NVLink Fusion matters, first recall what the original NVLink is.

4.1 NVLink: NVIDIA's historically closed high-speed interconnect

NVLink characteristics:

  • Ultra-high bandwidth (a single link far exceeds PCIe)

  • Ultra-low latency

  • Supports collective operations

  • Deeply integrated with CUDA / NCCL

Until now, only NVIDIA's own GPUs could use NVLink.

No ASIC could be paired with GPUs to train large models.


4.2 NVLink Fusion: a new spec supporting mixed ASIC + GPU architectures

NVLink Fusion is an open form of NVLink newly introduced in 2025:

It allows:

  • GPU ←→ GPU

  • GPU ←→ ASIC (e.g., Trainium 4)

  • ASIC ←→ ASIC (if the vendor supports it)

This means:

Trainium 4 can participate directly in the GPU fabric, just like a GPU.

Not through PCIe, and not over Ethernet.

But with NVLink-class latency and bandwidth.


4.3 What does this mean? (technical impact)

✔ GPUs + Trainium can jointly run All-Reduce

✔ MoE gating can be distributed cooperatively across both

✔ Model parameters can be shared

✔ No more PCIe bottleneck

✔ A single training job can use both accelerator types at once

For the first time ever, cloud players are starting to open chip interconnects to one another.


4.4 Why would NVIDIA do this?

The reasons are pragmatic:

  • AI models are so large that GPUs alone can no longer meet hyperscalers' cost needs

  • Hyperscalers need more "cheaper but good enough" ASICs

  • If NVLink stayed closed, cloud players would build their own ASIC ecosystems

  • GPUs would gradually be shut out of the AI fabric

NVIDIA doesn't want an "ASIC-only" future.

So NVLink Fusion is a "mutual-benefit strategy":

  • NVIDIA preserves the GPU's central position

  • AWS can turn Trainium into an extension accelerator for GPUs


5. NeuronLink: the changing role of AWS's native interconnect

Before Trainium 4, AWS's interconnect system was mainly:

PCIe (host I/O) + NeuronLink (Trainium interconnect)

NeuronLink's characteristics:

  • AWS-proprietary

  • Much faster than PCIe

  • Used for Trainium ↔ Trainium

  • Supports distributed training

But it has the same problem: it is closed.

With AWS supporting NVLink Fusion, NeuronLink's role becomes:

✔ "Intra-Trainium clusters" still use NeuronLink

✔ GPU ↔ Trainium interconnect moves to NVLink Fusion

✔ Large-scale fabrics will be dominated by the NVSwitch / NVLink architecture

NeuronLink won't disappear, but it will become:

a "local high-speed interconnect" limited to Trainium clusters.

6. UALink, PCIe, CXL: the battlefield of interconnect standards

In recent years, AI interconnect standards have proliferated:

  • NVIDIA: NVLink / NVLink Fusion

  • AWS: NeuronLink

  • AMD: UALink

  • Intel: CXL / Xe Link

  • PCI-SIG: PCIe Gen6/Gen7

Below is a brief explanation of the differences.


6.1 PCIe: universal I/O, but not suited to large-model training

PCIe's role:

  • CPU ↔ accelerator

  • Control, DMA, data movement

Drawbacks:

  • High latency

  • Insufficient bandwidth

  • No support for collective ops

  • Cannot do model parallelism

Conclusion:

PCIe is a necessity for AI, but it is not a high-speed interconnect.

6.2 CXL: for memory pooling, not AI training

CXL's strengths:

  • Memory pooling

  • Host-coherent memory

However:

  • Latency is still high

  • Not suited to massive parameter synchronization

  • Not suited to deep learning distributed training

So CXL is positioned to solve memory problems in CPU servers, not GPU/ASIC training.


6.3 UALink: an AMD-led "open AI interconnect"

UALink aims to:

  • let accelerators from different vendors share one fabric

  • counter NVIDIA's closed NVLink

  • build an open ecosystem

But it is still at an early stage, and whether it can challenge NVLink remains unknown.


6.4 Interconnect standards compared

Standard

Led by

Use

Latency

Bandwidth

Openness

PCIe

Multi-vendor

Host I/O

High

Medium

Open

CXL

Intel

Memory pooling

High

Medium

Semi-open

UALink

AMD

AI Fabric

Low-medium

High

Open

NeuronLink

AWS

Trainium Fabric

Low

High

Closed

NVLink Fusion

NVIDIA

Heterogeneous Fabric

Lowest

Highest

Semi-open (to hyperscalers)

7. Why will interconnect change the AI data center?

7.1 The bigger the model, the more interconnect matters over the GPU itself

GPT-4, Claude, Gemini, Sora, OpenAI's video models…

More and more models are moving toward:

  • Multimodality

  • Long context (>200k)

  • Mixture-of-Experts (MoE)

  • High-frequency inference for agentic AI

Every one of these needs extremely high cross-chip bandwidth.

So the bottleneck of future data centers:

❌ is not the GPU

✔ is the interconnect

✔ and even optical I/O


7.2 Hyperscalers cannot rely entirely on GPUs

Reasons:

  • Cost is too high

  • Supply is insufficient

  • The ecosystem is too closed

  • Power consumption is too high

  • Multimodal models need different specialized accelerators

That is why AWS, Google, Meta and Microsoft are all building ASICs.


7.3 NVLink Fusion makes hybrid architectures the new mainstream

Before:

GPUs could only train alongside other GPUs.

Going forward:

GPU (general-purpose) + ASIC (specialized) + hybrid fabric

This will become the new supercomputer standard.


7.4 Optical demand will accelerate from 800G → 1.6T → 3.2T

Because:

  • Interconnect bandwidth is exploding

  • GPU clusters are getting bigger

  • ASIC + GPU topologies are more complex

  • NVSwitch / UALink / hybrid fabrics need denser optical links

So the impact on the supply chain is huge (more below).


8. Blueprint for next-gen data centers: GPU × ASIC × Optical I/O

Over the next 3–5 years, data center architecture will shift from:

GPU-centric → fabric-centric

This will affect:

  • Chip architecture

  • Switch ASIC

  • Optical modules

  • Silicon photonics PICs

  • Advanced packaging

  • Cooling and power systems

  • Rack design

  • The entire supply chain

The true supercomputer of the large-model era is:

a giant fabric woven with optical interconnects — not a pile of GPUs.

9. Impact on Taiwan's supply chain: who benefits most?

Taiwan will become even more critical in next-generation AI data centers.

Here are the main categories of beneficiaries.

9.1 Optical module makers (800G / 1.6T / 3.2T) — the biggest winners

AWS, NVIDIA, Google and Meta will all sharply increase purchases of:

  • 800G SR8 / DR8

  • 1.6T DR8 / FR4

  • 3.2T (future)

With every interconnect upgrade, optical module shipments grow by multiples.

9.2 Silicon photonics — the core technology of the optical interconnect generation

As CPO / optical I/O go mainstream, SiPh becomes the basis for:

  • Optical engines

  • High-density optical I/O

  • PICs (photonic ICs)

  • Multimodal sensing

  • Low-power high-speed interconnect

Taiwan's SiPh ecosystem stands to gain enormous opportunities.


9.3 Lasers (EML / DFB / CW) — the scarcest key component

Optical module demand for AI servers surges → lasers are the most supply-constrained

Taiwan's supply capacity is limited today, but it will hold an important strategic position.


9.4 ABF, ceramic substrates, advanced packaging

Large numbers of GPUs, ASICs and switch ASICs all need:

  • Larger carriers

  • Higher bandwidth

  • Lower reflection / loss

Taiwan excels at the PCB / IC package supply chain.


9.5 GPU server ODMs / power / cooling

Taiwanese companies remain:

  • the world's largest server manufacturing base

  • the deepest pool of expertise in liquid cooling and AI racks

  • the most mature in data center power systems

As compute grows 10× in the future, cooling and power will be the next wave of critical technologies.


10. Conclusion: the era of heterogeneous accelerators has officially arrived

The core of an AI supercomputer is no longer just the GPU, but:

  • Hybrid GPU × ASIC compute

  • High-speed interconnect standards such as NVLink Fusion

  • Large-scale optical communications and silicon photonics

  • Fabric-centric data center architecture

  • A hyperscaler-led ecosystem of heterogeneous accelerators

The AWS Trainium 4 + NVLink Fusion combination signals:

→ AI compute is no longer the exclusive domain of a single company

→ Interconnect matters more than the GPU itself

→ Data centers will move toward "optical-interconnect-first" architectures

It also means:

The GPU era isn't over, but the era of "only one kind of GPU" is. Heterogeneous accelerators will be the main theme of next-generation AI supercomputers.

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page