Is the GPU Era Ending? AWS Trainium 4, NVLink Fusion and the Next AI Interconnect Standards
1. Introduction: a new compute war has begun in the AI era
For the past decade, we got used to equating "AI = GPU," as if training large models were NVIDIA's exclusive territory.
From GPT-3 and Stable Diffusion to recent multimodal models and agentic AI, NVIDIA has nearly monopolized the entire compute supply chain.
But starting in 2025, things look different.
AI models keep getting bigger and more multimodal, and inference volume is exploding with AI agents. "Compute" is no longer the only bottleneck — the real bottleneck has become interconnect.
GPU FLOPS are rising fast, but GPU-to-GPU interconnect bandwidth is not keeping pace, so efficiency gains in large-model training and inference are starting to slow.
AWS's Trainium 4 and NVLink Fusion, announced this year, mark a major turning point in the history of AI accelerators:
For the first time, a non-NVIDIA chip is allowed to connect directly to NVIDIA's NVLink fabric.
It signals that AI is no longer a single compute architecture, but an era where heterogeneous accelerators from many vendors coexist.
It also signals that the GPU's absolute dominance is being pried open, bit by bit, by other chipmakers and cloud companies.
From an engineering and data center architecture perspective, this article breaks down:
How performance evolves from Trainium 3 → 4
What NVLink Fusion really is
Where NeuronLink fits
The roles of UALink / PCIe / CXL
Why interconnect matters more than GPU compute
Why this will redefine the optical communications and silicon photonics industries
The far-reaching impact on Taiwan's supply chain
This article is not just about chips; it is about the overall architectural evolution of next-generation AI data centers.
2. GPUs got faster, but interconnect didn't keep up: AI's real bottleneck is "connection"
2.1 Exploding GPU FLOPS ≠ exploding training speed
Take NVIDIA GPUs as an example:
Hopper → Blackwell → Rubin: FLOPS grow 2–4× per generation
HBM bandwidth rises in step
Memory capacity keeps increasing
The compute looks formidable — so why hasn't training time dropped linearly?
Why do bigger models get harder to train?
The answer:
Interconnect *hasn't kept up.
2.2 Large-model training depends on GPUs constantly exchanging data
Large-model training isn't "one GPU" running; it's hundreds or thousands of GPUs exchanging data nonstop:
All-Reduce: gradient synchronization
All-Gather: MoE gating
Pipeline parallel: layer-by-layer transfer
Tensor parallel: matrix partitioning
FSDP / ZeRO: memory offloading
For these operations, 90% of the bottleneck is in the interconnect, not the GPU itself.
2.3 PCIe is growing far more slowly than GPUs
Interface | Bandwidth | Growth pace |
PCIe Gen4 → Gen5 | +2× | Every 3 years |
GPU FLOPS (H100 → B100) | +4× to 8× | Every 2 years |
GPU count (cluster) | Heading toward 512–4096 GPUs | Doubling rapidly |
Model scale (GPT-3 → GPT-4) | 175B → ∼1T | Exponential growth |
If interconnect doesn't keep up, even the strongest GPU is useless.
That is the core reason AWS had to build Trainium and NeuronLink — and ultimately adopt NVLink Fusion.
3. AWS's Trainium strategy: the leap from Trainium 3 → Trainium 4
AWS began building its own AI accelerators in 2021: Inferentia (inference) and Trainium (training).
The 2025 launch of Trainium 3 and Trainium 4 is key to AWS's effort to break its sole dependence on GPUs.
3.1 Trainium 3: AWS's first 3nm AI chip
Trainium 3 is the version currently in commercial service. Highlights include:
Built on a 3nm process
Big jump in per-chip performance (3× the training throughput of the previous generation)
40% better energy efficiency (performance per watt)
Paired with Trn3 UltraServer, integrating up to 144 chips
Suited to MoE, multimodal and long-context models
Trainium 3's goal is not to beat Blackwell, but:
to deliver "good enough" training performance at lower cost, so AWS data centers can fully control their supply chain.
3.2 Trainium 4: the key that truly unlocks next-gen AI architecture
Trainium 4's gains go far beyond "performance" in the traditional sense:
FP4 performance expected to reach 6× Trainium 3
FP8 performance up about 3×
Memory bandwidth up 4× over the previous generation
Major increases in interconnect bandwidth and parallelism
But the most important highlight is:
3.3 ✨ Trainium 4 supports "NVLink Fusion"
Something that hasn't happened in the AI accelerator industry for years:
For the first time, NVIDIA allows non-NVIDIA chips onto NVLink.
Trainium 4 becomes:
not a standalone cluster
not limited to NeuronLink
but able to join GPUs in a "Unified GPU-ASIC Fabric"
This completely changes data center design.
4. The NVLink Fusion revolution: heterogeneous accelerators enter the GPU fabric for the first time
To understand why NVLink Fusion matters, first recall what the original NVLink is.
4.1 NVLink: NVIDIA's historically closed high-speed interconnect
NVLink characteristics:
Ultra-high bandwidth (a single link far exceeds PCIe)
Ultra-low latency
Supports collective operations
Deeply integrated with CUDA / NCCL
Until now, only NVIDIA's own GPUs could use NVLink.
No ASIC could be paired with GPUs to train large models.
4.2 NVLink Fusion: a new spec supporting mixed ASIC + GPU architectures
NVLink Fusion is an open form of NVLink newly introduced in 2025:
It allows:
GPU ←→ GPU
GPU ←→ ASIC (e.g., Trainium 4)
ASIC ←→ ASIC (if the vendor supports it)
This means:
Trainium 4 can participate directly in the GPU fabric, just like a GPU.
Not through PCIe, and not over Ethernet.
But with NVLink-class latency and bandwidth.
4.3 What does this mean? (technical impact)
✔ GPUs + Trainium can jointly run All-Reduce
✔ MoE gating can be distributed cooperatively across both
✔ Model parameters can be shared
✔ No more PCIe bottleneck
✔ A single training job can use both accelerator types at once
For the first time ever, cloud players are starting to open chip interconnects to one another.
4.4 Why would NVIDIA do this?
The reasons are pragmatic:
AI models are so large that GPUs alone can no longer meet hyperscalers' cost needs
Hyperscalers need more "cheaper but good enough" ASICs
If NVLink stayed closed, cloud players would build their own ASIC ecosystems
GPUs would gradually be shut out of the AI fabric
NVIDIA doesn't want an "ASIC-only" future.
So NVLink Fusion is a "mutual-benefit strategy":
NVIDIA preserves the GPU's central position
AWS can turn Trainium into an extension accelerator for GPUs
5. NeuronLink: the changing role of AWS's native interconnect
Before Trainium 4, AWS's interconnect system was mainly:
PCIe (host I/O) + NeuronLink (Trainium interconnect)
NeuronLink's characteristics:
AWS-proprietary
Much faster than PCIe
Used for Trainium ↔ Trainium
Supports distributed training
But it has the same problem: it is closed.
With AWS supporting NVLink Fusion, NeuronLink's role becomes:
✔ "Intra-Trainium clusters" still use NeuronLink
✔ GPU ↔ Trainium interconnect moves to NVLink Fusion
✔ Large-scale fabrics will be dominated by the NVSwitch / NVLink architecture
NeuronLink won't disappear, but it will become:
a "local high-speed interconnect" limited to Trainium clusters.
6. UALink, PCIe, CXL: the battlefield of interconnect standards
In recent years, AI interconnect standards have proliferated:
NVIDIA: NVLink / NVLink Fusion
AWS: NeuronLink
AMD: UALink
Intel: CXL / Xe Link
PCI-SIG: PCIe Gen6/Gen7
Below is a brief explanation of the differences.
6.1 PCIe: universal I/O, but not suited to large-model training
PCIe's role:
CPU ↔ accelerator
Control, DMA, data movement
Drawbacks:
High latency
Insufficient bandwidth
No support for collective ops
Cannot do model parallelism
Conclusion:
PCIe is a necessity for AI, but it is not a high-speed interconnect.
6.2 CXL: for memory pooling, not AI training
CXL's strengths:
Memory pooling
Host-coherent memory
However:
Latency is still high
Not suited to massive parameter synchronization
Not suited to deep learning distributed training
So CXL is positioned to solve memory problems in CPU servers, not GPU/ASIC training.
6.3 UALink: an AMD-led "open AI interconnect"
UALink aims to:
let accelerators from different vendors share one fabric
counter NVIDIA's closed NVLink
build an open ecosystem
But it is still at an early stage, and whether it can challenge NVLink remains unknown.
6.4 Interconnect standards compared
Standard | Led by | Use | Latency | Bandwidth | Openness |
PCIe | Multi-vendor | Host I/O | High | Medium | Open |
CXL | Intel | Memory pooling | High | Medium | Semi-open |
UALink | AMD | AI Fabric | Low-medium | High | Open |
NeuronLink | AWS | Trainium Fabric | Low | High | Closed |
NVLink Fusion | NVIDIA | Heterogeneous Fabric | Lowest | Highest | Semi-open (to hyperscalers) |
7. Why will interconnect change the AI data center?
7.1 The bigger the model, the more interconnect matters over the GPU itself
GPT-4, Claude, Gemini, Sora, OpenAI's video models…
More and more models are moving toward:
Multimodality
Long context (>200k)
Mixture-of-Experts (MoE)
High-frequency inference for agentic AI
Every one of these needs extremely high cross-chip bandwidth.
So the bottleneck of future data centers:
❌ is not the GPU
✔ is the interconnect
✔ and even optical I/O
7.2 Hyperscalers cannot rely entirely on GPUs
Reasons:
Cost is too high
Supply is insufficient
The ecosystem is too closed
Power consumption is too high
Multimodal models need different specialized accelerators
That is why AWS, Google, Meta and Microsoft are all building ASICs.
7.3 NVLink Fusion makes hybrid architectures the new mainstream
Before:
GPUs could only train alongside other GPUs.
Going forward:
GPU (general-purpose) + ASIC (specialized) + hybrid fabric
This will become the new supercomputer standard.
7.4 Optical demand will accelerate from 800G → 1.6T → 3.2T
Because:
Interconnect bandwidth is exploding
GPU clusters are getting bigger
ASIC + GPU topologies are more complex
NVSwitch / UALink / hybrid fabrics need denser optical links
So the impact on the supply chain is huge (more below).
8. Blueprint for next-gen data centers: GPU × ASIC × Optical I/O
Over the next 3–5 years, data center architecture will shift from:
GPU-centric → fabric-centric
This will affect:
Chip architecture
Switch ASIC
Optical modules
Silicon photonics PICs
Advanced packaging
Cooling and power systems
Rack design
The entire supply chain
The true supercomputer of the large-model era is:
a giant fabric woven with optical interconnects — not a pile of GPUs.
9. Impact on Taiwan's supply chain: who benefits most?
Taiwan will become even more critical in next-generation AI data centers.
Here are the main categories of beneficiaries.
9.1 Optical module makers (800G / 1.6T / 3.2T) — the biggest winners
AWS, NVIDIA, Google and Meta will all sharply increase purchases of:
800G SR8 / DR8
1.6T DR8 / FR4
3.2T (future)
With every interconnect upgrade, optical module shipments grow by multiples.
9.2 Silicon photonics — the core technology of the optical interconnect generation
As CPO / optical I/O go mainstream, SiPh becomes the basis for:
Optical engines
High-density optical I/O
PICs (photonic ICs)
Multimodal sensing
Low-power high-speed interconnect
Taiwan's SiPh ecosystem stands to gain enormous opportunities.
9.3 Lasers (EML / DFB / CW) — the scarcest key component
Optical module demand for AI servers surges → lasers are the most supply-constrained
Taiwan's supply capacity is limited today, but it will hold an important strategic position.
9.4 ABF, ceramic substrates, advanced packaging
Large numbers of GPUs, ASICs and switch ASICs all need:
Larger carriers
Higher bandwidth
Lower reflection / loss
Taiwan excels at the PCB / IC package supply chain.
9.5 GPU server ODMs / power / cooling
Taiwanese companies remain:
the world's largest server manufacturing base
the deepest pool of expertise in liquid cooling and AI racks
the most mature in data center power systems
As compute grows 10× in the future, cooling and power will be the next wave of critical technologies.
10. Conclusion: the era of heterogeneous accelerators has officially arrived
The core of an AI supercomputer is no longer just the GPU, but:
Hybrid GPU × ASIC compute
High-speed interconnect standards such as NVLink Fusion
Large-scale optical communications and silicon photonics
Fabric-centric data center architecture
A hyperscaler-led ecosystem of heterogeneous accelerators
The AWS Trainium 4 + NVLink Fusion combination signals:
→ AI compute is no longer the exclusive domain of a single company
→ Interconnect matters more than the GPU itself
→ Data centers will move toward "optical-interconnect-first" architectures
It also means:
The GPU era isn't over, but the era of "only one kind of GPU" is. Heterogeneous accelerators will be the main theme of next-generation AI supercomputers.




Comments