OFC 2026 - 256 Gb/s DWDM Optical I/O in 3D-Stacked EIC/PIC Silicon Photonics Platform - NVIDIA
Introduction: AI Factories and Optics Converge
As 2026 AI factories scale toward 512K-GPU clusters and 40 MW power envelopes, traditional electrical interconnect and pluggable optical modules are hitting their physical limits. Interconnect already accounts for 7–8% of total system power, a huge waste for compute-hungry AI infrastructure.
At OFC 2026, NVIDIA unveiled its latest research: 256 Gb/s DWDM optical I/O. By tightly integrating the electronic IC (EIC) and photonic IC (PIC) through 3D-stacked packaging, it targets the shoreline density bottleneck and brings power down to 2.6 pJ/bit. This marks optical interconnect's move from pluggable slots into a new in-package era.
Core Technology Deep Dive: 3D Packaging and Clock-Forwarded Architecture
1. The 3D-Stacked COUPE Platform: Heterogeneous Integration of 7nm and 65nm
NVIDIA's packaging is based on TSMC's COUPE (Compact Universal Photonic Engine) platform.
Heterogeneous node mix: the EIC uses a 7nm FinFET CMOS process for high-performance logic, while the PIC uses a 65nm SOI silicon photonics process.
Cu-Cu Hybrid Bonding (SoIC): instead of conventional micro-bumps, the dies are bonded face-to-face using copper-to-copper hybrid bonding (SoIC). This sharply reduces parasitic capacitance (C_p + C_ESD), directly improving receiver (RX) sensitivity and cutting laser power by 30% to 40%.
2. Clock Forwarding with a Bandpass Filter (BPF)
Conventional clock and data recovery (CDR) circuits consume a lot of power at high speed. NVIDIA chose a clock-forwarded architecture to simplify the circuitry and shrink die area.
Challenge: conventional clock forwarding doubles up the uncorrelated jitter of the data and clock channels, degrading performance.
Solution: a bandpass filter (BPF) on the clock channel. It preserves the clock signal while filtering out-of-band uncorrelated noise. Experiments show stable transmission at BER < 10⁻¹² at 32 Gb/s.
3. Key Metrics (Quantitative Benchmark)
According to NVIDIA's measured data, the system shows a decisive lead across every dimension:
Parameter | Measured value | Notes |
Per-lane rate | 32 Gb/s | 8 data + 1 clock lanes |
Total throughput (per fiber) | 256 Gb/s | Using DWDM |
Energy efficiency (total) | 2.6 pJ/b | Includes laser, circuit and heater power |
Shoreline bandwidth density | 6x improvement | Vs. embedded-clock architecture |
Areal bandwidth density | 1.33 Tb/s/mm² | Up to 20x density improvement |
Power breakdown (2.59 pJ/b total):
TX circuits and thermal control: 0.67 pJ/b
RX circuits and clock network: 0.59 pJ/b
Micro-ring heaters: 0.57 pJ/b
Laser source (DFB module): 0.76 pJ/b (10.2% efficiency)
Supply Chain and Market Impact: Silicon Photonics' "Golden Decade"
What NVIDIA showed isn't just a lab prototype but a mature DWDM micro-ring resonator (MRR) solution. It has three far-reaching implications for the supply chain:
A shift of power in the packaging supply chain: as COUPE and SoIC become central to optical communications, TSMC's role will expand from wafer foundry to the gateway for integrated optoelectronic packaging, creating a technical barrier for traditional optical module assemblers.
Laser sources become separate and standardized: in NVIDIA's design, the laser source is separate from the SPE (Silicon Photonic Engine). That helps with laser thermal management and replaceability, and gives laser chip makers such as Lumentum and Coherent a new external laser source (ELS) market standard.
A precursor to the zero-DSP era: through clock forwarding and analog front-end (AFE) optimization, NVIDIA achieved very low power at 32 Gb/s. That differs from today's push for maximum speed at 100G/224G, but for short-reach, high-density interconnect inside GPU clusters, a low-latency, low-power analog-only path may hold the stronger commercial edge.
Simple Tech Trend View: Interconnect Is Compute
NVIDIA has revealed a clear strategy: scaling compute no longer depends only on the transistor count inside the GPU, but on how cheaply and densely photons can carry chip-to-chip communication.
2.6 pJ/bit is a formidable number. For comparison, today's electrical interconnect over the same distance runs around 1–2 pJ/bit, but optics offers nearly unlimited bandwidth scaling and several times the reach. NVIDIA's choice to validate the prototype at 32 Gb/s rather than chase 224G PAM4 reflects an engineering philosophy that prioritizes energy-efficient density (performance per watt per mm²) over peak per-lane speed.





















Comments