top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

NVIDIA's Silicon Photonics Bombshell at ISSCC 2026: 3D Stacking and Low-Latency DWDM Links

2 days ago
2 min read

With AI model parameter counts growing exponentially, GPU-to-GPU communication bandwidth has become the ceiling on performance. At ISSCC 2026, NVIDIA's paper (Paper 23.1) revealed its thinking on next-generation interconnect: stop blindly chasing per-lane speed and embrace a high-density, low-latency 3D silicon photonics architecture.


1. A Major Strategic Pivot: Why 32G NRZ Beats 224G PAM4

The industry is broadly moving toward 224G, but NVIDIA points out that each doubling of data rate costs roughly 4dB of receiver sensitivity, along with a severe bandwidth-limitation penalty.


  • The price of data rate: Conventional high-speed electrical IO at 224G has a raw BER of only about 1E-4 to 1E-6 and must rely on complex FEC (forward error correction) to meet system requirements, which inevitably adds 10ns of latency.

  • DWDM strikes back: NVIDIA uses DWDM (dense wavelength division multiplexing) to pack 8 32G NRZ data channels into a single fiber. This sustains a total throughput of up to 256 Gbps while cutting latency to <1ns, with no need for complex FEC or ML processing.


2. The Core Component: Pushing Microring Resonators to the Limit

NVIDIA chose the microring as the core DWDM building block. Although it is extremely sensitive to process and temperature, the advantages it brings are irreplaceable:

  • High selectivity, small size: With high wavelength selectivity (high Q), a microring can serve simultaneously as modulator, multiplexer and filter, and its tiny footprint suits CMOS drive.


  • Crosstalk and spectrum management: NVIDIA spaces 9 microrings with a radius of just 5µm (8 data + 1 clock) evenly across the 1310nm band, achieving channel spacing of about 200GHz, and precisely models and resolves crosstalk from the spectral long tail.



3. The Secret Sauce: Forwarded Clocking with a Band-Pass Filter (BPF)

This is the most original part of the work. Traditional embedded clock (EC) architectures are limited by CDR bandwidth and are exposed to jitter risk.


  • Jitter tracking and filtering: NVIDIA uses a forwarded clock (FC) architecture with a band-pass filter (BPF) that filters out uncorrelated jitter while tracking correlated jitter.


  • Injection locking: Every channel at the receiver gets the forwarded clock and applies a second stage of noise filtering via an injection-locked oscillator (ILO), which is critical for an energy-efficient DWDM link.


4. Packaging and Integration: The Brute-Force Elegance of 3D Stacking

NVIDIA packages this technology as "Optics on Interposer".


  • Hybrid bonding: A 7nm electronic IC (EIC) and a 65nm silicon photonics IC (PIC) are 3D-stacked and placed directly on the interposer next to the GPU.


  • Performance benchmark: This structure achieves a remarkable areal density of 1.33 Tb/mm^2, greatly improving use of GPU shoreline compared with conventional approaches, which is key to rack-scale AI interconnect.


Conclusion: NVIDIA Is Defining the Standard for AI Interconnect

The ISSCC 2026 data show NVIDIA's measured energy efficiency at just 2.51-2.59 pJ/bit (including circuits, clocking and thermal tuning), with a raw BER below 1E-11. While the industry is still wrestling with 224G signal integrity, NVIDIA has shown how optical 3D stacking and clever clocking can strike a near-perfect balance among latency, power and density.

This is not just a win for optical communications, but another showcase of NVIDIA's vertical integration capabilities.

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page