top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

VLSI 2026 | Technical Paper Analysis | NVIDIA Squeezes -17.3dBm Sensitivity Out of a Differential TIA: A Counterintuitive Design for Direct-Detection Optical Receivers

2 days ago
10 min read

In AI clusters, a large share of optical-link power is not spent on the receiver itself but on the laser that has to be lit up to feed it. The 32Gb/s optical receiver NVIDIA presented at VLSI 2026 (C20.2) puts to work the photodiode's cathode-side current, which is normally wasted to ground, and combines it into a differential transimpedance amplifier (Differential TIA). The cost is double the TIA power and area; the payoff is roughly a √2 improvement in sensitivity. Final results: PD-referred OMA sensitivity of -17.3dBm (32Gb/s) / -18.9dBm (28Gb/s), energy efficiency of 0.484pJ/b, with a 7nm FinFET EIC 3D-stacked on a 65nm silicon photonics PIC (Cu-Cu hybrid bonding). This is not a showpiece receiver — it is a carefully calculated trade of circuit area for system-level laser power.


1. Paper Background: Why a Single Receiver Is Worth a Close Look

The work comes from NVIDIA (teams in Santa Clara, Ridgefield and Durham), with Georgios Kalogerakis as first author, presented at the 2026 IEEE Symposium on VLSI Technology and Circuits as the second paper (C20.2) in the C20 session on next-generation optical transceivers. In that same C20 session NVIDIA presented two 32Gb/s 3D-stacked receivers: one based on a differential TIA (this paper, C20.2) and one based on Self-Timed Decision Feedback Equalization (Self-Timed DFE) (C20.4). Both share the same goal — squeezing a few more dB of receiver sensitivity out of short-reach optical links whose power is dominated by the laser. This article covers the differential-TIA route.

Why does this matter now? As optical engines move ever closer to the compute chip (near-package, in-package), link budgets get very tight. Every 1dB of receiver sensitivity improvement can save a chunk of front-end laser power, or relax the optical link loss budget. The value of a receiver is shifting from "how good its local pJ/bit looks" to "whether it can cut the laser overhead of the entire link."

Figure 1: SNR advantage of a differential TIA (right) over a single-ended TIA (left) — using the AC current from both PD terminals doubles the signal while noise grows only by √2 | Source: A 32Gb/s Optical Receiver utilizing a Differential TIA… (VLSI 2026, C20.2) - Figure 1
Figure 1: SNR advantage of a differential TIA (right) over a single-ended TIA (left) — using the AC current from both PD terminals doubles the signal while noise grows only by √2 | Source: A 32Gb/s Optical Receiver utilizing a Differential TIA… (VLSI 2026, C20.2) - Figure 1

2. The Core Problem: Why Direct Detection "Wastes" Half the Current

First, the problem this paper tackles, in one sentence: direct-detection optical receivers conventionally use only one terminal of the PD. In coherent detection systems, balanced PDs are inherently differential, so pairing them with a differential TIA is natural. But in the direct-detection links used massively in data centers, a single PD is typical, and the TIA attached to it is essentially single-ended — it takes only the photocurrent from the PD anode, while the AC current at the cathode is swallowed by the bias circuit.

The paper asks the opposite question: what if the cathode-side AC photocurrent were used too? The theory is straightforward — the signal at the differential output roughly doubles, while the two TIAs contribute uncorrelated random thermal noise that grows only by √2. SNR therefore improves by √2, and sensitivity improves by the same factor. There is a key caveat: the √2 benefit applies only to TIA thermal noise, not to PD shot noise or laser relative intensity noise. So why is it still worth doing? Because noise in data center optical receivers is dominated by TIA thermal noise to begin with. In this application, the limitation simply doesn't hurt.


3. Why Nobody Did This Before: The Cathode Bias Bottleneck

If it's so good, why has direct detection always been single-ended? Because running the cathode at high speed requires a high reverse bias, far above the voltage a TIA input can accept. Past attempts each had a fatal flaw: first, doing single-ended-to-differential conversion directly at the anode node — which yields worse SNR than single-ended; second, AC-coupling the cathode node directly — which solves the bias incompatibility but severely limits the receiver's low-frequency cutoff (LFC); third, enlarging the AC-coupling capacitor to push the LFC down — but the large capacitor's parasitic capacitance eats back the differential TIA's benefit; fourth, using a stacked differential TIA — workable, but it needs two regulators plus a deep n-well on the cathode side, raising both cost and complexity. In other words, the challenge was never "whether to use the cathode" but "how to use the cathode without sacrificing low-frequency response or adding regulators."


4. The Core Solution: A 650fF Wideband Level Shifter

NVIDIA's answer comes down to two design choices. First, make the PD cathode bias circuit present a higher impedance than the cathode-side TIA input across all frequencies of interest, so the AC photocurrent flows into the TIA instead of being diverted by the bias circuit. Second, place a wideband level shifter in front of the cathode-side TIA to bring the bias voltage down to a level the TIA can accept. The level shifter is implemented as a parallel R_LS, C_LS network. As long as R_LS is lower than the output impedance of the cathode bias device, the circuit works all the way down to very low frequencies without needing a large C_LS. The paper uses a C_LS of just 650fF — small enough to be built as a MOM capacitor without touching the lower metal layers, keeping parasitic capacitance below 0.5% of C_LS. This step is the soul of the paper: it breaks the deadlock of "large AC-coupling capacitor → parasitics destroy the differential benefit."

Figure 2: The proposed differential TIA circuit — an input common-mode loop sets the PD cathode voltage, and the cathode side passes through a wideband level shifter into the second TIA | Source: VLSI 2026, C20.2 - Figure 2
Figure 2: The proposed differential TIA circuit — an input common-mode loop sets the PD cathode voltage, and the cathode side passes through a wideband level shifter into the second TIA | Source: VLSI 2026, C20.2 - Figure 2

The two TIAs are nearly identical, with one deliberate asymmetry: the feedback resistor R_f1 of the cathode-side TIA is 5% higher, compensating for the high-frequency loss caused by the extra capacitance at the cathode node compared with the anode. The cathode-side DC cancellation loop sets the required level-shift current I_LS so that V_cathode − I_LS·R_LS = 0.5VDD. With V_cathode at 1.65V, R_LS at 16kΩ and VDD at 0.85V, the target current is about 76μA. These numbers show the TIA was designed under the discipline of "single supply, minimal extra current."


5. Impedance Balancing: Anode, Cathode and the Bias PMOS


Figure 3: Input impedance of the anode-side / cathode-side TIAs versus the impedance of the PD cathode bias PMOS — the bias circuit impedance must exceed the TIA input impedance | Source: VLSI 2026, C20.2 - Figure 3
Figure 3: Input impedance of the anode-side / cathode-side TIAs versus the impedance of the PD cathode bias PMOS — the bias circuit impedance must exceed the TIA input impedance | Source: VLSI 2026, C20.2 - Figure 3

This figure shows the physical condition that makes the design work: the impedance presented by the cathode bias PMOS must exceed the cathode-side TIA input impedance across the whole band, so that AC current preferentially enters the TIA. It is the quantified version of the earlier statement that "the bias circuit must be high-impedance enough," and it is why this differential TIA works without a large AC-coupling capacitor.

Figure 4: Simulated impedance and transimpedance frequency response — |ZT| curves for the individual TIAs, the differential TIA and the full analog front end (AFE), showing a 2.5MHz low-frequency cutoff | Source: VLSI 2026, C20.2 - Figure 4
Figure 4: Simulated impedance and transimpedance frequency response — |ZT| curves for the individual TIAs, the differential TIA and the full analog front end (AFE), showing a 2.5MHz low-frequency cutoff | Source: VLSI 2026, C20.2 - Figure 4

This figure shows two things: the top half proves that the cathode bias circuit impedance does exceed the TIA input impedance (the design premise holds); the bottom half plots the transimpedance |ZT| versus frequency for the individual TIAs, the combined differential TIA and the full analog front end. The 2.5MHz LFC mentioned in the paper is read from here — and the paper explicitly notes that 2.5MHz is an area-constrained trade-off; nothing in the architecture prevents it from reaching a few hundred kHz. That is an honest hint: there is still room to push the LFC lower.


6. The Full Circuit: Offset Must Be Cleared Before Sampling


Full receiver circuit — multiple current DACs cancel the offset between the two TIA paths before data is sampled by the deserializer (DES) | Source: VLSI 2026, C20.2 - Figure 5
Full receiver circuit — multiple current DACs cancel the offset between the two TIA paths before data is sampled by the deserializer (DES) | Source: VLSI 2026, C20.2 - Figure 5

This figure shows the complete chain from PD to bit decision. Two points stand out: first, each CMOS TIA output also has its own wideband level shifter that lowers the common-mode voltage before feeding the following differential gain stage; second, the many current DACs in the figure exist to clean up mismatch offset between the two TIA paths before the data is sampled by the deserializer (DES). Reaping the benefits of a differential architecture requires the two paths to be symmetric enough — these DACs are what restore that symmetry.

7. The Measurement Setup: Two Lithium Niobate Modulators Plus Microring Filtering

Figure 6: Sensitivity measurement setup — two external lithium niobate (LiNbO₃) modulators generate the optical data and half-rate clock, coupled into the PIC via vertical grating couplers | Source: VLSI 2026, C20.2 - Figure 6
Figure 6: Sensitivity measurement setup — two external lithium niobate (LiNbO₃) modulators generate the optical data and half-rate clock, coupled into the PIC via vertical grating couplers | Source: VLSI 2026, C20.2 - Figure 6

This figure shows how the chip is fed data and clock. Two external LiNbO₃ modulators generate the optical data signal (13dB extinction ratio) and a half-rate clock signal (8.7dB extinction ratio), coupled into the PIC through vertical grating couplers. The data signal first passes through a microring resonator (MRR) filter before reaching the high-speed PD, while the clock goes straight into its PD. Worth noting: the MRR in the data path means this receiver is designed for microring-based DWDM links, not as an isolated single-channel demo.


8. Sensitivity Results: Where -17.3dBm Comes From, and Why It Trails 28Gb/s

Figure 7: Measured sensitivity — -18.9dBm at 28Gb/s and -17.3dBm OMA at 32Gb/s (PD-referred, BER=10⁻¹²) | Source: VLSI 2026, C20.2 - Figure 7
Figure 7: Measured sensitivity — -18.9dBm at 28Gb/s and -17.3dBm OMA at 32Gb/s (PD-referred, BER=10⁻¹²) | Source: VLSI 2026, C20.2 - Figure 7

This figure shows that at a BER target of 10⁻¹², the receiver's PD-referred OMA sensitivity is -18.9dBm (28Gb/s) and -17.3dBm (32Gb/s). 32Gb/s is about 1.6dB worse than 28Gb/s, which the paper attributes to a larger vertical eye-closure penalty — 0.87dB at 28Gb/s, jumping to 1.85dB at 32Gb/s. This is not worse noise; it is the eye being squeezed vertically under limited bandwidth. It indicates that the receiver's real ceiling is front-end bandwidth, not the differential architecture itself.

Bathtub curves at 28/32Gb/s measured 1dB above sensitivity — both show roughly 0.28UI opening; above are simulated differential eye diagrams at the analog front-end output | Source: VLSI 2026, C20.2 - Figure 8
Bathtub curves at 28/32Gb/s measured 1dB above sensitivity — both show roughly 0.28UI opening; above are simulated differential eye diagrams at the analog front-end output | Source: VLSI 2026, C20.2 - Figure 8

This figure shows that at an optical power 1dB above their respective sensitivities, both 28Gb/s and 32Gb/s retain about 0.28UI of horizontal opening. The differential eye diagrams above it are the visual version of the earlier "vertical eye closure" numbers — you can see directly that the 32Gb/s eye loses more in the vertical direction. A 0.28UI opening at BER=10⁻¹² has margin, meaning this is not a result that barely scrapes by.


9. How Much Does Differential Beat Single-Ended: 0.9dB and 0.5dB

Differential TIA vs. single-ended configuration sensitivity at 32Gb/s — differential leads by 0.9dB / 0.5dB for PRBS7 / PRBS15 | Source: VLSI 2026, C20.2 - Figure 9
Differential TIA vs. single-ended configuration sensitivity at 32Gb/s — differential leads by 0.9dB / 0.5dB for PRBS7 / PRBS15 | Source: VLSI 2026, C20.2 - Figure 9

This figure shows a key controlled experiment: the receiver can be configured in single-ended mode (pulling the cathode to the highest supply and turning off the cathode-side TIA), so differential vs. single-ended can be compared fairly on the same silicon. Differential mode leads by 0.9dB with PRBS7 and 0.5dB with PRBS15. In theory √2 corresponds to about 1.5dB, so the measured 0.5–0.9dB is discounted — by extra capacitance in the cathode path, residual mismatch between the two paths and non-idealities in the level shifter. But this "compare against itself" design actually makes the numbers more credible.


10. Power, Area and Peer Comparison: The Weight of 0.484pJ/b

Power breakdown and EIC die photo — at 32Gb/s it draws 14.45mA from a 1.05V supply (regulated on-chip to 0.81V), plus 0.14mA from the 2V PD supply; core area 3700μm² | Source: VLSI 2026, C20.2 - Figure 10
Power breakdown and EIC die photo — at 32Gb/s it draws 14.45mA from a 1.05V supply (regulated on-chip to 0.81V), plus 0.14mA from the 2V PD supply; core area 3700μm² | Source: VLSI 2026, C20.2 - Figure 10

This figure shows the receiver's actual overhead: at 32Gb/s it draws 14.45mA from a 1.05V supply (regulated on-chip to 0.81V) plus 0.14mA from the PD's 2V supply, for an energy efficiency of 0.484pJ/b and a core area of 3700μm². Yes, the TIA doubled, but overall efficiency stays below 0.5pJ/b — the evidence that trading area for sensitivity for system laser power pays off.

Table I: Comparison with existing differential / pseudo-differential receivers — this work, in 7nm FinFET, achieves the best energy efficiency (0.484pJ/b) and leading sensitivity (-17.3dBm @ 32Gb/s) | Source: VLSI 2026, C20.2 - Table I
Table I: Comparison with existing differential / pseudo-differential receivers — this work, in 7nm FinFET, achieves the best energy efficiency (0.484pJ/b) and leading sensitivity (-17.3dBm @ 32Gb/s) | Source: VLSI 2026, C20.2 - Table I

This table shows where the work sits among its peers. The competitors all use older or larger process nodes — 16nm FinFET, 130nm SiGe BiCMOS, 55nm SiGe BiCMOS, 180nm CMOS — while this work is the only 7nm FinFET entry. On energy efficiency, the others range from 1.02 to 8.91pJ/b versus 0.484pJ/b here; sensitivity of -17.3 / -18.9dBm is also near the front of the pack. The message is clear: stacking an advanced CMOS node + 3D stacking + a differential architecture lets direct-detection receivers take a big step forward in both efficiency and sensitivity at once.


11. Technical Highlights: The One or Two Things That Are Truly New

If you remember only two things: first, a wideband level shifter as small as 650fF makes "differential direct detection" work on a single supply, with no extra regulators and no loss of low-frequency response. That is the move that sidesteps every flaw of earlier approaches. Second, it formally treats receiver sensitivity as a lever on system-level laser power rather than an isolated RX metric. The paper's argument runs throughout: 1dB better sensitivity lets you save a chunk of front-end laser power. In that framework, doubling TIA power is not a drawback but a down payment for cutting a much larger slice of the laser budget.


12. Industry Implications: How Far From Production, and Who Benefits

The chip uses a platform that is already running: a 7nm FinFET EIC 3D-stacked on a 65nm silicon photonics PIC via Cu-Cu hybrid bonding — the shared foundation of NVIDIA's near-package optical interconnect (the other receiver in the same session, the C20.4 self-timed DFE design, uses the same 7nm EIC + 65nm PIC platform). Who benefits? The supply chain for hybrid bonding, 3D stacking and silicon photonics wafers — because this route turns packaging from a "back-end step" into a design boundary that determines bandwidth, noise and sensitivity. How far from production? Given that the platform is in production, the architecture is validated and the numbers are competitive, this looks more like an engineering-optimization stage than a proof-of-concept stage. If anything needs work, it is the area-constrained 2.5MHz LFC.


Conclusion

The single most important takeaway from NVIDIA's differential-TIA receiver: it proves that in direct-detection optical links, recovering the "wasted other half" of the PD is a trade whose math adds up. The theoretical √2 gain is discounted to 0.5–0.9dB in measurement, but in short-reach links dominated by laser power, those few dB translate directly into laser overhead saved at the system level — while overall efficiency stays at 0.484pJ/b. In the broader arc of technology history: as optical interconnect moves from pluggable modules to near-package, system-level optical I/O, receiver sensitivity is pushed back to the center of system power, and advanced CMOS nodes plus 3D stacking make differential direct detection — once unaffordable — viable on energy efficiency for the first time.

References

[1] G. Kalogerakis, S. Song, L. Xu, N. Mehta, A. Rekhi, N. Nedovic, B. G. Lee, B. Zimmer, S. Tell, Y. Nishi, X. Chen, W. Lopes, T. Greer, C. T. Gray, "A 32Gb/s Optical Receiver utilizing a Differential TIA with -17.3dBm Sensitivity in a 3D-stacked Silicon Photonics Platform," 2026 IEEE Symposium on VLSI Technology and Circuits (VLSI), C20.2, 2026.

[2] Other related work in the same C20 session: Intel, "An 800 Gbps/Fiber Silicon Photonic Microring-Based DWDM Transceiver in an Open-Cavity Package" (C20.1); University of Washington, "A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning" (C20.3); NVIDIA, "A 0.416-pJ/b, 32-Gb/s 3D-Stacked Optical NRZ Receiver with -18.5-dBm OMA Sensitivity Using Self-Timed Decision Feedback Equalization" (C20.4).


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page