top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

VLSI 2026 | Paper Analysis | NVIDIA Uses a "Self-Timed DFE" to Push Optical Receiver Sensitivity to -18.5 dBm at 0.416 pJ/b

2 days ago
8 min read

At VLSI 2026 (paper C20.4), NVIDIA presented a 32 Gb/s optical NRZ receiver that 3D-stacks a 7nm CMOS electronic IC (EIC) and a 65nm silicon photonics IC (PIC) using hybrid bonding. With a new technique called self-timed decision feedback equalization (STDFE), it reaches -18.5 dBm OMA sensitivity at BER<10⁻¹² with PRBS31, at an energy efficiency of just 0.416 pJ/b and an area of 3360 μm². The point isn't speed — it's sensitivity: every 1 dB of sensitivity saved at the receiver translates directly into laser power saved at the system level. The most elegant part is that it costs almost nothing: the entire STDFE loop adds less than 1 mA.

1. Background: the "receiver piece" of NVIDIA's optical I/O push

When people talk about co-packaged optics (CPO), attention mostly goes to the "emit and transmit" parts — laser sources, waveguides, micro-ring modulators — while the receiver is treated as a supporting player. Yet in an optical link's power budget, laser power is usually the largest item, and what determines how bright the laser must be is precisely how small a signal the receiver can hear. All authors of this paper are from NVIDIA (teams in Santa Clara, Durham and Ridgefield); it was presented at the 2026 IEEE Symposium on VLSI Technology and Circuits as paper C20.4. C20.2 in the same session took the differential transimpedance amplifier (differential TIA) route; C20.4 deliberately chose the seemingly contrarian but actually clever path of a "low-bandwidth front end plus equalization."

2. The core problem: sensitivity is the lever on laser power

The paper's motivation in one sentence: receiver OMA sensitivity = the lever on system laser power. OMA (optical modulation amplitude) sensitivity is the minimum optical signal swing required at the receiver input for a target BER. Given the link loss and extinction ratio (ER) from the external laser source (ELS) to the receiver, the better the receiver sensitivity, the lower the required laser power. Improving sensitivity is therefore not "a circuit team's internal matter" — it directly cuts the most expensive power item at the system level. The paper distills receiver design goals into four points: small input parasitic capacitance CIN, large transimpedance resistance R1, large sampler input swing, and matched delay between data and clock paths.

3. From pain point to concept: how "delay matching" unexpectedly produced a free equalizer

Figure 1 shows the causal chain of "how sensitivity determines laser power" and the receiver's various noise sources. The top half draws the link model: the external laser reaches the receiver through link loss; the minimum usable optical power min(PIN) plus link loss gives the minimum laser power min(PTX), which divided by wall-plug efficiency (WPE) yields the electrical power actually paid. The key insight: design goals can be condensed into "smaller CIN, larger R1, larger data swing vd." This figure translates circuit metrics into the system laser bill and is the starting point for understanding why the paper is worth doing.

Figure 1: Receiver OMA sensitivity determines the minimum laser power needed to close the link (top); design considerations for improving sensitivity (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 1
Figure 1: Receiver OMA sensitivity determines the minimum laser power needed to close the link (top); design considerations for improving sensitivity (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 1

Figure 2 shows how limiting amplification suppresses noise at the center of the eye, along with the baseline TIA circuit for comparison. The top shows eye diagrams at nodes v3 and vfe under transient noise simulation: v3 swings less than 100 mV with noise σ of about 6 mVrms, while vfe, amplified to rail-to-rail, swings more than 800 mV with noise σ at the eye center below 3 mVrms. This "rail-to-rail output suppresses center noise" effect is essentially equivalent to the "noise-free decision" concept in a DFE. It reveals an overlooked free resource: the data must be amplified to rail-to-rail anyway for delay matching, and that rail-to-rail signal happens to be a clean decision source for a DFE.

Figure 2: Eye diagrams of v3 and vfe under transient noise simulation, showing noise suppression from limiting amplification; vfe (rail-to-rail TIA output) can be treated as a noise-free decision (top); simplified baseline TIA circuit (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 2
Figure 2: Eye diagrams of v3 and vfe under transient noise simulation, showing noise suppression from limiting amplification; vfe (rail-to-rail TIA output) can be treated as a noise-free decision (top); simplified baseline TIA circuit (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 2

Figure 3 shows the core STDFE concept and how it opens the eye. The top compares v3 eye diagrams in three cases: baseline (R1=0.9k, STDFE off) as the 1× reference; increasing R1 to 1.8k with STDFE still off raises eye opening to 1.18× (low-frequency gain rises but bandwidth drops, adding ISI); with R1=1.8k and STDFE on, eye opening jumps to 1.55×, more than 1.5 times the baseline. The approach drives an inverter with the rail-to-rail TIA output, forming current summation at node v2, which combines naturally with the inverter-based Cherry-Hooper amplifier on the main path. Because the DFE timing is "derived from the data itself," it is called a self-timed DFE.

Figure 3: v3 eye diagrams with and without STDFE (top); the self-timed decision feedback equalization (STDFE) concept, using the rail-to-rail TIA output as the decision and current summation via an inverter-based Cherry-Hooper amplifier (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 3
Figure 3: v3 eye diagrams with and without STDFE (top); the self-timed decision feedback equalization (STDFE) concept, using the rail-to-rail TIA output as the decision and current summation via an inverter-based Cherry-Hooper amplifier (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 3

4. Circuit implementation: TIA and clock-forwarding architecture

Figure 4 shows the TIA circuit as actually implemented, including the high-R1 low-bandwidth front end and STDFE. Conceptually STDFE needs only one inverter from vfe to v2; but to optimize the STDFE loop independently without touching the main signal path, the designers instead tap the signal from node v4, amplify it, and drive a current-injection inverter at v2. The STDFE loop delay is close to 1 UI, and a process- and temperature-adaptive supply is used to reduce delay variation. The circuit uses R1=1.8k, plus a DC loop built from an ultra-low-gm OTA and a 5.3 pF thick-oxide MOS capacitor to achieve a very low high-pass cutoff frequency — which removes the overhead of DC-balanced encoding.

Figure 4: The proposed TIA circuit, with a low-bandwidth front end (high R1) and STDFE | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 4
Figure 4: The proposed TIA circuit, with a low-bandwidth front end (high R1) and STDFE | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 4

Figure 5 shows the simplified receiver circuit and clock-forwarding architecture. A half-rate clock is received by a clock receiver, which then drives clock distribution across multiple lanes. The same receiver circuit can be configured as either a data receiver or a clock receiver: the rail-to-rail TIA output enters a delay-matching block (DLY), with one branch for data and one for clock; in data mode the data path is enabled, and a feedback inverter is added to improve signal quality into the deserializer (DES). Clock distribution uses an injection-locked oscillator (ILO). This architecture is designed for many lanes in parallel.

Figure 5: Simplified receiver circuit and clock-forwarding architecture | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 5
Figure 5: Simplified receiver circuit and clock-forwarding architecture | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 5

5. Measurement results: sensitivity, BER and timing margin


Figure 6: Measurement setup (top) and optical eye diagrams of data and clock before the PIC at 32 Gb/s (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 6
Figure 6: Measurement setup (top) and optical eye diagrams of data and clock before the PIC at 32 Gb/s (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 6

Figure 6 shows the measurement setup: two external Mach-Zehnder modulators (MZMs) modulate lasers with a clock and a PRBS pattern respectively, at extinction ratios of 9.4 dB and 12.7 dB; a 50:50 splitter allows the optical eye to be observed before the PIC. The clock signal enters the PD directly via waveguide, while the data signal passes through a micro-ring resonator (MRR) before reaching the PD; the laser wavelengths measured were 1308.23 nm and 1300.00 nm. The inclusion of the MRR echoes NVIDIA's positioning on DWDM optical paths.

Figure 7: Measured receiver OMA sensitivity at 32 Gb/s with PRBS7 (clock receiver input OMA of -18 dBm) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 7
Figure 7: Measured receiver OMA sensitivity at 32 Gb/s with PRBS7 (clock receiver input OMA of -18 dBm) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 7

Figure 7 shows the sensitivity improvement of STDFE over the baseline with a PRBS7 pattern. At BER<10⁻¹², STDFE achieves a 2 dB sensitivity improvement over the baseline receiver, at the cost of less than 1 mA of current overhead (clock receiver input OMA of -18 dBm). This is the most persuasive figure in the paper: 2 dB of sensitivity for under 1 mA — at the system level, those 2 dB translate almost directly into laser power that can be cut.

Figure 8: Measured receiver OMA sensitivity at 32 Gb/s with PRBS7 and PRBS31 (clock receiver input OMA of -18 dBm) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 8
Figure 8: Measured receiver OMA sensitivity at 32 Gb/s with PRBS7 and PRBS31 (clock receiver input OMA of -18 dBm) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 8

Figure 8 shows BER curves under different PRBS patterns and gives the receiver's final sensitivity figures. At BER<10⁻¹², PRBS31 reaches -18.5 dBm OMA sensitivity and PRBS7 reaches -18.7 dBm. PRBS31 is the more demanding long pattern, and sensitivity drops only 0.2 dB, showing the equalization remains robust under long-range data correlation. A PRBS31-level figure is the threshold real link specs require — meeting it is what qualifies a design for product.

Figure 9: Bathtub BER measurement at 32 Gb/s with PRBS31 and 1 dB link margin (data receiver input -17.5 dBm OMA, clock receiver input -18 dBm OMA) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 9
Figure 9: Bathtub BER measurement at 32 Gb/s with PRBS31 and 1 dB link margin (data receiver input -17.5 dBm OMA, clock receiver input -18 dBm OMA) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 9

Figure 9 shows the bathtub curve with 1 dB of link margin, quantifying the timing eye opening. With 1 dB of link margin reserved (data receiver input -17.5 dBm OMA, clock receiver input -18 dBm OMA) and a PRBS31 pattern, the bathtub BER curve shows 27% UI of horizontal eye opening at BER<10⁻¹². This means that even with an extra 1 dB of system margin, nearly 30% of the UI remains as timing window, rather than all margin being used up.

6. 3D stacking and peer comparison: Figure 10 and Table I

Figure 10: Die photo of the EIC and the proposed receiver (top); cross-section of the 3D-stacked EIC and PIC (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 10
Figure 10: Die photo of the EIC and the proposed receiver (top); cross-section of the 3D-stacked EIC and PIC (bottom) | Source: A 0.416-pJ/b 32-Gb/s 3D-Stacked Optical NRZ Receiver… (VLSI 2026, C20.4) - Figure 10

Figure 10 shows die photos of the EIC and the receiver, and a cross-section of the EIC/PIC 3D stack. The EIC is 7nm FinFET and the PIC is 65nm SOI; the key benefit of hybrid bonding is minimizing parasitic capacitance at the receiver input — exactly the "smaller CIN" from Figure 1. The receiver occupies 3360 μm² and draws 13.2 mA from a 1.01 V supply at 32 Gb/s (regulated down to 0.84 V). 3D stacking is the physical prerequisite for shrinking CIN, which allows a larger R1, which in turn improves sensitivity.

As for the paper's Table I, it places this receiver among its peers: compared with JSSC'18 (-12.4 dBm, 1.41 pJ/b), ISSCC'23 (-11.4 dBm, 0.96 pJ/b) and SSCL'24 (-17.0 dBm, 0.0848 pJ/b), this work has the best sensitivity of the four at -18.5 dBm, and is the only one achieved under the most demanding PRBS31 pattern; its energy efficiency of 0.416 pJ/b (including TIA, DLY and DES) sits in the middle, with CIN of 45 fF, responsivity of 1.0 A/W, and no inductors. The industry signal is clear: NVIDIA chose to bet on "sensitivity" rather than on "the energy-efficiency number alone," because sensitivity is the variable that actually brings the system laser bill down.

Conclusion

The value of this paper isn't whether it is the world's fastest receiver or has the prettiest efficiency number; it's that it demonstrates a "free-found equalizer" design philosophy: since the signal must be amplified to rail-to-rail anyway for clock/data delay matching, reuse that rail-to-rail output as a noise-free DFE decision and build the equalization with a few inverters and current summation, for a total cost under 1 mA. The result is -18.5 dBm PRBS31 sensitivity, 0.416 pJ/b efficiency and a 3360 μm² footprint, all built on hybrid-bonded 3D stacking of a 7nm EIC and a 65nm PIC. STDFE turns "sensitivity" from an internal circuit-team metric into the single most worthwhile investment on the whole optical link's balance sheet.

References

Paper title: A 0.416-pJ/b, 32-Gb/s 3D-Stacked Optical NRZ Receiver with -18.5-dBm OMA Sensitivity Using Self-Timed Decision Feedback Equalization. Authors: Li Xu, Sanquan Song, Nikola Nedovic, Georgios Kalogerakis, Nandish Mehta, Angad Rekhi, Brian Zimmer, Stephen G. Tell, Yoshinori Nishi, Xi Chen, Ward Lopes, Benjamin G. Lee, Thomas H. Greer III, John W. Poulton, C. Thomas Gray (NVIDIA). Conference: 2026 IEEE Symposium on VLSI Technology and Circuits, paper C20.4, 2026 (DOI: 10.1109/VLSITECHNOLOGYANDCIR65830.2026.11577496).


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page