VLSI 2026 | Paper Analysis | Marvell Pushes a 3nm FinFET 212.5Gbps SerDes to 55dB Reach and 2.05pJ/bit: The Longest-Reach, Most Efficient Yet
At VLSI 2026, Marvell presented a 3nm FinFET 4-lane PAM4 transceiver running 212.5Gbps per lane that withstands 54.7dB of bump-to-bump channel loss, the longest-reach electrical link published at this data rate. The efficiency is even more striking: the analog side consumes only 2.05pJ/bit, also the lowest published at this rate. The transmitter measures 36dB SNDR, 0.97 RLM and 50fs integrated jitter, all top-tier. What makes this chip important is not any single number but that it is both the "longest" and the "most efficient" at once, two goals that usually fight each other in SerDes design.
1. Background: Why One SerDes DSP Deserves the Whole Optical Industry's Attention
This is not a paper about lasers or silicon photonics. It is about something more fundamental and more easily overlooked, yet underpinning the entire 1.6T generation: how to move electrical signals between chips, and across PCBs and packages, as far as possible with as little power as possible. The paper, "A 2.05pJ/bit, 212.5Gbps DSP based Transceiver with 55dB Reach in 3nm FinFET," was presented by Marvell Semiconductor at the 2026 IEEE Symposium on VLSI Technology and Circuits. The authors span four design centers in Santa Clara, Zhubei (Taiwan), Bangalore and Fishkill (New York), and the chip was taped out on TSMC's 3nm process. The industry has traditionally solved this problem by elimination: sacrifice reach for low power, or burn power for long reach. This paper's position is that all three dimensions can win at once, as long as the architecture is smart enough.
2. The Core Question in One Sentence
At 212.5Gbps, how can a shared-PLL, DSP-based architecture achieve both the world's longest reach (54.7dB) and best efficiency (2.05pJ/bit) without paying for it in jitter, linearity or area?
3. Architecture Overview: Figures 1-3, the Shared PLL and the CDR It Forced Into Existence
Figure 1 shows the top-level configuration of the 4-lane transceiver. Four transmitters (TX) and four receivers (RX) are driven by two shared PLLs rather than one PLL per lane. Sharing PLLs buys area efficiency and avoids the performance problems of crosstalk between multiple PLLs. The analog PLL contains an LC-VCO with dual coils and a dual-tail second-harmonic resonance, designed specifically to generate a low-jitter clock. The shared PLL is the key step to saving area and power, but it pushes the hard problem onto the receiver's clock and data recovery (CDR), which is where the paper's later technical innovations begin.

Figure 2 shows the full receiver (RX) block diagram, spanning the analog front end and DSP subsystem. This is an ADC-based receiver: the incoming analog signal is handed to digital signal processing (DSP) for equalization. The equalization front end is a three-stage CTLE (continuous-time linear equalizer), followed by 16 sampling networks with ADC buffers, each driving 8 sub-ADCs, for a total of 128 time-interleaved SAR ADCs running at up to 875Ms/s each. On the digital side, the DSP consists of a 29-tap feed-forward equalizer (FFE), a 1-tap decision feedback equalizer (DFE), and finally a maximum likelihood sequence detector (MLSD). This heavy-duty configuration is exactly why it can handle a brutal 54.7dB-loss channel.

Figure 3 shows the per-lane RX clocking architecture, the CDR design that the shared PLL forced into being. With the PLL shared across the chip, each RX lane needs its own way to fine-tune the sampling phase, and the paper's answer is one phase interpolator (PI) per lane. Each RX lane uses a 2-stage, 8-phase delay-locked loop (DLL) to drive a CML phase interpolator; the DLL first generates quadrature phases and then expands them into 8 phases spaced 45 degrees apart. The paper calls out a hidden killer in high-speed receivers: as the PI code rotates, CDR nonlinearity shows up as deterministic pattern jitter. Their solution uses phase-detection circuits to estimate the 90-degree and 45-degree phase errors and corrects them with de-skew buffers.

4. Proof of Phase Calibration: Figure 4
Figure 4 tests whether that phase calibration actually works in the most direct way: measuring bit error rate (BER) across different ppm frequency offsets. Without phase calibration, BER degrades noticeably as the ppm offset grows; with calibration, that ppm-dependent degradation curve is flattened. In real systems the reference clocks at the two ends always differ by some ppm, so a CDR that collapses at high ppm is unusable in actual data center deployments. This figure nails down that the shared-PLL architecture holds up in practice.

5. Transmitter Anatomy: Figures 5-6, the 21-tap FIR and the 8-bit DAC
Figure 5 shows the TX architecture block diagram. The transmitter rests on two pillars: equalization and clocking. TX DSP equalization uses a 21-tap FIR filter to compensate long-tail inter-symbol interference (ISI); the core DAC is 8-bit (6-bit binary + 2-bit thermometer). On the clock path, the differential PLL clock first passes through a CML buffer with duty-cycle correction (DCC), is converted to quadrature phases that feed the PI, and then goes through a divide-by-2 stage; these divided clocks drive the final multiplexer and a 64:4 serializer. Getting SNDR, RLM and jitter all to top-tier levels on the transmit side comes from stacking three things: DSP equalization, a high-resolution DAC, and jitter suppression across the entire chain.

Figure 6 shows the TX-DAC circuit details and how 1UI pulses are generated. Quadrature clocks first enter a timing-skew adjustment block, forming AC-coupled 1UI-wide pulses inside the DAC; these pulses sample and multiplex data to the pre-driver, which finally drives the output stage. The 4:1 DAC consists of time-interleaved differential 8-bit DAC slices whose output currents are summed into the final output current, with bandwidth optimized by a T-coil and resistive load. This current-mode DAC driver approach is the physical basis for achieving both high bandwidth (>64GHz) and low jitter.

6. Measurements: Figures 7-9, Eye Diagram, Reach and Jitter Tolerance
First, some context on clock quality: with a clock pattern, integrated RMS jitter measured at 26.6GHz is only 50fs (100Hz to Nyquist, 4MHz CDR), including all contributions from the PLL, clock distribution and transmitter; measured deterministic jitter (Dj) and total jitter (Tj) are 264fs and 971fs respectively.

Figure 7 shows the TX eye at 212.5Gb/s with a PRBS-13 pattern. Measured SNDR reaches 36dB (36.07dB), RLM is 0.97, Jrms is 6.6mUI and J3u is 54mUI; transmitter bandwidth exceeds 64GHz with a nominal output swing of 900mV. The RLM of 0.97 is especially important: PAM4 has four levels, and RLM measures how evenly those levels are spaced, with 1 being ideal, so 0.97 means very clean four-level signaling. A transmitter that pushes SNDR, RLM and jitter all to this level effectively leaves the signal integrity budget for the channel.

Figure 8 shows the receiver's results on the harshest channel, and this is the paper's signature number. At 54.7dB loss, pre-FEC BER is 1.7e-7. Testing used multiple variable trace lengths to cover different reach and loss targets, and multi-lane concurrent testing to emulate real crosstalk conditions; the 54.7dB result was measured with all four lanes active, not by cheating with a single lane. 54.7dB is the longest electrical reach published at this rate, and it determines how long copper cables or PCB traces in next-generation AI racks can run before having to switch to optics early.

Figure 9 shows jitter tolerance (JTOL) results. JTOL was tested on a 30dB channel at different ppm frequency offsets, and the measured curves show solid margin against the spec mask. This directly echoes Figure 4: phase calibration keeps the CDR standing under large ppm offsets, and JTOL is its stress-test report card. Passing JTOL means the chip is not just impressive under ideal lab conditions but reliable under the clock drift and jitter injection of real deployments.
7. The Silicon and Peer Comparison: Figure 10 and Table I

Figure 10 shows the die micrograph of the chip fabricated on TSMC 3nm. It corresponds to Table I: the chip's analog area including the PLL is 0.62mm² per lane. 0.62mm² is on the smaller side among its peers, which is what makes "high-density integration" more than a slogan.
Table I places this chip alongside three ISSCC 2025 transceivers at the same rate (212.5Gb/s), flattened into text below.

Two rows matter in this table. On reach, 54.7dB is the longest of the group; on power, 2.05pJ/bit is the lowest. The other three each make trade-offs: [3] has the best BER (1.8e-10) but only 40dB reach and a bloated 1.63mm² area; [2] has the second-lowest power (2.2pJ/bit) but reaches only 46dB. Only this chip stands on both peaks, "longest + most efficient," and its 54.7dB was measured with all four lanes active. The 3nm process certainly helps a lot, but the architecture (shared PLL, DSP equalization, DLL-PI CDR) is what brings the two together.
8. Technical Highlights: The Two Innovations Worth Remembering
First, a shared PLL paired with a per-lane DLL-PI CDR. A shared PLL inherently saves area and power, but at the cost of each lane losing independent clock fine-tuning. The paper restores that capability with a per-lane DLL driving a CML PI plus phase calibration, and additionally eliminates the deterministic pattern jitter caused by PI rotation. The two measurements in Figure 4 and Figure 9 double-validate this innovation. Second, heavy DSP equalization with 128 SAR ADCs + a 29-tap FFE + MLSD. This is what gives it the confidence to take on 54.7dB. When channel loss is so high that the eye is nearly closed, no amount of analog equalization is enough; the signal has to be digitized and recovered with a deep enough FFE plus sequence detection, and it is the efficiency of 3nm that makes this affordable.
9. Industry Link: How Close Is This Chip to Production, and Who Benefits
This is a silicon paper from Marvell, not a purely academic proof of concept: measurement data, four-lane concurrent testing and a wafer-level loopback test mechanism are all in place, meaning it is already at the threshold of productization. The direct beneficiaries are 800G/1.6T optical modules and large switch ASICs. The transceiver supports both ultra-long-reach electrical links and direct optical integration, so it can serve as the retimer/SerDes for in-rack copper and traces or as the DSP for the electrical segment in front of an optical engine. For a supply chain racing from 800G to 1.6T and planning 3.2T, a SerDes with longer electrical reach and lower power per lane directly widens the design space for system topology.
Conclusion
This chip's historical position is clear: it is the first electrical transceiver of the 212.5Gbps generation to win both longest reach and best efficiency at once, using 3nm plus an architecture smart to the point of being rebellious. Its signal to the industry is not "SerDes got a bit faster again" but "the PPA triangle no longer has to be sacrificed." The old iron rule, that reach costs power and low power means shorter reach, has been broken by the combination of shared PLL + DLL-PI CDR + heavy DSP equalization. And with electrical reach pushed as far as 54.7dB, where to draw the line between optics and electronics has to be renegotiated.
References
Paper: A 2.05pJ/bit, 212.5Gbps DSP based Transceiver with 55dB Reach in 3nm FinFET. Authors: M. Gambhir, A. Mostafa, R. Chen, F. Chu, Z. Guo, X. Han, M. Hasan, A. Hassan, E. Hsiao, A. Lahiri, Z. Li, P. Liu, F. Lu, K.-M. Lu, P. Ramakrishna, M. Shannon, U. Shukla, A. Singh, M. Singh, Y.-P. Su, D. Visani, D. Zhou, H. Wang, K. Chang (Marvell Semiconductor Inc.). Conference: 2026 IEEE Symposium on VLSI Technology and Circuits, 2026. DOI: 10.1109/VLSITECHNOLOGYANDCIR65830.2026.11577411.




Comments