Paper Analysis | Intel Takes VCSEL CPO Below 1 pJ/b: A Full Circuit Breakdown from 4×50G NRZ to 108G PAM4
While everyone is talking about silicon photonics CPO, Intel presented a paper at CICC 2026 going the other way: using cheap, low-power multimode VCSELs to build co-packaged optics (CPO) with a 2.9 pJ/b 4×50G NRZ link and a 0.9 pJ/b 108G PAM4 link. This isn't a single-chip showpiece but a "circuit techniques overview" that lays out every pitfall VCSEL CPO hits from transmitter to receiver, and proposes a clever trick: cascading two 2-tap stages into a 3-tap FFE. We read it section by section and figure by figure, and close with STT's view: VCSEL is not an old technology made obsolete by silicon photonics, but an underrated low-power path for short-reach CPO.
1. Background: why CPO has gone from "crying wolf" to a must-have
This paper comes from Intel's team in Hillsboro, Oregon, presented at the 2026 IEEE Custom Integrated Circuits Conference (CICC), with authors including Sashank Krishnamurthy and Susnata Mondal. It is positioned as an overview, consolidating the team's VCSEL CPO results previously scattered across ISSCC, VLSI and JSSC into one complete narrative, and adding new measurements of an on-chip 3-tap FFE.
The driving force fits in one sentence: large language models (LLMs) are pushing data centers to disaggregate compute and memory. Nodes need high-bandwidth, low-latency interconnect, but conventional electrical interconnect hits a bottleneck beyond 50 GBaud and needs power-hungry DSPs for equalization. The solution is to move the optical engine next to the XPU/switch: from pluggable optical modules, through near-package optics as a transition, to the end state, co-packaged optics (CPO), which places the optical engine and XPU/switch in the same package, shortening the electrical channel and minimizing loss.

There is an easily overlooked division of labor here: single-mode silicon photonics targets hundreds of meters to kilometers, while multimode VCSELs focus on reaches within a few tens of meters, trading for CPO with lower power and cost. This paper sits on the VCSEL side, and no matter how good the VCSEL device is, it can't run fast without matching high-speed circuit techniques. That is exactly the gap this paper fills.
2. Two prototypes: first, be clear on what we're dissecting
The first is a 4-channel, 50 Gb/s-per-channel NRZ CPO transceiver (TRX), with the optical engine connected to electrical TX/RX ICs (acting as an XPU proxy) through a 12 mm in-package electrical channel. The transmitter is a VCSEL driver (VCDRV), the receiver is a transimpedance amplifier front end (TIA-FE), and the VCSEL/PD arrays are wire-bonded to the package. The second is a 108 Gb/s PAM4 direct-drive optical engine.

The two take different approaches to fiber coupling. The first uses a mechanical optical interface (MOI): active alignment has low loss (1–3 dB) but is expensive, while passive alignment is cheap but lossy (5–6 dB). The second switches to direct optical wiring (DOW), 3D-printed polymer waveguides that shrink volume by 4× and height by 3×, with about 3 dB loss. At 250 µm pitch, crosstalk at 32 GHz is below −30 dB, harmless for 64G NRZ, but at 128G PAM4 this level of crosstalk is enough to visibly degrade the eye.
3. VCSEL and driver: equalization is the real battleground
A VCSEL is a directly intensity-modulated device, but its electro-optic transfer function is a first-order electrical response cascaded with an underdamped second-order complex pole, causing in-band gain and group-delay peaking. Ideally a complex zero compensates both magnitude and phase at once, but an FFE with a single post-cursor tap can only create a real zero.

Core insight: a 3-tap FFE can synthesize a complex zero, but its maximum zero frequency is tied to the damping factor. For a typical VCSEL with ζ<0.5, the zero tops out at about 14.4 GHz at 50 GBaud, below most VCSEL resonance frequencies (>20 GHz). Intel's solution is a CZ-CTLE: adding an inductor in series with the degeneration resistor to synthesize a complex zero, with tunable R, L and C. A complementary implementation improves linearity by more than 5×, and an active version lets the maximum zero frequency scale with process fT while saving inductor area.

Broadband amplification: topology B (shunt inductive peaking) wins, with group-delay distortion < 4 ps and gain peaking < 0.75 dB at 1.5× bandwidth extension. The output driver stage uses a CH-type driver, delivering higher transconductance at similar bandwidth for a larger optical modulation amplitude (OMA).

4. Receiver TIA: fighting package parasitics
The TIA-FE IC is flip-chipped onto the package and the PD is wire-bonded, separated by a short 0.6–0.8 mm interconnect. This short transmission line has asymmetric terminations at its two ends (wire-bond inductance LBW in series with PD capacitance CPD on one side, the TIA-FE RC on the other), causing large in-band group-delay distortion and ringing.

The only thing the designer controls is the on-chip matching network: a series inductor LS absorbs the parasitics so the CLC π network approximates an artificial transmission line, pushing group delay down to ~8 ps for NRZ and ~5 ps for PAM4. There is also a hidden linearity vs. noise trade-off: VCSEL RIN rises with optical power, so SNR eventually approaches the RIN limit. Since optical noise dominates at high power, they can afford a low-gain (low-RF) TIA in exchange for better linearity, and perform single-ended-to-differential conversion (SE2D) early to suppress even-order distortion.
5. Receiver equalization and the new 3-tap FFE: cascading two 2-tap stages into one 3-tap
A low-RF TIA with an input peaking inductor produces magnitude and group-delay peaking, and here a real-zero CTLE actually makes things worse; the group-delay dip of a CZ-CTLE is needed to cancel it. The front end's 3-dB bandwidth is about 20 GHz, and a 1-UI pulse leaves one precursor (stronger) and one postcursor ISI tap. For the first 50G NRZ chip, a 1/4-rate 2-tap Cherry-Hooper FFE compensating the precursor is enough.

PAM4 can't tolerate residual postcursor ISI, so both precursor and postcursor must be compensated. Three samples must be simultaneously valid for 1 UI, each held for 3 UI, which requires four-phase 25/75% duty-cycle clocks, a real headache. Intel's solution is elegant: split one 3-tap equalizer into two cascaded 2-tap stages. It holds as long as the tap strengths satisfy |b₋₁b₁| < 0.25, and only four-phase 50% duty-cycle clocks are needed throughout.
[Insert Fig. 7] Caption: Fig. 7 — 2-tap Cherry-Hooper FFE and 3-tap FFE timing with 25/75% duty-cycle clocks
PAM4 can't tolerate residual postcursor ISI, so both precursor and postcursor must be compensated. Three samples must be simultaneously valid for 1 UI, each held for 3 UI, which requires four-phase 25/75% duty-cycle clocks, a real headache. Intel's solution is elegant: split one 3-tap equalizer into two cascaded 2-tap stages. It holds as long as the tap strengths satisfy |b₋₁b₁| < 0.25, and only four-phase 50% duty-cycle clocks are needed throughout.

More importantly, it points to a path that extends to N taps: as long as an N-tap equalization polynomial can be factored into N real-coefficient 2-tap FFEs, it can be stacked using 50% duty-cycle clocks, avoiding power-hungry DSPs.
6. Measured results: two prototypes plus the new on-chip FFE
Everything is in 22 nm FinFET CMOS. First, the 4×50G NRZ CPO TRX: IC area is only 0.19/0.13 mm². The standalone TX runs 4×64G NRZ with 3× the OMA of prior work and 1.3 pJ/b efficiency (9 mA), dropping to 1.1 pJ/b at 5 mA. At 4×50G and 7.5 mA it achieves 0.26 UI eye opening and −6 dBm sensitivity; without FFE, the 50G eye is closed at 10⁻¹². End-to-end link efficiency is 2.9 pJ/b (3× better than prior work), and the standalone VCDRV reaches 80G NRZ.

Second, the 108G PAM4 direct-drive optical engine: the VCDRV reaches 2.5 mW outer OMA at 128G PAM4 with 0.31 pJ/b efficiency; the direct-drive link reaches a pre-FEC BER of 2.4×10⁻⁴ at 108G PAM4, and its 0.9 pJ/b efficiency is the best reported.

What's new in this paper: the cascaded 3-tap FFE actually implemented on chip (22 nm FinFET). At 100G PAM4, enabling the FFE opens the eye width by > 1 ps; with the FFE off, the eye closes and the system can't exceed 88 Gb/s. Optical engine efficiency is 1 pJ/b, with the FFE data path and local clocking adding 0.25 and 0.15 pJ/b respectively.

7. The thermal test: the real long-term variable for VCSEL CPO

Three findings: slope efficiency drops as substrate temperature rises (the main challenge in maintaining high OMA); bandwidth changes more slowly with temperature, degrading less than 10% at 85°C; and VCSEL+driver eye height degrades about 36% at 56G and 55°C, but because the complex-pole behavior is stable, the equalizer never needs retuning. In other words, temperature hurts "optical power", not the "response shape".
8. Conclusion
This paper shows that VCSEL CPO is not a technical dead end, but a path drowned out by silicon photonics hype that is highly competitive in short-reach, low-power scenarios. Its place in technical history has three layers: an integrated circuit methodology (CZ-CTLE + matching network + low-RF TIA), the cascaded 3-tap FFE (avoiding power-hungry DSPs and extensible to N taps), and an honest treatment of temperature.
STT's view: For the past two years the CPO narrative has been monopolized by silicon photonics, and VCSELs are often treated as an old guard about to be retired. But this paper does the math: 0.9 pJ/b at 108G PAM4 is among the most power-efficient numbers in the public literature, and VCSELs are far ahead of single-mode in cost and integration maturity. In XPU-to-XPU and accelerator-to-memory scenarios that demand "within a few tens of meters, ultra-low power, ultra-low cost", VCSEL CPO's price-performance is hard to ignore. But three limits deserve a sober look: 250 µm pitch is a ceiling (flip-chip optical attach is still future work), PAM4 linearity and crosstalk margins are thin, and most critically, temperature (if thermal management isn't solved, volume-production reliability is in doubt). The paper's real value is in clearly marking the remaining miles before VCSEL CPO reaches volume: equalization, packaging and thermals.
Reference: S. Krishnamurthy, S. Mondal, J. Qiu, S. Yamada, Z. Zhou, J. Kennedy, J. Jaussi, and M. Mansuri, "50-GBaud+ VCSEL-Based Co-Packaged Optical Links: An Overview of Circuits and Systems," 2026 IEEE Custom Integrated Circuits Conference (CICC), Paper 19-1, Intel Corporation. DOI: 10.1109/CICC65509.2026.11509602.
This article is a technology and industry trend analysis based on a publicly published IEEE paper and does not constitute investment advice.




Comments