top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

Tech Analysis | Nature Electronics Opens the Full CPO Ledger: From a 1.4× Bandwidth Wall to a 100 fJ/bit Monolithic Endgame

2 days ago
28 min read


  • This review, published in Nature Electronics in August 2026, is the most complete ledger to date that takes CPO all the way from devices through packaging to systems. The author list spans the University of Virginia, Yonsei, UIUC, NTU Singapore and MIT, and the eighth author is listed under SK hynix AI Infra Optimization Team—a memory maker has formally taken a seat at the optical interconnect table.

  • There is only one core number: compute grows 3.0× every two years and AI models 8.8×, while interconnect bandwidth grows only 1.4×. That is the quantitative definition of the bandwidth wall, and the sole reason CPO exists.

  • The most valuable thing about this review is not "CPO is good," but that it plots the pJ/bit and Tbps/cm² of five generations—2D, 2.5D, 3D TSV, 3D hybrid bonding and monolithic—on a single chart: from 10–20 pJ/bit down to below 100 fJ/bit, and bandwidth density from 100 Gbps per package up to 50–100 Tbps/cm², with three walls in between—thermal, yield and standards.

  • Below we annotate and break down, in place, every one of the 24 sub-panels in Figures 1–4 plus Tables 1–2. If you want to skim, jump straight to the panel you care about.

1. The weight of this review, and the problem it sets out to solve

First, to be clear, this is not an experimental paper but a Review article: received on 18 November 2025, accepted on 8 July 2026, published online on 19 August, and printed in Nature Electronics Volume 9, pages 853–867.

A review's value depends on who writes it, and this author list deserves a second look: Jeehwan Kim (MIT, a leading figure in remote epitaxy and 2D-material transfer), Kyusang Lee (UVA, thin-film heterogeneous integration), Sang Hoon Chae (NTU), and Seunghoon Hong of the SK hynix AI Infra Optimization Team. In other words, this is the materials/heterogeneous-integration school's full position on CPO, not a roadmap written by optical-communications insiders for other insiders.

The problem it tackles fits in one sentence:

Now that the compute bottleneck has shifted from "can we compute fast enough" to "can we move data fast enough," unless the interconnect architecture is fundamentally redesigned, any progress in transistor scaling or logic throughput will yield only marginal system-level benefit.

That sentence appears in the second paragraph of the paper and is the spine of all 15 pages. HBM stacks DRAM with TSVs and does reach multi-terabits-per-second and pJ/bit-class efficiency inside the package—but that benefit exists only within the memory–logic package boundary. Once data has to leave the package over PCB traces, board connectors and DAC copper cables, resistive and capacitive losses deteriorate sharply with frequency and distance. That cliff—great inside the package, poor outside it—is where the whole paper starts.

2. Figure 1: three pieces of evidence for the bandwidth wall

Figure 1a: 8.8× vs 1.4×—the quantitative definition of the bandwidth wall

What the figure shows: four historical growth curves from 1995 to 2025, each normalized to its earliest data point. Blue is hardware compute throughput (FLOPS), red is DRAM bandwidth (GB/s), yellow is interconnect bandwidth (GB/s), and grey is total AI-model training compute (right axis, FLOPs).

Key numbers: hardware 3.0× per 2 years, DRAM 1.6× per 2 years, interconnect 1.4× per 2 years, AI models 8.8× per 2 years. Several anchors are marked—AlphaGo, GPT-4.5 and Grok-4 on the grey band, and NVIDIA B300, HBM3e and NVLINK 5.0 at the ends of the blue, red and yellow lines respectively.

Industry significance: this is the most honest bandwidth-wall chart I have seen. It does not say interconnect isn't improving; it says interconnect's slope is less than half that of compute and only one-sixth of AI-model demand. The slope gap compounds exponentially over time, so any fix based on "just add more copper lanes" only postpones the breaking point. CPO is not a performance bonus; it is a slope correction.

Figure 1a: growth slopes from 1995–2025 for compute throughput (3.0×/2 yr), DRAM bandwidth (1.6×/2 yr), interconnect bandwidth (1.4×/2 yr) and AI-model training compute (8.8×/2 yr). Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 1a
Figure 1a: growth slopes from 1995–2025 for compute throughput (3.0×/2 yr), DRAM bandwidth (1.6×/2 yr), interconnect bandwidth (1.4×/2 yr) and AI-model training compute (8.8×/2 yr). Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 1a


Figure 1b: plotting "bandwidth density × energy efficiency" against distance as a falling line

What the figure shows: a composite figure of merit: bandwidth density × energy efficiency, in (Gbps/mm) ÷ (pJ/bit), with maximum interconnect distance (metres) on the x-axis from 10⁻⁴ m to beyond 10² m, divided by vertical dashed lines into In-package / On-board / Off-board regions.

Key numbers: the solid red line is the projected trend for electrical interconnects, anchored at 1 Tbps/mm bandwidth density with 1 pJ/bit efficiency. Measured points are laid out by distance: In-package has UCIe Advanced and HBM (highest, about 10⁴), On-board has NVLink C2C and PCIe Gen5, and Off-board has PCIe Gen5 and pluggable-optics Ethernet SR4. The dashed blue line is the projected optical trend and the CPO target zone, clearly above the red line at long reach; a board-mounted optical backplane sits in the middle.

Industry significance: this is the chart in the review most worth screenshotting and sharing. It turns "when does optics beat copper" from a slogan into a crossover point you can look up—optics doesn't win at every distance; rather, the red line falls so fast that the crossover is moving into the package. UCIe Advanced still leads comfortably in-package today, but beyond on-board, electrical interconnect is in near free fall—which is exactly the physical reason CPO pulls optics into the package. We discussed this "copper wall moving into the rack" theme in our Computex 2026 Keynote: The Copper Wall Is Moving Into the Rack coverage; this chart effectively gives it academic coordinates.

Figure 1b: using "bandwidth density ÷ energy per bit" as the figure of merit, the trend crossover between electrical (red line) and optical (dashed blue) interconnects across In-package / On-board / Off-board distances. Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 1b
Figure 1b: using "bandwidth density ÷ energy per bit" as the figure of merit, the trend crossover between electrical (red line) and optical (dashed blue) interconnects across In-package / On-board / Off-board distances. Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 1b

Figure 1c: a three-tier deployment map for optical interconnects

What the figure shows: hierarchical deployment of optical interconnects across three spatial scales. On the far left is CPO (μm–mm), with HBM and an ASIC side by side on a substrate, a photonic engine alongside, fibre exiting the package side, and the SerDes and optical–electrical conversion locations marked; in the middle is intra-node (cm–m), with DRAMs / CPUs / GPUs stacked in three layers and linked by high-speed copper cable; on the far right is inter-rack (m–km), high-speed links between server racks.

Key numbers: the three tiers sit on a continuous close-to-far axis, matching the paper's three-tier description—CPO suppresses parasitic loss and latency at the μm-to-mm scale, intra-node spans cm to m and connects processors and memory within a single compute node, and inter-rack covers data-centre scale from metres to kilometres.

Industry significance: the value of this figure is that it helps each supplier locate which tier it sits in. FAU, optical-engine and interposer players are in the left tier; DAC/AEC/backplane players in the middle; 800G/1.6T pluggable and DCI players on the right. The paper explicitly notes that the middle tier (intra-node) is where the copper-vs-optics tug-of-war is fiercest today, and it is the main battlefield for scale-up networks. Most of the opportunity for Taiwanese suppliers lies in the left and middle tiers, not the right.

Figure 1c: three-tier hierarchical deployment of optical interconnects, from package level (μm–mm) to intra-node (cm–m) to inter-rack (m–km). Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 1c
Figure 1c: three-tier hierarchical deployment of optical interconnects, from package level (μm–mm) to intra-node (cm–m) to inter-rack (m–km). Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 1c

3. Table 1: a health check on five key components

What the table shows: for a hybrid optical interconnect link, the five key components (laser, modulator, detector, waveguide, coupler)—their benchmark metrics, current state of the art, integration maturity, coupling/packaging approach, and thermal/operational risks.

Key numbers:

  • Component: laser; approach: III–V bonding; key metrics: >5 mW, linewidth ~40 kHz; maturity and risk: wafer bonding advancing, alignment-sensitive, CTE-mismatch risk

  • Component: laser; approach: transfer printing; key metrics: 1–5 mW, <1 MHz; maturity and risk: lab-scale, adhesive bonding thermally limits CW output

  • Component: laser; approach: direct epitaxy; key metrics: <1 mW, 1–5 MHz; maturity and risk: early-stage, defect-limited, drift risk

  • Component: modulator; approach: Si carrier depletion; key metrics: >50 GHz, multi-pJ; maturity and risk: CMOS-foundry ready, but with tuning overhead

  • Component: modulator; approach: thin-film LN; key metrics: >100 GHz, sub-pJ; maturity and risk: wafer bonding maturing, fibre attach difficult

  • Component: modulator; approach: plasmonic / ferroelectric; key metrics: >100 GHz, <1 fJ/bit; maturity and risk: early-stage, stability issues

  • Component: detector; approach: Ge-on-Si PD; key metrics: 265 GHz, 0.3 A/W; maturity and risk: wafer-scale, absorption vs bandwidth trade-off

  • Component: detector; approach: Si APD; key metrics: ~40 GHz, ~0.4 A/W; maturity and risk: wafer-scale, excess noise

  • Component: waveguide; approach: silica; key metrics: <0.1 dB/m, low density; maturity and risk: mature but large footprint

  • Component: waveguide; approach: SOI; key metrics: 1–3 dB/cm, high density; maturity and risk: on the CMOS roadmap, tuning overhead

  • Component: waveguide; approach: SiN; key metrics: ~0.1 dB/m, high density; maturity and risk: back-end compatible, heater drift

  • Component: coupler; approach: edge; key metrics: <1 dB, ±2 μm; maturity and risk: high yield, but alignment-sensitive

  • Component: coupler; approach: grating; key metrics: 3–6 dB, ±1 μm; maturity and risk: wafer-scale capable, polarization-dependent

  • Component: coupler; approach: freeform; key metrics: 1–2 dB; maturity and risk: early-stage, complex fabrication

Industry significance: the rows to underline in this table are the three coupler rows. Edge couplers lose <1 dB but demand ±2 μm alignment; grating couplers are listed at ±1 μm and can be done at wafer scale, at the cost of 3–6 dB insertion loss and polarization dependence. What does 3 to 6 dB mean? Optical power drops by half to three-quarters outright. The entire CPO debate over packaging yield and link budget is, at its core, a trade-off between these two rows—which is why SENKO framed CPO's mass-production bottleneck at OCP as "that 0.15 dB". Also note the <1 fJ/bit in the plasmonic/ferroelectric row: that is today's theoretical ceiling for modulator efficiency, but the maturity column honestly says early stage.

Table 1: benchmark metrics, state of the art, integration maturity and thermal risks for the five key optical-interconnect components (laser / modulator / detector / waveguide / coupler). Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Table 1
Table 1: benchmark metrics, state of the art, integration maturity and thermal risks for the five key optical-interconnect components (laser / modulator / detector / waveguide / coupler). Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Table 1

4. Figure 2: breaking an optical link into seven parts

Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 2
Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 2

Figure 2 is the review's "component textbook page": seven sub-panels, each matching one of the seven stages of a complete link.

Figure 2a: an end-to-end view of the link

What the figure shows: a complete high-bandwidth optical data link. On the left, the Tx side forms the electrical domain from ASIC + SerDes, an E/O converter turns it into an optical carrier, and a wavelength multiplexer combines channels into the fibre; on the right, the Rx side runs in reverse—a demultiplexer separates wavelengths, an O/E converter returns them to electrical, and SerDes restores parallel data for the ASIC.

Key numbers: the figure deliberately draws N parallel sets of ASIC/SerDes/converters, matching the paper's description of "multiple optical chiplets placed side by side on the accelerator package."

Industry significance: this figure's use is to help you confirm which box your own product sits in. Typical positions for Taiwanese suppliers: optical-engine makers in the E/O and O/E boxes, passive-component makers in the multiplexer box, and packaging houses spanning all of them. The real high margins are in the E/O and mux boxes—which happen to be the two hardest boxes to yield today.


Figure 2b: SerDes—the stage everyone wants to kill but nobody has killed yet

What the figure shows: the architecture of a serializer/deserializer (SerDes): on the left, Tx compresses multiple parallel data lanes into one high-speed serial stream; on the right, Rx splits it back into parallel. Below are the corresponding waveforms—parallel 0/1/0/1 stacked into one densely timed serial waveform.

Key numbers: the paper cites 112 Gb/s SerDes still at 1–2.3 pJ/bit. By comparison, silicon photonic transmitters themselves have already reached 0.7 pJ/bit.

Industry significance: this is the most awkward stage in the whole link. The optical part already uses less power than the electrical part, yet SerDes still eats 1–2.3 pJ/bit. That is exactly why the "micro-LED eliminates SerDes" route in Figure 4f exists—and why, when Lightmatter Passage splits its 4.6 pJ/bit into "2.6 pJ optical + roughly 2 pJ co-packaged SerDes," that latter 2 pJ stands out so glaringly.


Figure 2c: NRZ vs PAM-4—eye diagrams that make the cost of spectral efficiency clear

What the figure shows: a comparison of two modulation formats. On the left, NRZ (non-return-to-zero) has only two levels, 0 and 1, giving one clean, wide eye; on the right, PAM-4 (four-level pulse amplitude modulation) uses four levels, stacking into three visibly smaller eyes, with the 11 / 10 / 01 / 00 symbol mapping marked.

Key numbers: PAM-4 has 2× the spectral efficiency of NRZ (twice the bits at the same baud rate), at the cost of a compressed noise margin—three small eyes instead of one big one.

Industry significance: this figure explains why DSP and FEC weigh ever more heavily in high-speed links. Later in the paper there is a hard number: coherent equalization power scales almost linearly with tap count, rising from about 1 pJ/bit at 4 taps to more than 5 pJ/bit at 16 taps. In other words, part of the bandwidth PAM-4 saves is bought back with DSP power. That is also why routes like LPO/LRO, which "remove the DSP," keep attracting bets.


Figure 2d: the Mach–Zehnder modulator—writing electrical signals into light through interference

What the figure shows: how a Mach–Zehnder modulator (MZM) works: a continuous-wave laser at the lower right feeds light in, which is split into two arms; each arm receives the electrical data input, and phase modulation makes the two arms interfere where they recombine, producing an intensity-modulated signal at the output.

Key numbers: the paper notes that interferometric phase modulation in an MZM delivers high linearity and bandwidth beyond 100 GHz, at the cost of a larger footprint that limits integration density. By contrast, carrier-depletion silicon modulators (p–n or p–i–n junctions) are CMOS-compatible and scalable, but the plasma dispersion effect is inefficient in silicon, requiring longer phase-shift sections and higher drive voltages.

Industry significance: this figure is the premise for understanding why TFLN is so sought after. With the same MZM structure, switching to Pockels-effect ferroelectric materials such as thin-film lithium niobate (TFLN) or barium titanate (BaTiO3, BTO) lets you run above 100 GHz at sub-volt drive. The paper states plainly that TFLN is today's benchmark for integrated high-speed modulation, while BTO has an even larger electro-optic coefficient and suits ultra-compact, low-power modulators; however, BTO must be grown epitaxially on SrTiO3-buffered silicon at temperatures beyond the CMOS back-end thermal budget, and remains pre-commercial.


Figure 2e: VCSEL and photodiode + TIA—the two ends of the active optoelectronics

What the figure shows: a pair of active optoelectronic devices. The top half is a driver-driven VCSEL (vertical-cavity surface-emitting laser), emitting light vertically from its surface; the bottom half is a photodiode paired with a TIA (transimpedance amplifier), where absorbed light becomes photocurrent that is then amplified into a voltage signal.

Key numbers: the detector figures in Table 1 map to this panel—Ge-on-Si photodetectors at 265 GHz bandwidth and 0.3 A/W responsivity; Si APDs at about 40 GHz and about 0.4 A/W, trading internal gain for sensitivity at the cost of excess noise and higher bias.

Industry significance: this is the most easily overlooked yet most pragmatic panel. The paper makes two points clearly. First, monolithic or co-packaged photodiode–TIA integration reduces parasitic capacitance and improves overall efficiency, but demands strict thermal and impedance matching; second, in practice short-reach links still rely mainly on p–i–n plus CMOS TIA, with APDs and coherent front ends explored to extend sensitivity and reach. This means the VCSEL route is not out—Furukawa-style 1060 nm single-mode VCSELs running 2 km are reasonable within this framework.


Figure 2f: the microring resonator—routing, filtering and multiplexing with one ring

What the figure shows: three roles of a wavelength-selective microring resonator: on the left, a wavelength-selective add-drop structure; in the middle, the coupling between the microring and the bus waveguide, with frequency spectra drawn above and below showing a specific wavelength being picked out while the rest pass through.

Key numbers: the microring's advantage is extreme compactness and integration density; the cost, as the paper puts it bluntly, is thermal sensitivity. The quantification later in the paper: silicon microrings drift by tens of GHz per K, while semiconductor laser emission drifts by tens of picometres per K. By contrast, arrayed waveguide gratings (AWGs) have a large footprint but good thermal stability.

Industry significance: this figure is the technical core of the DWDM route and the epicentre of CPO's thermal-management problem. The microring has to work next to the ASIC, and the ASIC's temperature fluctuates—tens of GHz of drift per K means you need heaters and closed-loop control, and heaters themselves consume power and generate heat. The paper describes this as a feedback loop that "destabilizes optical power and spectral coherence." That is also why athermal designs (athermal resonators in SiN or TFLN) are specifically called out in the paper's outlook.


Figure 2g: grating coupler vs edge coupler—the first yield gate in CPO packaging

What the figure shows: two chip-to-fibre coupling strategies. The top half is a grating coupler, with fibre coming down vertically (vertical fibre coupling) onto a Si carrier stacked with electronics and waveguides; the bottom half is an edge coupler, with fibre entering horizontally from the side (in-plane alignment).

Key numbers: per Table 1—edge coupler <1 dB insertion loss, ±2 μm alignment tolerance; grating coupler 3–6 dB, ±1 μm, wafer-scale capable but polarization-dependent with limited optical bandwidth. The paper also gives a system-level figure: each chip-to-fibre interface contributes 1–3 dB of coupling loss.

Industry significance: this is the most painful step in taking CPO from lab to production line—and the one where Taiwan's supply chain has the best chance to stake a position. 1–3 dB per interface doesn't sound like much, but a CPO package has tens to hundreds of interfaces, and fibre-burn risk grows as per-channel optical power rises. The paper adds another cut: compared with integrated light sources, external-laser architectures pay an extra 1–5 dB of link budget at the same bandwidth, and that penalty becomes sharper as WDM scales to tens of channels.


5. Figure 3: five steps of co-packaging, each with a pJ/bit price tag

Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 3
Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 3

Figure 3 is the most valuable figure in the review. Its eight sub-panels fall into two groups: a–d are four distances describing "how far the optical engine is from the SoC," and e–h are five generations of "how PIC and EIC are stacked."

Figure 3a: pluggable optical modules—optics at the board edge, farthest from the SoC

What the figure shows: a cross-section of the traditional pluggable architecture: the SoC sits in the middle of the PCB, a long copper trace runs across the board (labelled Far), reaching the photonic engine only at the board edge, where the fibre exits.

Key numbers: the paper's quantitative description—centimetre-scale copper traces from centre to edge cause significant resistive loss, frequency-dependent signal distortion and power consumption, and at terabit rates board copper traces longer than a few centimetres incur tens of dB of attenuation.

Industry significance: this is today's 800G/1.6T pluggable world. It won't die immediately—the paper itself says electrical interconnects remain irreplaceable at short reach thanks to manufacturing maturity and CMOS compatibility—but its physical ceiling has now been drawn.


Figure 3b: on-board optics (OBO)—moving the TRx next to the SoC

What the figure shows: a cross-section of on-board optics: the transceiver (TRx) moves right next to the SoC (labelled Close), copper traces shrink dramatically, and embedded waveguides in the PCB carry optical routing.

Key numbers: the paper describes this as co-locating PIC and EIC directly on the board to reduce electrical loss and parasitic delay.

Industry significance: OBO is a stop the industry has already passed through and largely abandoned. Its problem was commercial, not physical—it shortened the copper but solved neither serviceability nor yield, while losing the modular advantages of pluggables. Its role in the review is transitional: proving that 3a to 3c is a continuous physical evolution rather than a leap.


Figure 3c: CPO—SoC and TRx share one interposer

What the figure shows: a cross-section of co-packaged optics: the SoC and TRx are packed together on an interposer, which is mounted on the PCB, with fibre exiting directly from the TRx. This is the paper's canonical CPO definition diagram.

Key numbers: this step squeezes interconnect length down to the μm-to-mm scale. The measured result cited by the paper: recent CPO implementations have demonstrated 50% lower power than pluggable transceivers, with scalability beyond 12.8 Tbps per link.

Industry significance: read that 50% carefully—it is a saving relative to pluggables, not proof of hitting the absolute efficiency target. Looking at Table 2, Broadcom Bailly is 6.9 pJ/bit and Davisson 5–7 pJ/bit, still an order of magnitude away from the paper's sub-pJ/bit target. First-generation CPO solves "save half the power," not "reach the theoretical limit". That is why the downward curve in Figures 3e–g is the real roadmap.


Figure 3d: vertical integration—TRx stacked directly on the SoC

What the figure shows: a cross-section of a heterogeneous integration platform: the TRx no longer sits beside the SoC but is stacked vertically directly on top of the SoC (arrows mark vertical integration), with the interposer and PCB still below and fibre exiting from the TRx on top.

Key numbers: the paper describes this configuration as achieving high-density electro-optical interconnects through vertical assembly and interposer-level routing. The related 3D bandwidth-density figures are given in Figure 3f.

Industry significance: 3d is the bridge from 3c to 3f, and also the configuration that most challenges thermal management. Stacking light-emitting, highly temperature-sensitive photonic devices directly above the hottest logic die saves distance physically but creates an engineering nightmare. The paper's whole later section on the "thermal–optical coupling feedback loop" is about the cost of this figure.


Figure 3e: 2D and 2.5D—wire bonding at 10–20 pJ/bit, micro-bumps at 2–5 pJ/bit

What the figure shows: a side-by-side comparison of two generations. Top, 2D: EIC and PIC side by side on the same substrate, with a close-up of wire bonding and a real package photo on the right. Bottom, 2.5D: EIC and PIC on an interposer, with a close-up of micro-bumps and solder alloys, and on the right a cross-sectional SEM showing TSVs, a silicon photonics interposer, TIA/driver and a multi-layer organic substrate.

Key numbers: this is where the figure is most ruthless—every tier is labelled directly with pJ/bit and bandwidth density.

  • 2D wire bonding: about 10–20 pJ/bit, about 100 Gbps per package

  • 2.5D micro-bump: about 2–5 pJ/bit, 1–2 Tbps/cm²

The paper adds: frequency-dependent impedance distortion and capacitive loading from wire bonds limit achievable frequencies for high-speed applications to below 30 GHz; 2.5D micro-bumps have 5–40 μm pitch and TSVs 0.2–5.0 μm diameter, supporting transmission above 100 Gbps.

Industry significance: from 2D to 2.5D, efficiency improves 4–5× and bandwidth density by more than an order of magnitude. This is the step already in volume production today—both Broadcom Bailly and Davisson are 2.5D. For the supply chain, this means CoWoS-class interposer capacity is the physical bottleneck of CPO's first wave. That judgement—advanced packaging becoming the gatekeeper of AI compute—TSMC already spelled out at OCP 2026.


Figure 3f: 3D TSV and hybrid bonding—the 0.1 pJ/bit, 50 Tbps/cm² tier

What the figure shows: two kinds of 3D stacking. Top, 3D TSV: EIC stacked on EIC/interposer, with a close-up of Cu-filled through-silicon vias and a back-side TSV cross-sectional SEM on the right. Bottom, 3D hybrid bonding: a close-up of a direct-bond interface of interleaved metal and dielectric, with an SEM of a row of hybrid-bonding pads on the right.

Key numbers:

  • 3D TSV: about 0.5–2 pJ/bit, >5 Tbps/cm²

  • 3D hybrid bonding: about 0.1–0.5 pJ/bit, 10–50 Tbps/cm²

The paper adds that conventional micro-bump die-to-die pitch is typically >10 μm, with each bump adding tens of femtofarads of parasitic capacitance; hybrid bonding replaces them with direct Cu–Cu contacts embedded in dielectric, shrinking interconnect length to tens of micrometres with sub-micrometre pitch—wafer-to-wafer bonding has demonstrated about 400 nm, more than an order of magnitude smaller than micro-bump pitch. Combined with TSV vertical routing, hybrid bonding can support >5 Tbps/mm² of bandwidth density in 3D photonic–electronic stacks.

Industry significance: this is the single most important conclusion in the review—hybrid bonding is not a yield optimization; it is a phase change in energy efficiency. The reason is spelled out: it eliminates the micro-bump parasitic capacitance that dominates switching energy at die boundaries. And ultra-short interconnects reduce high-frequency loss, allowing retimers and complex DSP circuitry to be dropped, further improving total power—that, not just physical shrinkage, is the real source of 0.1 pJ/bit. Taiwanese companies' hybrid-bonding positions will pay off directly on the CPO line.


Figure 3g: monolithic integration—the endgame below 100 fJ/bit at 50–100 Tbps/cm²

What the figure shows: monolithic integration: EIC and PIC are no longer two chips but co-fabricated on one substrate in a unified process flow. The close-up shows transistors and waveguides coexisting in the same Si layer, with a microring resonator marked; the SEM on the right shows an optical bus waveguide, analogue front-end circuits and a microring modulator on one die.

Key numbers: <100 fJ/bit, >50–100 Tbps/cm². This is the extreme of the whole of Figure 3.

Industry significance: the paper is honest about this panel—it is the ultimate convergence, but depends on breakthroughs in materials integration and process co-optimization. The only approach at the highest process maturity today is CMOS-compatible silicon photonics (germanium photodetectors plus high-speed silicon modulators co-fabricated with logic transistors); bringing III-V, BTO/LN or 2D semiconductors into a monolithic flow is still held back by material incompatibility, thermal limits and process co-optimization. In industry terms: this panel is a post-2030 story, but it is the directional indicator for the hybrid-bonding tier.



Figure 3h: three radar charts that sum up the trade-offs of all five generations

What the figure shows: three overlaid radar charts with seven axes: technology maturity, cost (low), bandwidth, energy efficiency, thermal management, fabrication accessibility and footprint (small). The five integration strategies are colour-coded—2D (grey), 2.5D (dark grey), 3D TSV (yellow), 3D hybrid (red) and monolithic (blue).

Key numbers: the shapes themselves are the conclusion. The 2D/2.5D polygons lean toward maturity, cost and fabrication accessibility; the 3D polygons tilt toward bandwidth and energy efficiency but pull in on thermal management and fabrication accessibility; the blue monolithic polygon almost touches the outer ring on bandwidth, energy efficiency and footprint, but collapses the most on maturity and cost.

Industry significance: this chart should be printed and pinned on the desk of everyone making CPO decisions. Its message is not "which is best," but that these five generations will not replace one another; they will coexist for a long time, each serving a different cost/performance band. The paper's own conclusion: 2D offers simplicity and low cost but limited bandwidth; 2.5D offers moderate density and manufacturability; 3D pushes bandwidth density and efficiency to the limit at the cost of yield and thermal complexity; monolithic promises the highest performance but hinges on materials-integration breakthroughs. That is also why the CPO / NPO / XPO panel at OCP concluded "this is not a battle of routes".


6. Table 2: ten industrial platforms already running, and their real report cards

What the table shows: ten representative industrial CPO and optical I/O platforms, with columns for architecture, bandwidth, latency, energy, integration approach and reach.

Key numbers (summarized from the original table):

  • Company: Broadcom; platform: Bailly (TH5); architecture: switch CPO + ASIC; bandwidth: 51.2 Tbps; latency: sub-10 ns/hop; energy: 6.9 pJ/bit (5.5 W per 800G port); integration: 2.5D; reach: rack

  • Company: Broadcom; platform: Davisson (TH6); architecture: switch CPO + ASIC; bandwidth: 102.4 Tbps; latency: sub-10 ns/hop; energy: 5–7 pJ/bit (estimated); integration: 2.5D; reach: rack

  • Company: NVIDIA; platform: Quantum-X Photonics (COUPE SiP); architecture: switch CPO; bandwidth: 115.2 Tbps; latency: sub-10 ns/hop; energy: 5.6 pJ/bit (9 W per 1.6T port); integration: 3D; reach: rack

  • Company: Ayar Labs; platform: TeraPHY; architecture: optical I/O chiplet (UCIe); bandwidth: 8 Tbps bidirectional; latency: ~10 ns per chiplet hop; energy: <5 pJ/bit; integration: chiplet (SiP); reach: mm–km

  • Company: Intel; platform: OCI; architecture: optical I/O chiplet (PIC+EIC); bandwidth: ~4 Tbps bidirectional; latency: <10 ns plus ToF; energy: 5 pJ/bit; integration: co-packaged + laser; reach: rack

  • Company: Ranovus; platform: scale-up engines; architecture: optical engines (CPO); bandwidth: multi-Tbps; latency: low; energy: low (laser); integration: co-packaged; reach: not listed

  • Company: Lightmatter; platform: Passage M1000; architecture: 3D photonic interposer; bandwidth: 114 Tbps (4,000 mm²); latency: sub-10 ns; energy: 4.6 pJ/bit wall-plug; integration: chip-on-wafer (3D); reach: 10 m–km

  • Company: Celestial AI; platform: Photonic Fabric (PFLink); architecture: chiplet + system fabric; bandwidth: 14.4 Tbps; latency: sub-10 ns; energy: not listed; integration: chiplet + SiP; reach: >50 m

  • Company: Avicena; platform: LightBundle; architecture: parallel micro-LED transceivers; bandwidth: ~1 Tbps (304 ch × 3.3 Gbps); latency: micro-LED native; energy: ~1 pJ/bit; integration: CMOS (16 nm) + micro-LED array; reach: <5 m

  • Company: Microsoft; platform: MOSAIC; architecture: parallel micro-LED transceivers; bandwidth: ~1 Tbps (460 ch × 2 Gbps); latency: a few ns (no DSP/FEC); energy: 6.6 pJ/bit (5.3 per end); integration: CMOS + micro-LED (pluggable); reach: <50 m

Industry significance: three observations worth underlining.

First, NVIDIA is the only switch platform in the table marked as 3D co-packaged. Both Broadcom generations are 2.5D. Against the pJ/bit curve in Figures 3e–f, this architectural difference matters far more than the bandwidth numbers.

Second, the optical I/O chiplets from Ayar Labs and Intel follow an entirely different business logic—not putting optics into the switch, but into the XPU, with reach from millimetres to kilometres. Ayar's 8 Tbps bidirectional and <5 pJ/bit, plus UCIe compatibility, are the entry ticket to scale-up fabrics.

Third, the two micro-LED rows are the most overlooked but strongest signals in the table. Avicena's 1 pJ/bit is the table's efficiency champion, and Microsoft MOSAIC's few ns with no DSP/FEC is its latency champion. The cost is noted in the text: micro-LED emission spectra are wider than 10 nm (versus under 1 pm for lasers), and dispersion caps reach at tens of metres. But in CPO's centimetres-to-metres distance band, dispersion is simply not a problem—which sets up Figure 4f.


Table 2: bandwidth, latency, energy and integration approach for ten industrial CPO and optical I/O platforms, including Broadcom, NVIDIA, Ayar Labs, Intel, Lightmatter, Celestial AI, Avicena and Microsoft. Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Table 2
Table 2: bandwidth, latency, energy and integration approach for ten industrial CPO and optical I/O platforms, including Broadcom, NVIDIA, Ayar Labs, Intel, Lightmatter, Celestial AI, Avicena and Microsoft. Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Table 2

7. Figure 4: the photonic interposer, and a death sentence for SerDes

Figure 4 is the review's outlook figure; its six sub-panels depict what things look like beyond 2030.

Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 4
Source: Kim, B. et al. Nat. Electron. 9, 853–867 (2026) - Figure 4

Figure 4a: CPO grown directly on the EIC

What the figure shows: two ways of integrating optics into the EIC, side by side: on the left, 3D integrated optics with EIC (a photonic layer 3D-stacked on the electronic logic die); on the right, monolithic integrated optics with EIC (photonics and electronics co-fabricated monolithically). Both sit on chiplets, an electrical interposer and a substrate, with fibre exiting from the side.

Key numbers: this panel corresponds to the efficiency figures of Figures 3d/3f/3g—0.1–2 pJ/bit for 3D integration and below 100 fJ/bit for monolithic.

Industry significance: this figure pushes the definition of "CPO" one layer further inward. Today the industry uses CPO to mean "the optical engine and ASIC in the same package"; this figure means "the optical layer is itself part of the ASIC." The former is a packaging-house business; the latter is a foundry business. This shift in who owns the definition will decide where value concentrates over the next decade.


Figure 4b: a photonic interposer network—optically linking an XPU pool and a memory pool

What the figure shows: a photonic interposer acting as the intra- and inter-node network: on the left an XPU pool, on the right a memory pool, each sitting on its own photonic interposer, with PIC/EIC and fibres marked alongside and an optical path connecting the two pools.

Key numbers: the paper describes this panel as "connecting compute and memory chiplets across intra-node and inter-node clusters." In Table 2, Lightmatter Passage M1000's 114 Tbps / 4,000 mm² is the first industrial implementation of the concept.

Industry significance: this figure is the physical basis of memory disaggregation, and the reason SK hynix people appear on the author list. If the memory pool can be separated from the XPU pool by optics, HBM's physical constraint of "having to sit right next to the die" loosens—whether that is a threat or an opportunity for HBM depends on who builds that interposer first. The logic behind Marvell's acquisition of Celestial AI is in this figure too.


Figure 4c: a multi-chip module cross-section—this figure is the full spec sheet

What the figure shows: a complete cross-section of a multi-chip module: at the top on both sides are Tx/Rx optical fibres; side by side in the middle are chiplets (ASIC / CPU / HPU / HBM); the layer below is a photonic interposer with embedded waveguides, pierced by TSVs, with the paths of electrical signals marked.

Key numbers: this panel integrates all the earlier numbers—TSV diameter 1–10 μm, hybrid-bonding pitch down to about 400 nm, and 3D photonic–electronic stack bandwidth density >5 Tbps/mm².

Industry significance: this is the closest thing to a "product cross-section" in the whole paper. It combines three supply chains that are currently separate: advanced packaging (TSV / hybrid bonding), silicon photonics (embedded waveguides) and fibre assembly (FAU). Whoever can deliver all three at once gets pricing power over this module—which is why CoWoS and CPO capacity are becoming ever harder to discuss separately.


Figure 4d: optical TSVs—sending light vertically upstairs

What the figure shows: a close-up cross-section of an optical TSV: a photonic membrane stacked on a photonic interposer, connected by a vertical optical via, with a curved optical path (a 90-degree waveguide turn) redirecting horizontally travelling light vertically.

Key numbers: the paper is direct—optical TSVs enable low-loss optical communication between stacked layers of photonic devices, overcoming the bandwidth and thermal limits of conventional electrical TSVs. Transferable membrane thickness ranges from hundreds of nanometres to several micrometres.

Industry significance: this is the panel that is "still far off but critical." In today's 3D stacks, electrical signals can travel vertically (TSVs), but optical signals cannot—light only runs within a horizontal layer. Once optical TSVs work, 3D stacking goes from "one optical layer plus many electrical layers" to "many optical layers", which is the prerequisite for the volumetric waveguide network in Figure 4e. Note that the paper also says scaling still requires solving optical alignment, loss control, thermal management and manufacturing yield.


Figure 4e: 3D waveguides—volumetric optical routing

What the figure shows: a cube structure: multiple layers of 3D waveguides stacked and embedded inside a photonic interposer, forming a volumetric waveguide network.

Key numbers: the paper describes this panel as an ultra-dense routing topology that supports thousands of wavelength channels with minimal crosstalk and loss, enabling densely wavelength-multiplexed optical channels and scalable spatial routing.

Industry significance: this figure is an "optical version of multi-layer metal routing (BEOL)." Today's PICs are planar, and routing density is limited by planar area; once waveguides can be stacked in layers, photonic-chip routing density will replicate the semiconductor BEOL growth curve. For materials and equipment makers, this is the farthest-off but largest piece of the pie—it needs low-loss, multi-layer-stackable waveguide materials compatible with the CMOS back-end thermal budget, with SiN and polymer waveguides both on the shortlist.


Figure 4f: micro-LED arrays plus photodetector arrays—the route that needs no SerDes

What the figure shows: massively parallel optical interconnect: on the left a micro-LED array, on the right photodetector arrays, each fronted by a lens array, with light coupled across free space through the lenses.

Key numbers: this panel maps to the hardest set of numbers in the text—an 800 Gbps micro-LED transceiver is projected to consume 3.1–5.3 W, versus 9.8–12 W for a mainstream 800 Gbps laser-based AOC, a 56–68% reduction; with denser arrays and 4–8 Gbps per channel it could scale further to 1.6–3.2 Tbps. Measured results in Table 2: Avicena LightBundle at 304 channels × 3.3 Gbps ≈ 1 Tbps, 1 pJ/bit, and Microsoft MOSAIC with a 100-channel prototype error-free over 20 m at 2 Gbps per channel, projected to scale to 460 channels and 800 Gbps.

Industry significance: this is the most counterintuitive panel in the review, and the one most worth adding to a watchlist. The micro-LED logic is not "light is faster," but "so many channels that you don't need SerDes"—each channel runs only 2–8 Gbps, slow enough to need no high-speed serialization, no DSP and no FEC, so the stage in Figure 2b that eats 1–2.3 pJ/bit disappears entirely. The paper's own conclusion: such systems could ultimately reach multi-terabit chip-to-chip bandwidth and femtojoule-per-bit efficiency, performance metrics unattainable with conventional architectures. The cost is reach—but at CPO scale, reach was never the problem.


8. Three walls not yet cleared: thermal, yield and standards

The paper's Challenges section is more candid than most reviews and deserves to be pulled out on its own.

Wall one: a positive thermal–optical feedback loop. Semiconductor laser emission drifts by tens of picometres per K, and silicon microring resonances by tens of GHz per K. CPO puts the laser and ASIC in the same thermal envelope, so local temperature swings reduce quantum efficiency and push wavelength drift, forming a feedback loop that degrades both optical power and spectral coherence. Worse, DSP, FEC and thermal tuning are themselves power sinks—more power means more heat, more heat needs more active cooling, and cooling consumes still more power. The paper calls this an adverse thermal–power feedback loop and notes it makes system-level efficiency degrade "faster than the link itself" as link bandwidth scales.

Wall two: optical gain and packaging yield. The main obstacle is integrating III-V gain media such as lasers and semiconductor optical amplifiers onto silicon platforms. External laser modules remain the pragmatic solution but are limited by size, alignment tolerance and high coupling loss. Wafer/die-scale bonding is the most mature integration path, but yield loss, defect density and process complexity keep hindering scale-up. Micro-transfer printing is a complementary solution; the paper cites the INSPIRE project's InP-on-SiN as an example and notes that InGaAs/InP photodetectors have been successfully transfer-printed onto SiN platforms. On the packaging side, the sticking point is that sub-micrometre alignment accuracy must be maintained in high-throughput assembly, plus delamination and long-term performance drift from CTE mismatch.

Wall three: ecosystem fragmentation. This is the section I think is most underestimated. The paper points out that photonic foundry platforms differ widely in material sets, component libraries, process parameters and packaging compatibility, and that photonic process rules often deviate from established CMOS manufacturing standards, inhibiting electro-optical co-integration. Whereas the EIC side has unified design flows, mature EDA and long-term roadmaps, the photonics side's "lack of convergence on wafer-level packaging standards, electro-optical EDA and reliability-qualification protocols" is explicitly named by the paper as a key delaying factor for CPO's entry into HPC infrastructure.

Incidentally, the paper presents glass interposers as a strong alternative for 2.5D, with concrete numbers: propagation loss below 0.05 dB/cm, an order of magnitude lower than silicon; TGV diameters below 30 μm, aspect ratios above 5:1 and pitch approaching 100 μm; stepper lithography on 300 mm glass wafers has demonstrated about 0.5 μm feature resolution; and validation in a 102.4 Tbps data-centre switch already exists. But three quantitative hurdles remain: CTE mismatch (ordinary soda-lime glass about 9 ppm/K vs silicon about 2.6 ppm/K; only display-grade borosilicate glass at 3.2–3.6 ppm/K improves die-to-substrate shear stress by more than 3×), TGV mechanical stability and fabrication, and sub-micrometre patterning over large areas.

9. Industry links: three practical implications of this review for the supply chain

First, advanced packaging is CPO's throttle valve. Every tier on the curve in Figures 3e–g, from 10–20 pJ/bit down to below 100 fJ/bit, is named after a packaging technology: wire bond, micro-bump, TSV, hybrid bonding, monolithic. This means the pace of CPO generations is capped by the volume-production maturity of packaging technology, not by optical components. Companies working on hybrid bonding, TGV and FOWLP are effectively building CPO's pace controllers.

Second, coupling and FAU are the shortest stave in the barrel. The three coupler figures in Table 1 (<1 dB / 3–6 dB / 1–2 dB), together with the text's "1–3 dB per interface" and "external lasers pay an extra 1–5 dB of link budget," pinpoint CPO's yield problem very clearly. The competitor list in this segment is relatively short, technical barriers are high, and ties to system makers are deep—it is one of the few links where Taiwan's supply chain can win pricing power.

Third, micro-LEDs and photonic interposers are two long-term tickets for the watchlist. The former (Figure 4f plus Avicena/Microsoft in Table 2) attacks SerDes, the efficiency poison, with 1 pJ/bit already measured; the latter (Figures 4b/4c) attacks the physical disaggregation of memory and compute, with Lightmatter Passage M1000's 114 Tbps / 4,000 mm² as the first industrial-grade sample. Neither route is in the mainstream view yet, but the paper uses its entire outlook section to vouch for them.

For a complete map of the optical-communications and CPO supply chain, read this alongside 2026 AI Infrastructure Must-Read: The Full Optical Communications and CPO Supply Chain Map—almost every box in this review has a matching supplier on that map.

10. Conclusion

The review's real contribution is not declaring that CPO will win—the industry voted on that with orders long ago. Its contribution is turning CPO from a vague architecture label into a technology ladder that can be quantified, ranked and priced:

10–20 pJ/bit (2D wire bonding) → 2–5 pJ/bit (2.5D micro-bumps) → 0.5–2 pJ/bit (3D TSV) → 0.1–0.5 pJ/bit (3D hybrid bonding) → below 100 fJ/bit (monolithic integration)

Industry today stands between the second and third rungs: Broadcom's Bailly and Davisson are 2.5D, NVIDIA's Quantum-X Photonics is 3D, and measured efficiency is still at 5–7 pJ/bit—still a full order of magnitude, and three walls, away from the paper's sub-pJ/bit target.

So if you take away just one line: the CPO race is no longer about "can the optics be built," but "can packaging reach volume at that tier." The three radar charts in Figure 3h say it most plainly—2D, 2.5D, 3D and monolithic will not replace one another; they will coexist for a long time, each serving a different cost/performance band. What really deserves tracking is not who announces CPO first, but who first gets hybrid-bonding yield to shippable levels, because that is the true phase-change point on the pJ/bit curve.

As for the two long-term tickets the mainstream has yet to price—micro-LED SerDes-free architectures and photonic-interposer memory disaggregation—they look distant now, but the six panels of Figure 4 tell you: the people writing the outlook have already placed their bets there.

References

  1. Kim, B., Choi, S. H., Zograf, G., Yoo, Y. J., Baek, Y., Kim, S., Kim, J., Kim, J., Chae, S. H., Kim, H., Hong, S. & Lee, K. Co-packaged optics for high-performance computing and artificial intelligence. Nature Electronics 9, 853–867 (2026). DOI: 10.1038/s41928-026-01681-6 (Review article; received 18 November 2025, accepted 8 July 2026, published online 19 August 2026)

  2. Affiliations: University of Virginia, Yonsei University, University of Illinois Urbana-Champaign, Nanyang Technological University Singapore, MIT Research Laboratory of Electronics, SK hynix Inc. (AI Infra Optimization Team)

  3. All component and platform data cited here come from the paper's Figures 1–4, Table 1 (p. 859) and Table 2 (p. 862); in Table 2, Broadcom Bailly (Tomahawk 5, March 2024) and Davisson (Tomahawk 6, October 2025) figures come from company product announcements, while other platform data follow the reference numbers in the original table.

  4. Image copyright: panels in Figure 3e are reproduced under CC BY 4.0 and CC BY NC ND licences respectively; panels in Figures 3f and 3g are adapted with permission from IEEE and Springer Nature. Please follow the original licences when reusing.

Related reading

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page