top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

VLSI 2026 | Paper Analysis | University of Washington Pushes a Single Wavelength to 320Gb/s: Coherent QAM with MRMs Beats PAM-4 at 0.69pJ/b

2 days ago
9 min read

The University of Washington (UW) presented a monolithic silicon photonic coherent optical transmitter at VLSI 2026 (session C20.3). Instead of the bulky Mach-Zehnder interferometer (MZI) traditionally used, it performs QAM-4 coherent modulation with microring modulators (MRMs): four channels at 80Gb/s each, delivering 320Gb/s aggregate on a single wavelength. The key number is 0.69pJ/b energy efficiency, lower than other MRM CPO solutions on the same process node. Why it matters: rather than scaling bandwidth by "adding ever more laser wavelengths," it raises the information density of each wavelength with coherent QAM, directly cutting the number of laser lines and fibers the system has to support. It is built on GlobalFoundries (GF) 45nm monolithic silicon photonics, with an on-chip digital PLL (DPLL) clock and fully time-multiplexed thermal tuning control loops.

1. Background: The UW Team, VLSI 2026 C20.3, and Why It Deserves Attention

First, what this paper is not. It is not yet another PAM-4 optical engine stacking up data rates, nor the traditional route of opening more dense wavelength-division multiplexing (DWDM) channels to make up bandwidth. It does something the silicon photonics community has long considered "impossible for MRMs": direct coherent QAM with microring modulators. The paper comes from Sajjad Moazeni's group at the University of Washington (first author Pengyu Zeng), presented at the 2026 IEEE Symposium on VLSI Technology and Circuits, session C20.3. Thanks to their small footprint and high efficiency, MRMs are widely regarded as the workhorse of DWDM CPO, but DWDM comes at the cost of many independent laser lines and fibers, with system complexity and cost rising linearly with wavelength count. This paper's value is offering another scaling path that does not rely on "adding wavelengths without limit."

Figure 1: Concept of coherent CPO (C2PO) using MRM optical transmitters, and its operation principle. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 1
Figure 1: Concept of coherent CPO (C2PO) using MRM optical transmitters, and its operation principle. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 1

2. The Core Problem in One Sentence

In one sentence: rather than endlessly increasing the number of wavelengths, raise the information density of a single wavelength with coherent QAM; but MRMs inherently couple amplitude and phase, making direct QAM hard, and this paper sidesteps that coupling with a RAMZI architecture. Conventional intensity-modulation direct-detection (IM-DD) carries only one dimension of information per symbol; coherent QAM uses amplitude and phase simultaneously across separate I/Q paths, so each symbol carries far more information and spectral efficiency rises sharply. The problem is that when an MRM's resonance is tuned, its phase shifts along with its amplitude, making it hard to draw a clean QAM constellation. The paper's answer is the ring-assisted Mach-Zehnder interferometer (RAMZI).

3. System Concept and Chip Architecture: Figure 1, Figure 2

Figure 1 shows the new paradigm of coherent CPO (abbreviated C2PO in the paper) and how RAMZI turns MRMs into coherent modulators capable of QAM. Each RAMZI uses "a pair of MRMs," thermally tuning the two rings' resonances to either side of the laser wavelength; a differential drive voltage then moves the two rings in opposite directions simultaneously, producing a signal with "fixed phase and only amplitude modulated" - neatly sidestepping the MRM amplitude-phase coupling problem. Two RAMZIs (each carrying independent data) are then combined with a 90-degree phase difference to form an offset-QAM modulator. The phase offset is deliberately added and tunable, so the receiver can perform DSP-free carrier phase recovery.

Figure 2 shows the full transmitter (Tx) chip architecture: four channels share one clocking core. The left half is the shared clocking core, with the photonics highlighted in blue; the right half is the core architecture of a single transmit channel. All four channels share the same optical input to generate modulated data, and in this design all channels share a single fiber to forward the local oscillator (LO) laser to the receiver, minimizing fiber-count overhead. Four channels, a single optical input, one shared LO forwarding fiber - these "shared" choices are the core design decisions that cut system-level fiber and laser counts.


Figure 2: The transmitter chip architecture with 4 channels and a shared clocking core (left) with photonics highlighted in blue, architecture of each transmitter core/channel (right). | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 2
Figure 2: The transmitter chip architecture with 4 channels and a shared clocking core (left) with photonics highlighted in blue, architecture of each transmitter core/channel (right). | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 2

4. Clocking and High-Speed Electronic Front End: Figure 3, Figure 4

Figure 3 shows the block diagram of the clocking core. It consists of a 20GHz digital PLL (DPLL) with an LC-based digitally controlled oscillator (LC-DCO) and a clock divider chain (CLK Div Chain), referenced to an external CML clock of about 625MHz. The 20GHz differential output is distributed to each channel over coplanar waveguide (CPW) transmission lines. Each channel's digital backend runs at 1.25GHz (handling PRBS and the control loops). Integrating the DPLL and high-speed clock distribution into monolithic silicon photonics means this is a complete SoC that generates and distributes its own clocks.


Figure 3: Block-diagram of the clock core with digital PLL (DPLL), LC-DCO and Clock Divider Chain (CLK Div Chain). | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 3
Figure 3: Block-diagram of the clock core with digital PLL (DPLL), LC-DCO and Clock Divider Chain (CLK Div Chain). | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 3

Figure 4 shows the circuit block diagram of a single channel's transmitter analog front end (AFE). To reach 40 GBaud on GF 45nm, the team customized the serializer latches and tri-state multiplexers (tri-state MUX), using transmission gates to connect the NMOS/PMOS inputs to the complementary outputs of the previous stage. This structural optimization effectively cuts fanout by about 2x, lowering the equivalent capacitive load and markedly increasing the latches' data-capture speed. Within each channel, the 20GHz clock is first converted into four-phase 10GHz CMOS clocks feeding the I/Q serializers, and only the final stage steps back up to 20GHz, minimizing routing parasitics. Hitting 40 GBaud on a mature, inexpensive 45nm node comes from circuit-level ingenuity, not an advanced process.


Figure 4: Block diagram of a single channel transmitter electronics. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 4
Figure 4: Block diagram of a single channel transmitter electronics. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 4

5. Photonic Components and Thermal Tuning Control Loops: Figure 5, Figure 6

Figure 5 shows the complete photonic components of the RAMZI-based QAM transmitter. Both the MRM resonances and the RAMZI bias points are held by closed-loop thermal phase shifters; the phase shifters are controlled by programmable pulse-density modulation (PDM), with thick-oxide NMOS devices driving the heaters. Each I/Q RAMZI has three power-monitoring photodiodes (PDs) - one on the drop port of each of the two MRMs, plus one on the second optical port of the RAMZI output combiner. Integrating both the sensing and actuation of thermal tuning on a single chip is the key to whether this kind of approach can actually be deployed.

Figure 5: Detailed photonic components of the RAMZI-based QAM Tx. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 5
Figure 5: Detailed photonic components of the RAMZI-based QAM Tx. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 5

Figure 6 shows the on-chip thermal tuning circuitry - the paper's signature "time-multiplexed thermal tuning." Photocurrents from the monitor PDs are time-multiplexed into a capacitive transimpedance amplifier (CTIA) integrating front end with a large dynamic range, then digitized by a 6-bit successive-approximation (SAR) ADC. A digital loop sweeps the heater values, first finding and then locking the average optical power to a pre-calibrated target level. The 6-bit SAR ADC and the digital backend reference clock run at 20GHz/16 = 1.25GHz. Time-multiplexing multiple PD currents onto one shared CTIA+ADC is the key technique that keeps the power and area overhead of the thermal tuning loops "minimal."

Figure 6: Block-diagram of on-chip thermal tuning circuitry (CLK is the digital backend's reference clock at the 20GHz/16=1.25GHz). | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 6
Figure 6: Block-diagram of on-chip thermal tuning circuitry (CLK is the digital backend's reference clock at the 20GHz/16=1.25GHz). | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 6

6. Measurement Results: Figure 7, Figure 8, Figure 9


Figure 7: Die photo of the multi-channel optical transmitter silicon photonic chip. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 7
Figure 7: Die photo of the multi-channel optical transmitter silicon photonic chip. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 7

Figure 7 shows the die photo of the C2PO transmitter. This is the physical die of the multi-channel optical transmitter silicon photonic chip, with electronics and photonics coexisting on GF 45nm monolithic silicon photonics for minimal parasitics. A single die holds four channels, the shared clocking core and four RAMZI photonic blocks - this monolithic integration density is a structural advantage over approaches that "separate electronics and photonics, then package them together."

Figure 8: RAMZI/MRMs thermal tuning measurements: initial calibration of MRMs to find optimum locking levels (top), measured optical power by ADCs showing locking while input laser wavelength is intentionally swept (bottom). | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 8
Figure 8: RAMZI/MRMs thermal tuning measurements: initial calibration of MRMs to find optimum locking levels (top), measured optical power by ADCs showing locking while input laser wavelength is intentionally swept (bottom). | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 8

Figure 8 shows the RAMZI/MRM thermal tuning measurements. The top half is the initial MRM calibration used to find the optimum locking level; the bottom half shows the optical power measured by the ADCs staying locked while the input laser wavelength is intentionally swept. For RAMZI to work properly, the team fixed the two MRMs' resonances about 100pm apart, then tuned the RAMZI phase shifter to maximize the observed optical modulation amplitude (OMA). The measurement used a DFB laser with about 1MHz linewidth and 16.5 dBm total optical power, with an estimated -5.5 dBm reaching each MRM. This is empirical proof that "the thermal tuning really does hold lock."

Figure 9: Measured optical eye-diagrams and power/area breakdown per channel. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 9
Figure 9: Measured optical eye-diagrams and power/area breakdown per channel. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 9

Figure 9 shows the measured optical eye diagrams and the per-channel power/area breakdown. The I and Q optical eyes were measured with a commercial optical receiver and a high-bandwidth sampling oscilloscope; the right side shows each channel's electrical power and area breakdown. Each RAMZI modulator measured 5.0 dB insertion loss (IL) and 7.9 dB extinction ratio (ER), corresponding to a normalized OMA of 0.26, or a normalized electric-field OMA of 0.33 for offset-QAM. Overall power is about 0.69pJ/b @ 40 GBaud - thanks to each channel's I and Q each running NRZ, which relaxes the baud rate relative to PAM-4 and therefore saves power. At 0.69pJ/b, coherent QAM and low power - long considered contradictory - sit on the same chip.

7. Test Platform and Packaging: Figure 10


Figure 10: Test setup diagram and photos showing chip-on-board packaging and an optical V-groove array. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 10
Figure 10: Test setup diagram and photos showing chip-on-board packaging and an optical V-groove array. | Source: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning (VLSI 2026, C20.3) - Figure 10

Figure 10 shows the test setup and photos of the board-level packaging. Light is coupled on and off the chip through an 8-channel, 250um-pitch V-groove array. This matches the low-fiber-count design described earlier, where "all channels share a single optical input and one shared LO forwarding fiber." Showing a practical packaging approach like chip-on-board plus a V-groove array signals that this is not just a chip-level result, but a step toward actually going into systems.

8. Technical Highlights: The Two Genuine Innovations

First, building a RAMZI from two MRMs to sidestep MRM amplitude-phase coupling and achieve coherent offset-QAM. Coherent modulation used to be almost exclusively the domain of large-footprint MZIs, with MRMs ruled out because their phase drifts along with amplitude modulation. UW moves the two rings in opposite directions to output pure amplitude modulation with fixed phase, then combines two RAMZIs at 90 degrees - achieving QAM in a small fraction of an MZI's area, with a deliberate offset so the receiver can recover carrier phase without DSP. Second, time-multiplexed thermal tuning resolves the MRM's biggest headache - thermal stability - with minimal overhead. Currents from multiple monitor PDs are time-multiplexed onto one shared CTIA integrating front end plus a 6-bit SAR ADC, and a digital loop automatically sweeps the heaters and locks to the target optical power. This is key to holding 0.69pJ/b.

9. Industry Relevance: How Far from Volume Production, and Who Benefits

The chip runs on GF 45nm monolithic silicon photonics, a mature, production-ready platform, with chip-on-board plus standard V-groove array packaging - no exotic process involved. That means it is not far from being "reproducible by industry"; the main hurdle is whether the ecosystem is willing to adopt coherent CPO. The beneficiaries are hyperscale data centers looking to squeeze more bandwidth out of the same laser and fiber budget: fewer laser lines and fewer fibers directly cut system cost and complexity. It also offers a clear upgrade path: today it is QAM-4 (each channel's I and Q running NRZ), and scaling to QAM-16 would double the rate again without tearing up the RAMZI and thermal tuning architecture.

Conclusion

The paper's positioning is clear: it proves MRMs need not be stuck in IM-DD forever and can deliver low-power coherent QAM on a mature silicon photonics process. By sidestepping MRM amplitude-phase coupling with RAMZI and thinning out the cost of thermal stability with time-multiplexed thermal tuning, UW brings four usually conflicting goals - "coherent," "low power," "small footprint" and "manufacturable process" - together in one chip at 0.69pJ/b and 320Gb/s per wavelength. It will not replace PAM-4 or DWDM overnight, but it gives CPO an empirically demonstrated path to "scale capacity without adding wavelengths endlessly." The real next steps to watch are whether it can climb to QAM-16, and whether DSP-free coherent receivers can mature alongside it.

References

Paper title: A 0.69-pJ/b 4×80-Gb/s MRM-Based Coherent Optical Transmitter with Time-Multiplexed Thermal Tuning. Authors: Pengyu Zeng, Marziyeh Rezaei, Daniel Sturm, Asha Rashmi Nayak, Scott Li, Sajjad Moazeni (University of Washington, Seattle, WA, USA). Conference: 2026 IEEE Symposium on VLSI Technology and Circuits, session C20.3, 2026. DOI: 10.1109/VLSITECHNOLOGYANDCIR65830.2026.11577385.


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page