top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

VLSI 2026 | Technical Analysis | When Packaging Becomes AI's Main Battleground: Decoding TSMC's 3.5D System Integration Blueprint (Figures 1–19 Walkthrough)

2 days ago
13 min read

The performance bottleneck of AI systems has long since moved beyond the transistors of any single chip to data movement—transferring data across chips and across packages can consume orders of magnitude more power than chiplet-to-chiplet communication inside a package. In this TSMC paper at the 2026 IEEE VLSI Technology and Circuits symposium, presented by Lee-Chung Lu, the answer points in one direction: move from 2D monolithic SoCs to 2.5D/3.5D heterogeneous integration, make the package itself the vehicle for AI scaling, and use the 3DFabric platform (SoIC, CoWoS, COUPE, SoW) to cover all four axes at once—compute, bandwidth, power delivery and thermals. The paper's numbers are hard: transistor count per package grows more than 48x from 2024 to 2029, HBM bandwidth grows 34x, 3D stacking at a 4.5µm bond pitch lifts the bandwidth-density-to-energy ratio by 10.8x, the COUPE optical engine breaks 200 Gb/s, and automated substrate routing is two orders of magnitude faster than human experts.

1. Paper Background: Not a Process Showcase, but TSMC Setting the Tone on “System Integration”

First, what this paper is not. It is not a process paper pitting N2 against A14 on who has the smaller transistor, nor a product launch betting on a single packaging technology such as CoWoS or SoIC. It is an overview paper presented by TSMC at the 2026 IEEE Symposium on VLSI Technology and Circuits, authored by Lee-Chung Lu, that sets the tone for how next-generation AI systems will scale. When TSMC declares that “advanced packaging is the pivotal technology for AI scaling,” the statement carries a weight no design company's roadmap can match—it determines how much NVIDIA, AMD, Broadcom and Google can pack into a single package over the next three to five years. The paper identifies three drivers behind AI's performance surge: advances in semiconductor technology, thermal/power/bandwidth optimization, and innovation in 3DIC design methodology together with ecosystem collaboration.

2. The Core Problem: Data Movement Is AI's Real Electricity Bill

In one sentence, the problem this paper tackles is this: as models grow, system power and latency are no longer set by computing but by moving data. Many stages of LLM training and inference are inherently data-intensive, shuttling massive model parameters back and forth between high-bandwidth memory (HBM) and the system-on-chip (SoC). Moving data between chips costs far more power, silicon area and latency than moving it on-chip; once data has to cross package boundaries, the power can be orders of magnitude higher than chiplet communication within a package. The solution follows naturally: make the package bigger and fit more HBM and chiplets into one integrated package so that data rarely has to leave it. That in turn makes the move to 3D chip stacking inevitable—at the cost of higher power density and thermal resistance from vertical stacking, which makes thermal-aware design a non-negotiable prerequisite.

3. Architecture Overview: The Three Pillars of the 3DFabric Platform (Figures 1–3)

Figure 1 shows two parallel tracks of semiconductor integration: the traditional monolithic path that keeps scaling along Moore's Law, and the 2.5D/3D heterogeneous integration path. A single die is limited by the reticle size, putting a physical ceiling on its transistor count; heterogeneous integration combines multiple dies to break through that limit. With heterogeneous integration, the total transistor count of a single system can be pushed beyond one trillion. This figure is the worldview of the entire paper: it declares that the continuation of Moore's Law has shifted from making transistors smaller to stacking and tiling dies together.

Parallel tracks of monolithic and heterogeneous integration | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 1
Parallel tracks of monolithic and heterogeneous integration | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 1

Figure 2 lays out the full technology map of TSMC 3DFabric—an overview of TSMC's heterogeneous integration toolbox. 3DFabric consists of three technology categories: first, advanced silicon process; second, SoIC 3D silicon stacking, which stacks dies vertically for ultra-high density and extremely short interconnects; third, advanced packaging solutions, including the 2.5D interposer-based InFO and CoWoS as well as SoW, which integrates the system directly on the wafer. These can be combined—SoIC stacks upward, CoWoS tiles sideways, and SoW turns an entire wafer into a system. This also explains why packaging capacity (especially CoWoS) has become the choke point of the entire AI supply chain.

Figure 2: TSMC 3DFabric (technology map of advanced silicon process / SoIC 3D stacking / CoWoS·InFO·SoW advanced packaging) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 2
Figure 2: TSMC 3DFabric (technology map of advanced silicon process / SoIC 3D stacking / CoWoS·InFO·SoW advanced packaging) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 2

Figure 3 shows how an HPC platform heterogeneously integrates a variety of 3DIC components, grounding Figure 2's abstract map in a real cross-section. Multiple components sit on the same CoWoS RDL interposer, including HBM, logic dies and various embedded components. Details worth flagging: the HBM base die has moved from a conventional DRAM process to TSMC's advanced logic process; LSI embedded in the interposer serves as a high-speed channel between chiplets; embedded IVR and enhanced eDTC improve power delivery efficiency; COUPE is integrated to combine electronic and photonic integrated circuits; and SoIC handles vertical stacking of the top and bottom dies. In short, it makes plain that an advanced package is really a miniature system.

HPC platform with heterogeneous integration | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 3
HPC platform with heterogeneous integration | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 3

4. Compute Scaling: CoWoS Tiles Sideways, SoIC Stacks Upward (Figures 4–7)

Figure 4 shows the compute scaling path of the CoWoS platform as the foundation of AI system integration. CoWoS is the vehicle for ever-larger packages, using reticle size as its unit of scaling. CoWoS has grown from 3.3x reticle size to 5.5x, with the roadmap heading to 9.5x, 14x and beyond. Every step up in reticle size means another step up in the logic and memory area a single package can hold.

Figure 4: Compute Scaling with the CoWoS platform | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 4
Figure 4: Compute Scaling with the CoWoS platform | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 4

Figure 5 shows SoIC, the “stack upward” route to compute scaling. SoIC enables vertical chiplet stacking through ultra-fine-pitch, direct die-to-die connections. Vertical stacking effectively multiplies the amount of compute logic within the same package footprint, sharply raising compute density per unit area; placing more processing units closer together shortens the physical distance data travels, cutting both latency and power. CoWoS solves “not enough area”; SoIC solves “what to do once the area runs out.”

Figure 5: Compute Scaling with the SoIC platform | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 5
Figure 5: Compute Scaling with the SoIC platform | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 5

Figure 6 shows the overall compute scaling trend—quantifying the combined effect of the two routes in Figures 4 and 5. From 2024 to 2029, compute per package is projected to grow more than 48x. Three things drive this together: process advancing from N7 to A14, the introduction of SoIC, and CoWoS scaling from 3.3x to over 14x reticle size, raising the number of SoCs integrated per package from 2 to 24. 48x is a number that will rewrite data center design, and it explains why power delivery and thermals have become hard constraints that must be solved in lockstep.

Figure 6: Compute scaling trend (N7→A14 + SoIC + CoWoS 14x) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 6
Figure 6: Compute scaling trend (N7→A14 + SoIC + CoWoS 14x) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 6

Figure 7 shows the HBM bandwidth scaling trend, the memory-side counterpart to compute scaling. From 2024 to 2029, HBM bandwidth is projected to grow 34x. Drivers include the standard evolving from HBM3 to HBM5E, I/O count per HBM rising from 1024 to 2048, a 5.7x increase in bitrate per I/O, and the number of integrated HBM stacks growing from 8 to 24; meanwhile, the HBM logic base die moves from a DRAM process to N3P. The ratio of 34x bandwidth to 48x compute is itself a signal—memory bandwidth grows slightly slower than compute, so pressure from the “memory wall” persists.

Figure 7: HBM bandwidth scaling trend | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 7
Figure 7: HBM bandwidth scaling trend | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 7

5. Bandwidth Scaling: A Three-Tier Architecture of Scale-in, Scale-up and Scale-out (Figures 8–11)

The paper breaks bandwidth scaling into three tiers: scale-in (between chiplets, raising bandwidth density and lowering latency), scale-up (within a single multi-chiplet package), and scale-out (high-bandwidth, energy-efficient transmission across packages, systems and racks). TSMC maps SoIC to scale-in, CoWoS to scale-up and COUPE to scale-out. Figure 8 quantifies the gains of scale-in and scale-up bandwidth scaling. In 2.5D integration, shrinking the uBump pitch from 45µm to 35µm and moving the process from N3P to A14 and beyond raises bandwidth density 1.3x and cuts energy to 0.7x, a 1.8x improvement in the bandwidth-density-to-energy ratio. In 3D chip stacking, a bond pitch as fine as 4.5µm—versus a 9µm/N7 baseline—raises bandwidth density 4x and cuts energy to 0.37x, lifting the bandwidth-density-to-energy ratio by a striking 10.8x.

Scale-in and scale-up for AI bandwidth scaling | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 8
Scale-in and scale-up for AI bandwidth scaling | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 8

Figure 9 shows UCIe bandwidth scaling and signal integrity validation results on the CoWoS platform. TSMC has validated 32Gb/s UCIe performance on CoWoS, with the RDL interposer delivering excellent power and latency, and eDTC further strengthening power integrity. It also studied UCIe 3.0 signal integrity at 64Gb/s: with a 45µm bump pitch and an effective shielding strategy, the eye diagram remains robust at 64Gb/s; and the 35µm bump pitch option—with shorter traces, smaller IP area and better energy efficiency—was confirmed to retain sufficient 64Gbps signal integrity as well.

Figure 9: UCIe bandwidth scaling (32/64Gb/s signal integrity validation of UCIe on CoWoS) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 9
Figure 9: UCIe bandwidth scaling (32/64Gb/s signal integrity validation of UCIe on CoWoS) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 9

Figure 10 shows the scale-out tier—the evolution of optical interconnect, i.e., COUPE's path to integrating optical transceivers directly into the package. For long-reach transmission across packages and racks, electrical interconnect can no longer keep up and must give way to optics. The figure traces optical links moving from the board edge, to near the chip, to co-packaged inside the package. COUPE integrates optical transceiving directly into the package for high-bandwidth, energy-efficient optical communication; SoIC, CoWoS and COUPE work together to form a hierarchical bandwidth scaling system from scale-in to scale-out.

Figure 10: Evolution of optical connections (optical interconnect evolution and in-package COUPE integration) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 10
Figure 10: Evolution of optical connections (optical interconnect evolution and in-package COUPE integration) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 10

Figure 11 details the micro-ring modulator (MRM) technology that pushes the COUPE optical engine PIC beyond 200+Gb/s. Through process-design co-optimization, COUPE breaks the 200 Gb/s data rate. The approach optimizes the PIC's MRM junction, co-designs the EIC driver with the PIC MRM impedance, and applies inductive peaking—together delivering a 1.8x improvement in overall system bandwidth. 200Gb/s single-lane optical modulation is concrete evidence that CPO is moving from concept to a manufacturable specification.


Figure 11: COUPE PIC 200+Gb/s Micro-ring Modulator (COUPE micro-ring modulator breaks 200Gb/s) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 11
Figure 11: COUPE PIC 200+Gb/s Micro-ring Modulator (COUPE micro-ring modulator breaks 200Gb/s) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 11

6. Power Delivery Network Optimization: Current Is the Real Silent Killer (Figures 12–14)

Figure 12 shows the AI package power scaling trend—the price of 48x compute written directly on the power curve. Soaring transistor counts push power straight up; with massive parallel computing enabled by computer architecture innovation stacked on top of SoIC 3D stacking innovation, total package power grows steeply. Design-technology co-optimization (DTCO) of power delivery and thermals therefore becomes a prerequisite for sustaining AI power scaling. The slope of the power curve is approaching the ceiling of conventional power delivery and cooling solutions.

Figure 12: Package power scaling trend | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 12
Figure 12: Package power scaling trend | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 12

Figure 13 shows the structure and optimization strategy of the power delivery network (PDN), focusing on suppressing AC voltage droop at low-to-medium current densities. Energy efficiency requires lowering the supply voltage (Vdd), but higher power density combined with lower Vdd means the package must draw much larger currents. At low-to-medium current densities (up to 4 A/mm²), optimization focuses on suppressing AC droop. TSMC offers two kinds of decoupling capacitors: on-chip high-density MIM decap with a capacitance density of 500 nF/mm² to suppress high-frequency (typically >100 MHz) supply noise, and in-package eDTC with a capacitance density of up to 2500 nF/mm² to provide a low-impedance current source at mid frequencies (10–100 MHz).

Figure 13: PDN network structure and optimization (power delivery network structure and decoupling capacitor optimization) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 13
Figure 13: PDN network structure and optimization (power delivery network structure and decoupling capacitor optimization) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 13

Figure 14 shows the evolution of voltage regulator styles, tackling the high-current-density end of the problem. Once current density exceeds 4 A/mm², I²R losses and electromigration can no longer be handled by passive components; the package input voltage must be raised to cut current sharply, which calls for integrated voltage regulators (IVR). TSMC's On-Wafer Inductor (OWL) technology builds inductors directly on the wafer and pairs them with a power management IC to form a buck converter, raising the package input voltage from the core voltage (e.g., ~0.7V) to >1.8V—reducing current density by more than 2.5x and DC losses by more than 6x.

Figure 14: Evolution of voltage regulator styles (voltage regulator evolution and IVR/OWL integration) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 14
Figure 14: Evolution of voltage regulator styles (voltage regulator evolution and IVR/OWL integration) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 14

7. Thermal Design Co-Optimization: The Last Bottleneck of Vertical Stacking (Figures 15–16)

Figure 15 shows thermal DTCO solutions for high power density—how packaging and design work together on heat dissipation. SoIC doubles the transistor count and directly raises power density, so heat can only be contained by attacking it from both the package and the design side. On the package side there are two levers: lidless packaging lets heat flow directly from the carrier to the heat sink, and high-thermal-conductivity (high-kappa) carriers further improve in-package heat conduction. On the design side there are three: hotspot spreading—the most effective design rule—spreads concentrated heat over a larger area, while inserting dummy bonds and dummy vias creates additional heat paths for the top and bottom dies.

Figure 15: Thermal DTCO solutions for high power density | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 15
Figure 15: Thermal DTCO solutions for high power density | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 15

Figure 16 shows thermal optimization of the HBM base die under SoIC vertical stacking—the toughest thermal-coupling scenario in 3D stacking. When the HBM logic base die and the compute die are vertically stacked with SoIC, complex thermal coupling arises among the compute SoC, the heat sink, the DRAM stack and the logic base die. Mitigation strategies include improving the thermal resistance of the heat sink and thermal interface material (TIM), lowering the power of both the SoC and the logic base die, and introducing thermal-aware floorplanning at the design stage. The real difficulty of HBM stacking is not bandwidth—it is heat.

Figure 16: HBM base die thermal optimization | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 16
Figure 16: HBM base die thermal optimization | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 16

8. Ecosystem Collaboration and EDA Automation: 3Dblox and Agentic AI (Figures 17–19)

Figure 17 shows the architecture of the 3Dblox language—a modular, hierarchical language TSMC built for 3DIC design automation. Launched in 2022, 3Dblox has become a global standard for 3DIC design, used to describe the complexity of 3D stacking. The language modularizes 3D components into chiplets, interfaces and connections, providing a unified language for logical and physical connectivity; it supports top-down design methodology, promotes chiplet reuse and improves interoperability across EDA tools. It has been donated to the IEEE Standards Association as P3537, “3Dblox — Chiplet Connectivity and Physical Properties Description Language.” With TSMC leading the description language for 3D stacking and placing it into IEEE, it has secured a strategic position at the top layer of the ecosystem.

Figure 17: The 3Dblox Language (modular hierarchical language, now IEEE P3537) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 17
Figure 17: The 3Dblox Language (modular hierarchical language, now IEEE P3537) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 17

Figure 18 shows the results of substrate auto-routing—productivity evidence of 3Dblox plus EDA collaboration. Advanced-package substrate design is full of special challenges—high-density escape routing, fine line width and spacing, differential-pair routing, plated through-hole (PTH) planning and length matching—making manual work extremely time-consuming. TSMC and Cadence co-developed the Allegro router, achieving routing within 10% of human quality on large industrial-scale designs while cutting routing time by two orders of magnitude compared with senior human designers. This signals that substrate design, long dependent on veteran experience, is being taken over by AI-assisted auto-routing.

Figure 18: Substrate auto-routing (<10% gap to human quality, two orders of magnitude faster) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 18
Figure 18: Substrate auto-routing (<10% gap to human quality, two orders of magnitude faster) | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 18

Figure 19 shows how agentic AI is permeating every stage of the 3DIC design flow. The paper lists several concrete applications—AI for high-speed interface channel optimization, a runset coding copilot to assist script generation, an EDA knowledge base giving engineers access to accumulated design knowledge, physical design agents pursuing optimal power/performance/area (PPA), and DRC agents accelerating verification cycles. The full 3D context defined by 3Dblox also makes possible compile_bumps automatic bump assignment, 3D ESD analysis and AI-driven global resource optimization. This foreshadows that the EDA industry's next battleground lies not in algorithms but in “agentification.”

Figure 19: 3DIC design with agentic AI | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 19
Figure 19: 3DIC design with agentic AI | Image source: Advancing Package and System Integration for Next-Generation AI (VLSI 2026, TSMC) - Figure 19

9. Technical Highlights: The Two Engineering Turning Points Worth Remembering

First, the energy-efficiency dividend of 3D stacking is quantified at 10.8x (Figure 8). “Vertical stacking saves power” used to be intuition; this paper nails the 10.8x gap in bandwidth-density-to-energy ratio to a number, using a 4.5µm bond pitch versus a 9µm/N7 baseline. That is the strongest argument for pushing the entire die-to-die interface from 2.5D to 3D—not for speed, but to avoid being crushed by power bills and heat at 48x compute. Second, design automation is upgrading from tools to agents (Figures 17–19). 3Dblox provides a unified 3D description language, substrate auto-routing runs 100x faster within a <10% gap, and agentic AI takes over PPA and DRC. The bottleneck of advanced packaging is shifting from “can the process build it” to “can design keep up.”

10. Industry Implications: How Far These Technologies Are from Volume Production, and Who Benefits

The roadmap in the paper lays out a fairly clear timeline. CoWoS is already in volume production at 5.5x reticle size and moving toward 9.5x/14x; SoIC and UCIe 32Gb/s are validated, with 64Gb/s signal integrity confirmed; COUPE has broken 200Gb/s; 3Dblox is already a global standard and has entered IEEE P3537; substrate auto-routing and AI agents have delivered results on industrial-scale designs. The beneficiaries line up: AI chip designers (NVIDIA, AMD, Broadcom, Google and others) gain a vehicle for 48x compute and 34x bandwidth; the big three HBM makers (SK hynix, Samsung, Micron) are pulled along by the spec roadmap in Figure 7; the CPO/silicon photonics supply chain is driven by COUPE; EDA vendors (Cadence is named explicitly) are repositioning in the agentification wave; and the power management IC supply chain is being reshuffled as IVR/OWL move power delivery into the package.

Conclusion

The real weight of this paper lies not in any single number, but in how it ties advanced packaging, power delivery, thermals and EDA automation together as one problem. These four used to belong to different teams, different suppliers and different conference agendas; at VLSI 2026, TSMC put them on the same table, declaring that the next chapter of Moore's Law lies not in transistor dimensions but in co-design for system integration. When compute per package grows 48x in five years, the power curve steepens in step, and vertical stacking pushes heat to the limit, optimizing any one link in isolation will be dragged down by another. For the entire AI supply chain, this paper draws the engineering map for packing a trillion-plus transistors into a single system over the next five years—and at the center of that map is the package.

References

Paper title: Advancing Package and System Integration for Next-Generation AI. Author: Lee-Chung Lu (Taiwan Semiconductor Manufacturing Company, TSMC, Hsinchu, Taiwan). Conference: 2026 IEEE Symposium on VLSI Technology and Circuits, 2026. DOI: 10.1109/VLSITECHNOLOGYANDCIR65830.2026.11577449. Further references: [1] Y.-J. Mii, "Semiconductor Industry Outlook and New Technology Frontiers," 2024 IEEE IEDM; [2] IEEE Standards Association, "3Dblox — Chiplet Connectivity and Physical Properties Description Language," P3537.


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page