How NVIDIA Achieves High-Compute Chips
Updated: 21 hours ago
NVIDIA's latest Blackwell compute chip reaches 20,000 TFLOPS. Had chip compute simply followed Moore's Law, it would have fallen far short of the high-performance computing needed for today's AI era.

CoWoS Breaks Through Moore's Law
Today, the way past Moore's Law is advanced packaging: tightly packaging chips together to shorten the signal path.
This advanced packaging is TSMC's CoWoS (Chip-on-Wafer-on-Substrate).

CoWoS overcomes these limits through advanced packaging. Its main features and advantages include:
1.Heterogeneous integration: CoWoS allows chips from different processes and technology nodes to be integrated together — for example, combining high-performance logic chips with high-density memory chips to form a system-in-package (SiP). This heterogeneous integration raises system performance and lowers power consumption.
2.Higher I/O density and bandwidth: CoWoS uses TSV (Through-Silicon Via) technology to dramatically increase I/O density and data bandwidth. This helps reduce latency and raise data transfer rates, further boosting overall system performance.
3.Shorter interconnect distance: In conventional packaging, chip-to-chip interconnect usually runs through the PCB (printed circuit board), which means longer distances and higher latency. CoWoS packages chips directly on the substrate, greatly shortening interconnect distance and thereby reducing signal latency and power consumption.
4.Better thermal performance: Advanced packaging usually comes with better thermal management. By packaging chips on the same substrate, CoWoS can manage heat more effectively, maintaining system stability and performance.
5.Flexible design and rapid iteration: CoWoS makes chip design more flexible, letting designers iterate and optimize quickly as needs change, shortening time-to-market.
Blackwell B200 Is the Best Example
The Blackwell-architecture GPU, NVIDIA B200, is built on TSMC's 4NP process node and uses multi-die packaging to combine two Blackwell GPU dies into a single giant chip. Each GPU die contains 104 billion transistors, so one NVIDIA B200 chip contains 208 billion transistors.
As a result, NVIDIA doesn't have to wait for a more advanced process node — with advanced packaging, it can quickly launch the B200 with higher compute.


Silicon Photonics and CPO Are the Next Step
News about TSMC's silicon photonics and CPO has been growing throughout 2024. The goal is to further solve the problem of insufficient I/O bandwidth.
NVIDIA has also announced a number of Optical I/O technologies. Silicon photonics and CPO look set to be key factors in pushing compute even higher.





Comments