top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

CES 2026: Specs of Every Key Chip and Technology in NVIDIA's Vera Rubin Architecture

57 minutes ago
3 min read

Based on NVIDIA CEO Jensen Huang's CES 2026 keynote and related technical details, here is a detailed rundown of the specs for each key chip and technology in the Vera Rubin architecture.

It is the product of "Extreme Co-design." To break through the physical limits of Moore's Law, NVIDIA chose to redesign all 6 key chips in the system at once.


1. Vera CPU: a super core built for AI collaboration

This CPU was built to pair with a powerful GPU. It is no longer a general-purpose ARM core but a highly customized one.

  • Core architecture: 88 physical cores (high-performance cores).

  • Threads: supports 176 threads (SMT, Simultaneous Multi-Threading), designed so each thread gets close to physical-core performance.

  • Performance gains: versus the previous-generation Grace CPU, 2x the performance, and 2x the performance per watt.

  • Role: dedicated to data preprocessing and orchestration in AI systems, removing the bottleneck of traditional CPUs under AI workloads.


2. Rubin GPU: an AI engine that breaks physical limits

The heart of the whole architecture, packaging two dies together.

  • Compute performance:

    • AI training performance: 3.5x Blackwell.

    • AI inference performance: 5x Blackwell.

  • Transistor count: only 1.6x Blackwell (a 5x performance gain from architectural innovation with just 60% more transistors).

  • Key technology: NVFP4 Tensor Core

    • Not just simple 4-bit: traditional FP4 precision is too low and prone to distortion. NVFP4 is a complete processing unit.

    • Adaptive Precision: it dynamically adjusts precision and structure (block-wise scaling) based on what each model layer needs. It runs ultra-low-bit (4-bit) for very high throughput where high precision isn't needed and switches back to high precision when it matters, delivering both speed and accuracy.

  • Manufacturing: uses TSMC's customized CoWoS-L packaging and an advanced process node (e.g., N3P).


3. NVLink 6 Switch: a data highway with more bandwidth than the internet

The switch is the key to making 72 GPUs operate like "one giant GPU."

  • Per-chip specs: a single switch chip has 400 Gbps ultra-high-speed SerDes, twice today's mainstream speed.

  • Rack-level bandwidth: the NVL72 rack backplane delivers up to 240 TB/s of cross-sectional bandwidth.

    • For comparison: that's more than 2x total global internet traffic (about 100 TB/s).

    • Purpose: ensures every GPU can talk to all the other 72 GPUs simultaneously at full speed, non-blocking.

4. ConnectX-9 SuperNIC: ultra-fast network interface

This is the NIC that connects the rack to the outside world, optimized for AI east-west traffic.

  • Throughput: up to 1.6 Tbps (terabits per second).

  • Co-design: deeply integrated with the Vera CPU to handle RDMA (remote direct memory access) and data-path acceleration for AI workloads, freeing up CPU resources.



5. BlueField-4 DPU and the "infinite memory" architecture

One of the biggest architectural changes in the keynote, aimed at the KV cache bottleneck created by long context.

  • Specs: integrates Grace-class CPU cores and ConnectX-9 IP, with 800 Gbps - 1.6 Tbps of throughput.

  • Breakthrough feature: Inference Context Memory Storage

    • Background: the longer an AI conversation runs, the larger the KV cache (context memory) grows, until it no longer fits in GPU HBM.

    • Solution: a 150 TB shared memory pool managed by BlueField-4 is deployed inside the rack.

    • Result: through BlueField-4's high-speed data movement, each Rubin GPU gets an extra 16 TB of ultra-fast memory on top of its own HBM, letting AI "remember" a lifetime of conversations.

6. Spectrum-X Ethernet Switch (Silicon Photonics)

The world's first AI Ethernet switch using COUPE (Co-packaged Optics) silicon photonics technology.

  • Specs: a single chip provides 512 ports at 200 Gbps.

  • Technical highlights:

    • Lasers come in: fiber no longer plugs into the switch front panel but connects directly to the chip package.

    • Advantages: sharply reduces signal transmission power and latency, ideal for hyperscale AI factories connecting thousands of racks.




Summary: Vera Rubin NVL72 rack specs

Integrating all of the chips above into one rack gives you Vera Rubin NVL72:

  • GPU: 72 Rubin GPUs.

  • CPU: 36 Vera CPUs.

  • Cooling: 100% liquid cooling with a 45°C inlet temperature (warm-water cooling), eliminating power-hungry chillers.

  • Weight: about 2.5 tons (including coolant).

  • Assembly: zero-cable blind-mate design, cutting assembly time from 2 hours to 5 minutes.

These specs show NVIDIA trying to force Moore's Law of compute performance to continue through full-stack innovation.




Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page