CES 2026: Specs of Every Key Chip and Technology in NVIDIA's Vera Rubin Architecture
Based on NVIDIA CEO Jensen Huang's CES 2026 keynote and related technical details, here is a detailed rundown of the specs for each key chip and technology in the Vera Rubin architecture.
It is the product of "Extreme Co-design." To break through the physical limits of Moore's Law, NVIDIA chose to redesign all 6 key chips in the system at once.
1. Vera CPU: a super core built for AI collaboration
This CPU was built to pair with a powerful GPU. It is no longer a general-purpose ARM core but a highly customized one.
Core architecture: 88 physical cores (high-performance cores).
Threads: supports 176 threads (SMT, Simultaneous Multi-Threading), designed so each thread gets close to physical-core performance.
Performance gains: versus the previous-generation Grace CPU, 2x the performance, and 2x the performance per watt.
Role: dedicated to data preprocessing and orchestration in AI systems, removing the bottleneck of traditional CPUs under AI workloads.

2. Rubin GPU: an AI engine that breaks physical limits
The heart of the whole architecture, packaging two dies together.
Compute performance:
AI training performance: 3.5x Blackwell.
AI inference performance: 5x Blackwell.
Transistor count: only 1.6x Blackwell (a 5x performance gain from architectural innovation with just 60% more transistors).
Key technology: NVFP4 Tensor Core
Not just simple 4-bit: traditional FP4 precision is too low and prone to distortion. NVFP4 is a complete processing unit.
Adaptive Precision: it dynamically adjusts precision and structure (block-wise scaling) based on what each model layer needs. It runs ultra-low-bit (4-bit) for very high throughput where high precision isn't needed and switches back to high precision when it matters, delivering both speed and accuracy.
Manufacturing: uses TSMC's customized CoWoS-L packaging and an advanced process node (e.g., N3P).

3. NVLink 6 Switch: a data highway with more bandwidth than the internet
The switch is the key to making 72 GPUs operate like "one giant GPU."
Per-chip specs: a single switch chip has 400 Gbps ultra-high-speed SerDes, twice today's mainstream speed.
Rack-level bandwidth: the NVL72 rack backplane delivers up to 240 TB/s of cross-sectional bandwidth.
For comparison: that's more than 2x total global internet traffic (about 100 TB/s).
Purpose: ensures every GPU can talk to all the other 72 GPUs simultaneously at full speed, non-blocking.

4. ConnectX-9 SuperNIC: ultra-fast network interface
This is the NIC that connects the rack to the outside world, optimized for AI east-west traffic.
Throughput: up to 1.6 Tbps (terabits per second).
Co-design: deeply integrated with the Vera CPU to handle RDMA (remote direct memory access) and data-path acceleration for AI workloads, freeing up CPU resources.

5. BlueField-4 DPU and the "infinite memory" architecture
One of the biggest architectural changes in the keynote, aimed at the KV cache bottleneck created by long context.
Specs: integrates Grace-class CPU cores and ConnectX-9 IP, with 800 Gbps - 1.6 Tbps of throughput.
Breakthrough feature: Inference Context Memory Storage
Background: the longer an AI conversation runs, the larger the KV cache (context memory) grows, until it no longer fits in GPU HBM.
Solution: a 150 TB shared memory pool managed by BlueField-4 is deployed inside the rack.
Result: through BlueField-4's high-speed data movement, each Rubin GPU gets an extra 16 TB of ultra-fast memory on top of its own HBM, letting AI "remember" a lifetime of conversations.

6. Spectrum-X Ethernet Switch (Silicon Photonics)
The world's first AI Ethernet switch using COUPE (Co-packaged Optics) silicon photonics technology.
Specs: a single chip provides 512 ports at 200 Gbps.
Technical highlights:
Lasers come in: fiber no longer plugs into the switch front panel but connects directly to the chip package.
Advantages: sharply reduces signal transmission power and latency, ideal for hyperscale AI factories connecting thousands of racks.




Summary: Vera Rubin NVL72 rack specs
Integrating all of the chips above into one rack gives you Vera Rubin NVL72:
GPU: 72 Rubin GPUs.
CPU: 36 Vera CPUs.
Cooling: 100% liquid cooling with a 45°C inlet temperature (warm-water cooling), eliminating power-hungry chillers.
Weight: about 2.5 tons (including coolant).
Assembly: zero-cable blind-mate design, cutting assembly time from 2 hours to 5 minutes.
These specs show NVIDIA trying to force Moore's Law of compute performance to continue through full-stack innovation.






Comments