top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

NVIDIA at OFC 2026: Slide-by-Slide Breakdown of Revolutionizing Networking for AI Factories

2 days ago
4 min read

In her OFC 2026 plenary talk, "Revolutionizing Networking for AI Factories," NVIDIA Vice President Dr. Alexis Björlin redefined the role of optical communications in the AI era. Björlin stated plainly that the data center is the computer, and the network is the core that defines that computer. As AI enters the era of "agentic scaling," low latency and high throughput in infrastructure have become the key metrics for turning intelligence into economic value.

1. Vision and Industrial Foundation

  • AI factories as infrastructure: Laid out the core business logic: inference is the workload, tokens are the new commodity, and compute is revenue.


  • A multi-layer infrastructure stack: Showed the vertical stack from energy, chips and infrastructure up to the model and application layers. AI is no longer just chips, but a vast system spanning digital biology, robotics and enterprise agents.

2. The Extreme Demands of Four Scaling Laws

  • How compute demand is evolving:

    • Pre-training scaling: Model size grows by 10x in parameters per year.


    • Post-training scaling: Intelligence improves through fine-tuning and quantization (e.g., NVFP4).


    • Test-time scaling: Models "think longer" and generate more reasoning tokens (e.g., DeepSeek R1).


    • Agentic scaling: A new era of "AI talking to AI," requiring low latency, large context windows and massive memory.


  • Open models are catching up: The gap between open models and leading labs is narrowing (e.g., NVIDIA's NeMoTron-3 Super); 9 of the top 13 models are now open.

3. The Economic Cost of Training Reliability

  • Reliability is a lifeline: In Meta's Llama 3 405B training, 16,000 GPUs running for 54 days experienced 466 job interruptions, 78% of them attributed to hardware. Even a single degraded link drags down whole-system throughput.


  • Quantifying the cost of downtime:

    • Operating cost is about US$3.4 per GPU per hour.


    • For a 512,000-GPU cluster, 10% downtime means a loss of US$4.1 million per day.

4. The Revolutionary Shift in Inference Workloads

  • From node scale to data center scale: In 2024 inference was considered a single-node task; by 2026 it has become a full data-center-scale workload, with compute demand 10,000x higher than in the ChatGPT era.

  • The impact of Mixture of Experts (MoE): MoE spreads experts across many nodes (e.g., DeepSeek R1 has 256 experts per layer), generating frequent all-to-all communication and placing unprecedented demands on network bandwidth and latency.

  • Disaggregated inference: Splits inference into Prefill (compute-bound) and Decode (memory-bound), maximizing throughput by dynamically allocating compute pools.

5. Extreme Co-Design and Multiplied Performance (Pages 13-15)

  • A revolution in token cost: Through co-design, the Blackwell platform delivers 50x better performance per watt than competitors and cuts token cost by 35x.

  • Hybrid architecture: Showcased NVIDIA Dynamo, pairing processors (Vera Rubin NVL72) with low-latency processors (Groq 3 LPX), efficiently interconnected through Spectrum-6 switches.

  • Economic payoff: The Rubin + LPX architecture is expected to open a revenue opportunity worth US$300 billion.

6. The Rise of Reasoning AI and Agents (Pages 16-18)

  • A shift in inference patterns: Reasoning models require multi-turn thinking, driving 5x annual growth in token generation.


  • The agent explosion: GitHub stars for agent projects such as OpenClaw are growing faster than Linux. From "chatting" to "thinking" to "acting," each inflection point adds 100x in new compute demand.

7. The Role of Optics in Scale-Up

  • Full-stack scaling of AI factories: Showed scaling across everything from chips and system software to models and context storage.


  • NVL576 and optical links:

    • Next-generation systems will make heavy use of optics for scale-up (rack-to-rack interconnect).


    • A 512,000-GPU cluster needs more than 1.2 million optical transceivers, and the optics alone draw 30MW (about 7% of total cluster power).


  • Compressed volume ramps: 1.6T modules will ramp to volume far faster than previous generations.

8. NVIDIA Photonics: The Silicon Photonics and CPO Core

  • Integrated silicon photonics platform:

    • Showcased a 1.6T silicon photonics CPO chip built on TSMC's CoWoS 3D stacking process.


    • It includes micro-ring modulators (MRM), an external laser source (ELS) and detachable fiber connectors.

  • Advantages of micro-ring modulators (MRM):

    • MRMs are extremely small (micron scale) and extremely low power.


    • Achieved a bit error rate (BER) below 1E-10 at 212.5 Gbps per lane.


    • Thermal stability: Overcame the MRM's temperature-sensitivity pain point, staying locked through a 50°C temperature swing.




    The next frontier: DWDM: NVIDIA sees dense wavelength division multiplexing (DWDM) as the next scaling frontier. Using micro-ring resonators across 8 wavelengths (50G each), it achieves 400G per fiber at ultra-high density, with a target efficiency of 3 pJ/bit.


    Platform roadmap: From Blackwell in 2024 to Rubin in 2026 and Feynman in 2028, each generation doubles CPO bandwidth.



NVIDIA Is Rewriting the Rules Around Light

  1. Optics moves inside the rack: NVIDIA's NVL576 system shows that optics is no longer just about "kilometers of distance" but about "centimeters of connection." When even intra-rack interconnect needs CPO, optical vendors shift from "component suppliers" to "the core of system integration."

  2. MRM becomes the key variable: NVIDIA is betting on micro-ring modulators (MRM) rather than Mach-Zehnder modulators (MZM) for their space and energy efficiency. If NVIDIA's thermal-stabilization algorithms are validated at scale, it will reshape the competition among silicon photonics technology paths.

  3. Inference networking as a revenue driver: The market once assumed inference needed fewer modules, but by disaggregating prefill and decode, NVIDIA has created a large volume of high-bandwidth connectivity demand. That means the "inference dividend" may drive the optics industry longer than the "training dividend."


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page