NVIDIA Presentation at OFC2025: Large Scale AI Systems With Photonic Connectivity
Updated: 20 hours ago
Content:
Large scale AI systems are highly optimized and specialized for characteristics of their workloads. The linkage from the AI workloads down to photonic interconnect is not readily apparent but is currently summarized as "use copper where you can, optical where you must". We explore how other attributes such as latency, channel error rates, packaging and power might cause increased photonic adoption.
Presenter: Larry Dennison, NVIDIA Corp.
🧠 How AI Compute and Its Requirements Are Evolving
Large foundation models: models such as GPT++ and O1-class models need not only massive compute to train, but also high-performance inference architectures.
Test-time compute: the inference stage needs stronger parallelism, scaling from 8-way tensor parallelism up to 576 GPUs, sharply raising demands on network latency and bandwidth.
Inference is getting more complex: LLM inference has grown from a few GPUs to tens or even hundreds of GPUs in parallel, so inference performance is no longer a single-card problem.
🔗 The Role of Photonic Connectivity
Why silicon photonics:
About 3.5x better power efficiency
About 10x better network resiliency
Lower latency and lower bit error rate (BER)
Key technology advances:
Developing a 1.6T optical engine
Using stacked PICs (photonic integrated circuits) and a complete packaging solution (including lasers, fiber interfaces, etc.)
To be used first in NVIDIA's switches
🕸 System Architecture Challenges and Adjustments
Communication latency becomes a key bottleneck: especially in high-dimensional GPU partitioning strategies such as 3D slicing, communication latency can no longer be fully hidden.
Switching rate must increase: sacrificing some per-port bandwidth to reduce latency is a necessary trade-off.
I/O density challenge: seeking higher packaging integration and die-to-die connectivity solutions.
Limits of electrical signaling: copper is still the mainstay today, but photonic interconnect has become an indispensable option.
📦 System Packaging and Deployment Strategy
CPO (Co-Packaged Optics): considering co-packaging the network interface with the GPU to cut latency and power, though reliability and flexibility still need to be evaluated.
Storage architecture evolution: because inference needs persistent context, technologies such as customized key-value stores and persistent KV cache are being introduced.
Security: traditional encryption and authentication mechanisms add latency and power; photonic interconnect systems need redesigned lightweight authentication schemes.
🛠 Open Technology Exploration and Experiments
Evaluating a range of photonic technologies, such as:
Single-mode fiber
DWDM/CWDM
Multimode fiber + VCSEL
Even considering micro/nano-photonics technologies
Emphasizing that "all technology options are on the table" to ensure scalability and reliability.
🔚 Conclusions and Trend Outlook
Photonic interconnect is happening: NVIDIA stresses that it is not only implementing it, but also driving the whole ecosystem.
Scaling is the core problem: efficiently integrating tens of times more GPUs into a single system is the key to future AI supercomputing.
Electrical architectures will struggle to carry future demand; adopting photonics is a key step toward AI at scale.






























Comments