Lighting the Path to Exascale AI: Photonics in High-Performance Clusters (OFC 2025)
Updated: 20 hours ago
As artificial intelligence enters the exascale era, photonics is fast becoming the foundation of future high-performance computing clusters. At an OFC 2025 panel session, leading industry experts took a deep dive into silicon photonics, optical interconnect and advanced packaging, revealing today's technical breakthroughs and the road ahead.
Why Photonics? Why Now?
As AI training scale and compute density grow rapidly, traditional copper interconnects are running into physical limits on power and signal integrity — especially for the increasingly demanding scale-up links between GPUs in the same rack. Future AI systems will need hundreds of Tbps of interconnect bandwidth, and a single rack may exceed 600kW, leaving today's pluggable modules struggling with space and thermal challenges.







Speaker Takeaways
🔧 Andy Bechtolsheim (Arista)
1. Optical modules need a 10x improvement in reliability, power and cost
Andy noted that exascale AI data centers may eventually need millions of optical modules, yet today's modules still fall short on stability and efficiency. He pointed to three core pain points:
Reliability: AI systems have extremely low tolerance for transmission errors — a single failed optical link can drag down overall compute efficiency. Current module failure rates are still too high, and soft errors (dirty connectors, laser aging, MPI reflections, etc.) cannot be fully avoided.
Power consumption: today's 800G modules with a retimed DSP design can draw about 25W each. If future data centers deploy hundreds of thousands of modules, total optics power could reach hundreds of MW — power taken directly away from GPUs, sacrificing usable compute.
Cost: high power and low reliability mean high operating costs and downtime risk. Future optical modules must move toward more cost-effective designs built for mass production.
Andy argued the industry must deliver a "10x improvement" on all three fronts to truly support AI-era network architectures.
2. The essential difference between LPO and CPO: not the technology, but the packaging
Andy stressed that LPO and CPO are fundamentally the same in how they transmit signals; the only difference is where and how the optics are packaged:
Item | LPO (Linear Pluggable Optics) | CPO (Co-Packaged Optics) |
Packaging location | Pluggable modules at the front panel | Optics co-packaged directly with the ASIC on the BGA |
Design flexibility | Higher; modules can be swapped and upgraded independently | Lower; tied to a single ASIC and package design |
Serviceability | Modules can be replaced quickly, suited to large-scale deployment and operations | If an optical component fails, the whole system or main board must be replaced |
Ecosystem | Many vendors, high interoperability | Few vendors; highly integrated but highly dependent |
In short, LPO is the practical solution that preserves open modular design, easy servicing and market diversity, while CPO holds more promise for bandwidth density and power control but must overcome serviceability and supply-chain integration issues.
3. Serviceability is key to enterprise optics deployment
Andy further noted that data center operations depend heavily on rapid deployment and fault replacement. His key observations:
With a CPO architecture, if one optical channel fails, the entire switch (even a 70-pound chassis) must be replaced — a huge burden in operations and downtime cost.
By contrast, LPO modules can be replaced in minutes without powering down, greatly improving reliability and RMA efficiency.
He also cautioned that CPO does not eliminate the root causes of errors (connector contamination, reflections, laser instability, etc.). So when a link fault is not caused by hardware damage, a CPO architecture may actually make the problem harder to fix quickly.





🌐 Fotini Karinou (Microsoft)
1. Traditional copper cannot support cross-rack scale-up networks
Speaking for Microsoft Azure, Fotini explained that as AI models reach trillion-parameter scale, the number of GPUs and the memory bandwidth needed to train them are soaring, so the resources inside a single rack are no longer enough — scale-up architectures must extend across multiple racks.
But this introduces a critical bottleneck:
Traditional copper interconnects can no longer meet the reach and power demands of hyperscale AI.
Copper offers low latency and low power, but its reach is limited; once high-speed links are needed between racks, signal attenuation and power problems escalate quickly.
In Microsoft's own testing, when scale-up extends across racks, existing copper technology cannot sustain acceptable latency and transmission efficiency — a major catalyst for the shift to optical links.
2. Architectural shift: from separate interfaces to a "unified physical layer"
In traditional AI compute architectures, GPU-to-host and GPU-to-GPU links usually use different types of interfaces (e.g., PCIe vs. NVLink). That is acceptable at small scale, but in large clusters it leads to inflexibility and performance bottlenecks.
Fotini recommended that future architectures move toward:
A "Unified Physical Layer": one general-purpose high-speed interconnect architecture that simultaneously supports:
GPU ↔ GPU communication (scale-up)
GPU ↔ memory transfers (High Bandwidth Memory)
Efficient data exchange and AI inference workloads
Such an interface needs the following properties:
Flexible: resources can be reallocated as applications change, without being locked to a hardware interface
Low latency: memory-class access latency, even better than conventional DRAM access
Highly compatible: integrates with different modules and packaging technologies (e.g., chiplet architectures)
3. Core technical targets: <4 pJ/bit, low latency, high reliability
Fotini added that for a unified physical layer to become reality, optical interconnect technology — especially silicon photonics — must meet much tougher performance thresholds:
Target | Requirement |
Energy per bit | Below 4 picojoules/bit to match or replace copper |
Latency | Below 500 ns (for RDMA or memory-access-class applications) |
Bit error rate (BER) | Better than 10⁻¹² to support reliable memory-class communication |
Bandwidth density | Support >100 Tbps per rack |
Reliability (RAS) | Highly stable, repairable and serviceable, able to handle 24/7 AI compute loads |
Fotini also emphasized that these new interconnect technologies should not be evaluated at the component level alone, but validated from a system perspective, including thermal design, packaging integration and fault recovery.






💡 Ashkan Seyedi (NVIDIA)
1. Every watt is "performance currency"
Ashkan highlighted NVIDIA's core view on allocating data center compute resources:
"Every watt represents compute that can be converted into revenue." In other words, power spent on moving data weakens the "main force" of AI inference and training.
Citing Jensen Huang at GTC, he said power should be concentrated on the GPUs' own inference and training, not on transporting data. That is why low-power, low-latency optical interconnect has become a core building block of the AI factory.
2. CPO dramatically raises interconnect density
Ashkan explained why Co-Packaged Optics (CPO) is key to AI factory design:
At the same power, CPO can deliver up to 3x the GPU interconnect density.
That lets a system connect more GPUs with the same energy → higher throughput → more tokens → more AI training and inference output.
Simply put, an optical interconnect design that raises performance density directly creates higher data center returns and profit.
3. Technology choices must be evaluated at the system level
Ashkan offered a practical perspective:
Even though many components such as TF-LN, BTO and III-V laser integration now have commercial offerings, are they worth integrating? Component performance alone isn't enough; packaging difficulty, thermal sensitivity, reliability and service strategy must all be weighed together.
For example, on-chip lasers may look convenient, but if reliability is shaky and thermal management is hard, they may end up slowing volume production and deployment.
He reminded the photonics ecosystem not to lose sight of system optimization in pursuit of technical showmanship. What data centers ultimately need are complete solutions that come online fast, consume little power and are easy to maintain.








🚀 Dave Lazovsky (Celestial AI)
1. "Photonic Fabric" — built for scale-up
Dave presented Celestial AI's Photonic Fabric technology, an optical interconnect platform designed specifically for AI scale-up compute architectures — think of it as an "optical NVLink."
The goal: as more and more GPUs need tight interconnection, deliver maximum data-exchange performance at minimum power.
Celestial AI does not compete with scale-out protocols such as Ethernet / InfiniBand; it focuses on high-speed connectivity within scale-up architectures (e.g., GPU ↔ GPU, memory ↔ accelerator).
2. Results already delivered
Photonic Fabric has achieved:
Power below 3.2 pJ/bit
Bandwidth density of 1 Tbps per square millimeter
BER (bit error rate) < 10⁻¹²
No DSP (digital signal processor) required at all
This level of power and error-rate control can support memory-class transfer needs, such as AI model memory disaggregation (RDMA), sharply reducing system power and footprint.
3. Thermally stable SiPh modulators simplify packaging
Celestial uses thermally stable SiPh modulators (e.g., GeSi modulator structures), which avoid high-temperature failures and make the overall package design more flexible, with these benefits:
Can be co-packaged with large silicon ASICs, shortening the electrical-optical path and reducing loss
Maintains very low error rates without a DSP → saving power and area
Packaging platforms such as OMIB (Optical Multi-chip Interconnect Bridge) enable die-level optical connectivity and system integration
This also shows that Celestial AI designs its photonic architecture starting from the system, rather than focusing only on component performance.





🌏 Charley Bu (Accelink)
1. The Chinese market: cost-effectiveness first
Charley noted that in China's AI data center build-out, customers care most not about high-performance components, but about:
"Whether a scalable, stable optical interconnect system can be built at the lowest possible cost"
As a result, many Chinese cloud operators are positive about both:
LPO (Linear Pluggable Optics)
CPO (Co-Packaged Optics)
— and solutions with good power control and competitive pricing will be adopted first.
2. 400G / 800G modules still have a long life cycle
Unlike the U.S. market's rapid shift to 1.6T optics, China still builds mainly with 400G and 800G modules. The reasons:
Mature module technology and a stable supply chain
Clearly lower cost than next-generation modules
Still very practical for China's private AI data centers and SME deployments
He expects Chinese demand for previous-generation modules to keep rising over the next few years, extending the overall module life cycle.
3. Immersion cooling cuts power by up to 40%
Charley added that to further control system power, several Chinese operators are adopting:
Immersion cooling
Applied to LPO / CPO modules and system cooling scenarios
Test results show up to 40% lower system energy consumption compared with conventional air-cooled facilities — a very attractive solution for large-scale AI inference data centers.
Accelink also showed working liquid-cooled modules at the OFC exhibition, indicating the technology has entered real deployment.




Industry Trend Insights
The panel sent one clear message: the AI system network has moved from a "supporting unit" to the "core bottleneck" and a "value driver". Photonics is no longer optional; it is a necessity. The winners will be technologies that balance the following metrics:
Power: < 4 pJ/bit
Bandwidth: > 100 Tbps/rack
Reliability and serviceability
Tight integration with ASIC packaging and readiness for volume production
As AI factories and cloud architectures undergo a full-scale upgrade, photonics will light the way to high-speed, high-efficiency computing.




Comments