top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

OCS (Optical Circuit Switching): The Next Optical Revolution in AI Data Centers

36 minutes ago
5 min read

Over the past two years, the AI boom has made the phrase "the data center bottleneck" sound like more than a cliché for the first time. The problem is not the GPU, not HBM, not PCIe — it is the network at the very bottom: the switching architecture. AI clusters keep growing, from a single room packed with GPUs to many rooms stitched together into an AI super-factory. And every time scale grows 4x, the cabling, power and latency of the switching network explode exponentially.


The whole industry now sees the same thing:

The core of next-generation AI networking is not faster SerDes, but smarter topology plus lower-power switching.

That is why OCS (Optical Circuit Switching) is drawing attention.



Why Is Everyone Taking OCS Seriously?

In my view, the reason is:

Light should not be forced into electrons and then back into light.

A traditional switch (i.e., OEO switching) does the following:

  1. Optical → electrical

  2. Switching handled inside the chip

  3. Electrical → optical

That sounds reasonable, but at AI-scale data volumes it turns into three big problems:


1. Electrical switching is the bottleneck

All traffic must pass through the switch chip, so:

  • When SerDes speeds can't keep rising, your entire network upgrade stalls

  • Chip power becomes unreasonably high

  • Switching equipment becomes big, hot and expensive


2. Real-time OEO processing is too costly

1,000+ racks and tens of thousands of data flows going in and out, each requiring OEO conversion — the cost scales linearly, but traffic grows exponentially.


3. Performance can't scale with the cluster

Everyone knows GPU scaling is hitting diminishing returns, but few notice that:

The network is the main reason diminishing returns arrive early.

OCS is like replacing one of the rail tracks with a frictionless one.


The Essence of OCS: "Optical Path Reconfiguration" That Skips O/E/O

It is not a "faster switch"

Nor is it a "new type of optical module"

OCS is:

  • A way to connect optical paths directly to optical paths

  • Every connection is a "dedicated optical path"

  • Data rate, protocol and modulation format don't matter


In other words:

You connect fiber A to fiber B without touching electronics in between.

This also means:

Fully Rate-Agnostic

→ No need to replace the OCS for 800G / 1.6T / 3.2T or beyond


Truly Low Power

→ No DSP, no SerDes, no switch chip


Highly Scalable

→ Want more ports? You're just connecting more fiber

→ No need to redeploy the spine layer


Ultra-Low, Stable Latency

→ No packet processing, no buffers, no ASIC pipeline


The Four Technology Paths of OCS

There is a lot of engineering detail here, but I'll summarize it as simply as possible:

1. MEMS: mature, but with limits

  • Rotating mirrors, ~25ms switching

  • Low insertion loss, large scale, first to commercialize

  • Adopted in Google TPU v4 / v5 / Ironwood

→ The most mature OCS in the near term, but with clear long-term bottlenecks.


2. Liquid crystal (LC / DLC): extremely high reliability

  • Derived from WSS technology

  • World-leading reliability

  • Slower switching (~100ms)

→ Suited to large-scale spines that need high reliability and stability.


3. Piezoelectric (Piezo / DLBS): stronger physical limits

  • No mechanical moving parts at all

  • Lowest insertion loss and return loss

  • Harder to scale to ultra-high port counts

→ Very promising long-term, but the challenge is scaling.


4. Silicon photonic waveguides (SiPh): the one to watch

  • Fastest switching (<100 µs)

  • Insertion loss can be compensated with SOAs

  • Greatest cost-reduction potential (CMOS process)

→ Very likely the ultimate solution.


Why OCS Is Really Taking Off: All Three AI Networks Are Hitting Limits

AI is not one network but three:

  • Scale-up (within a rack / pod)

  • Scale-out (within a data center)

  • Scale-across (between data centers)

OCS has found a place in all three.


1. Scale-up: TPU's 3D torus can't be held together without OCS

The evolution of Google's TPU topology shows why OCS is necessary.

TPU v4

  • 4,096 TPUs

  • 48 units of 136-port MEMS OCS

  • 3D torus topology relies on OCS to reconfigure optical paths

  • Both latency and power drop significantly

Ironwood

  • 9,216 TPUs (more than double the scale)

  • Requires double the OCS ports

  • Architecture fully retains OCS rather than CPO

This shows one thing:

Google isn't "trialing" OCS — it has made it part of the TPU system design.

2. Scale-out: redesigning the data center spine layer

Google embedded Apollo OCS into its Jupiter network, with results that shook the industry:

  • Latency down 10%

  • Throughput up 30%

  • Overall power down 40%

  • Cost down 30%

The impact of optical switching in large-scale data centers is far more pronounced than expected.

And this is not "partial traffic optimization"

It is Google's conclusion from real-world operation:

OCS in the spine layer = less pressure on the entire DC network

For hyperscalers, that is an incentive impossible to ignore.


3. Scale-across: NVIDIA's third network (across data centers)

With Spectrum-XGS, NVIDIA introduced a new concept:

Scale-across is the third critical network of the AI era.

When multiple data centers need to become "one AI cluster":

  • Massive DCI (long-haul optical links) is needed

  • Dynamically reconfigurable topology is needed

  • Low-latency synchronization across regions is needed

These needs map almost perfectly to OCS characteristics:

DCI requirement

OCS characteristic

High bandwidth

Rate-agnostic, scalable ports

Dynamic topology

Real-time optical path reconfiguration

Long distance

Low loss, works with C-band

Heterogeneous environments

Fully protocol-agnostic

Coherent has further announced:

→ A DCI-dedicated C-band OCS launching in 2026

This means:

Scale-across will be the next big OCS growth market.

OCS vs. CPO: Not Rivals, but the Twin Cores of AI Networking

This is the most commonly misunderstood part today.

  • Wrong view: OCS will replace CPO

  • Right view: CPO and OCS are complementary


NVIDIA's data makes it clearest:

Network architecture

Power

Pluggable

83 pJ/bit

Pluggable + OCS

50 pJ/bit

CPO

48 pJ/bit

CPO + OCS

31 pJ/bit (lowest)

This proves:

  • CPO handles high-speed, short-reach switching (rack / ToR / leaf)

  • OCS handles topology reconfiguration and spine / DCI

Combined → the end-state of the AI SuperFabric.


Supply Chain View: Which Positions Are Worth Watching?

I break the supply chain into three layers:

Layer 1: Core switching technology

MEMS

  • Globally dominated by one or two players (e.g., Calient, Lumentum)

Liquid crystal (LC / DLC)

  • Coherent has the most mature technology

Piezoelectric (DLBS)

  • Strongly pushed jointly by Polatis + Lingyun Photonics

Silicon photonic waveguides

  • iPronics

  • China's Taclink is catching up fast (offering 32×32 with in-house SOA capability)

The path most worth watching over the next decade: SiPh OCS


Layer 2: Optical / passive components

Including:

  • Fiber arrays

  • FAU

  • Lens arrays

  • Polarization components

  • Optical wedges (YVO₄)

  • Coupling assemblies

  • WDM(Z-block)


Layer 3: Complete / system-level OCS

  • Calient / Polatis (mature)

  • Coherent (liquid crystal systems)

  • TeraHop (an Innolight subsidiary)

  • Advanced Fiber Resources (entering via contract manufacturing)

  • Taclink (SiPh OCS prototype)

This layer will grow very fast, especially with:

  • Hyperscalers building their own OCS

  • Chinese AI data centers investing in all-optical networks


Finally: OCS Is Not About Being "Faster" but "Simpler"

Networking in the AI era is evolving from:

Electrical switching → hybrid opto-electrical → all-optical network

OCS is not a "faster switch"

but rather:

The core of a new architecture that makes network topology flexible, low-power and scalable again.

If you work on:

  • AI Infrastructure

  • GPU/TPU/ASIC clusters

  • Hyperscale data center architecture

  • Optical modules / SiPh

  • Cloud Networking

then the term OCS will show up more and more in your meetings and roadmaps over the next five years.


Summary

OCS is foundational infrastructure for next-generation AI data centers, playing a role similar to the GPU's in AI. It won't replace everything, but it will redefine the logic of everything.
  • It will be a must-have for TPUs / AI ASICs

  • It will reshape the spine layer

  • It will be the standard option for next-generation DCI

  • It will coexist with CPO — even become its best partner

  • It will drive a reshuffle of the optical supply chain

  • It will become a new battleground for the SiPh industry

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page