Technical Analysis | OCP 2026 White Paper: OCS Moves from Google's In-House Secret into the Whole Data Center Industry's Toolbox
The most-cited yet most-misunderstood document in optical communications over the past three years is Google's "Jupiter Evolving" paper from SIGCOMM 2022. It revealed that Google had already replaced its data center spine layer with Optical Circuit Switching (OCS) — cutting network power by 41% and CAPEX by 30%.
At the time, it looked like "Google using yet another in-house technology nobody else can replicate." But in April 2026, the Open Compute Project (OCP) released a white paper authored by five players with very different vantage points — iPronics, Lumentum, Ciena, Carnegie Mellon and Lumotive — elevating OCS to the level of industry consensus.
The most important signal in this white paper is not its technical content but the fact that it exists at all: OCS has gone from "Google's secret weapon" to "the next-generation data center backbone officially endorsed by the OCP community." The entire industry — including Taiwan's supply chain — should start putting it in the toolbox instead of treating it as a lab curiosity.
1. Why Now
EPS (Electrical Packet Switching) has carried the industry for 30 years, but in the AI training era it has hit a bottleneck. The white paper identifies three structural problems:
The first is power. Every EPS has to perform OEO (optical-electrical-optical) conversion, and that conversion is itself a power hog. The white paper puts a number on it: a large data center using 64 chassis-based 16-slot spine switches at 30 kW each (with 400G semi-coherent optics) burns 1.9 MW in the spine layer alone at full load — before counting cooling.
The second is latency and jitter. Collective operations in AI training (AllReduce, AllGather) are synchronous and blocking — if a single packet is late, the entire batch of GPUs has to wait. Packet buffering, queue contention and congestion control in EPS are all sources of latency variation. Reference [9] is blunt: in large clusters, network performance variability noticeably degrades application-level scalability.
The third is scalability. The radix of a single packet switch actually drops as per-port speed rises — port count can't keep up with bandwidth, so you stack more tiers, and every extra hop adds another OEO conversion, degrading both power and latency.
The OCS answer is brute force: just don't convert to electrical at all. Light enters an input port and is steered directly to an output port, staying entirely in the optical domain. No packet buffers, no OEO, no QoS processing. The price is losing packet-level control.
2. Think of OCS as a "Glass Window," Not a "Processor"
The white paper contains a particularly good analogy. It says OCS should be understood as "a clear glass window" — it simply guides the light through, without reading it, processing it, or even knowing what's being carried.
This makes EPS and OCS two completely different species. EPS is like a post office: every packet is opened, its address read, a forwarding decision made, and then it's resealed. OCS is like a railway switch: throw the lever and the whole train changes direction — what cargo is in the cars is none of its business.
This difference in species has four direct consequences:
Benefit 1: Because it doesn't read content, OCS is protocol-agnostic and rate-agnostic. The same OCS can carry 100G today, 400G tomorrow and 1.6T the day after, with no hardware swap. For hyperscalers, this means the spine layer can be deployed once and used for the long term, rather than ripped out every two or three years.
Benefit 2: No OEO conversion and no electronic processing means very low power, extremely low latency and near-zero jitter.
Cost 1: OCS is circuit switching — once a point-to-point connection is set up, it stays up, and it cannot dynamically fan out to multiple destinations the way packet switching can. So it is ill-suited to highly dynamic, bursty, unpredictable traffic.
Cost 2: Because it doesn't read packets, QoS, traffic inspection and telemetry are all impossible. Accumulated insertion loss also limits the number of OCS "hops" — the white paper states plainly that practical deployments today can only sustain about 1–2 hops; beyond that, SOAs (semiconductor optical amplifiers) are needed to make up the loss.
Understanding these four points is key to understanding why OCS is not a "replacement" for EPS but a "complement" — it takes over the part that hurts EPS most (long-reach, high-volume, slowly changing spine-layer connections) and leaves packet-level control to EPS.
3. Technology Landscape: Six Paths, Each with Its Own Battlefield
The white paper groups the physical technologies currently able to implement OCS into six categories. This comparison table is the most valuable part of the entire document.
Technology | Radix (port count) | Insertion Loss | Switching Time | Best Fit |
Robotic (mechanical fiber patch panel) | Large | Low | Seconds to minutes | Slow-changing, hyperscale |
MEMS (micro-electro-mechanical mirrors) | Medium | Medium | Milliseconds | Volume mainstream; used in Google TPU |
Liquid Crystal | Medium | Medium | Milliseconds | Requires polarization diversity |
Piezoelectric | Medium | Medium | Milliseconds | Precise alignment, long-term stability |
Silicon Photonics (SiPh MZI) | Low–medium | High | Nanoseconds to microseconds | Fastest switching, manufacturable |
Metasurface | Large | Medium | Microseconds | Compact, polarization-independent |
Robotic (mechanical fiber patch panel): large radix, low insertion loss, seconds-to-minutes switching; suited to slow-changing hyperscale fabrics
MEMS (micro-electro-mechanical mirrors): medium radix, medium insertion loss, millisecond switching; the volume mainstream (used in Google TPU)
Liquid Crystal: medium radix, medium insertion loss, millisecond switching; requires polarization diversity
Piezoelectric: medium radix, medium insertion loss, millisecond switching; precise alignment and long-term stability
Silicon Photonics (SiPh MZI): low-to-medium radix, high insertion loss, nanosecond-to-microsecond switching; fastest switching and manufacturable
Metasurface: large radix, medium insertion loss, microsecond switching; compact and polarization-independent
This table reveals three key industry calls:
Call 1: MEMS is the present tense. Google's TPUv4 Superpod uses 128-port MEMS-based OCS and has been running in production for years. MEMS maturity, reliability and port count all line up with industry needs, and it will remain mainstream for the next 3–5 years. Companies like Lumentum and Coherent are investing in capacity precisely for this line.
Call 2: Silicon photonics OCS is where the decisive battle will be fought. Why? Because its switching time is nanoseconds to microseconds — 1,000× faster than MEMS. Once AI training needs to "dynamically reconfigure the network topology between iterations," millisecond-class MEMS is too slow; only silicon photonics can keep up. The trade-offs are high insertion loss and limited port count, but both can be addressed with SOA integration and multi-stage cascading. This is the design battleground of the next decade.
Call 3: Metasurface is the dark horse. Lumotive's presence among the white paper's authors is no accident — with ±80° wide-angle beam steering and sub-wavelength pixels, they are building a compact, high-density, polarization-independent free-space OCS. This path is still at the narrative stage, but if the form factor truly materializes, it will upend how intra-rack interconnect is designed.
4. Five Use Cases, Three Quantified Milestones
The most practically useful part of the white paper breaks OCS applications into five concrete use cases with verifiable numbers. We'll focus on the three that matter most:
Case A: Google Jupiter — Spine-Layer Replacement
This is the most mature application and it's already happening. Google replaced all spine-layer EPS in its data centers with high-radix, slow-switching OCS. All leaf EPS are interconnected directly through a flat optical layer.
The numbers: 41% lower network power, 30% lower CAPEX, 5× higher throughput, and dramatically shorter deployment time. These are not simulation results — they are production figures Google published in the Jupiter Evolving paper [1].
More importantly, an OCS spine is "transparent" to leaf upgrades — as leaves move from 100G to 400G to 800G, the spine stays untouched. That is impossible in an EPS architecture.
Case B: Google TPUv4 Superpod — From Static to Dynamic AI Clusters
This is the landmark application where OCS first moved from "network backbone" into "AI compute fabric."
Google TPUv4 combines 64 Cubes (64 TPUs each) into one Superpod of 4,096 TPUs. Each Cube connects to three groups totaling 48 128-port MEMS OCS units (16 each for the X, Y and Z dimensions). This means the Superpod's internal topology isn't fixed — it can be dynamically reconfigured into a torus, ring, mesh or other shape based on each training job's communication pattern.

This figure shows the core concept of reconfigurable Cubes: the same physical hardware cluster can be carved into different sub-units depending on job requirements. How big is the payoff for LLM training? The white paper cites reference [11]: compared with a static configuration, the TPUv4 Superpod delivers a 3.3× performance gain on large language model training while cutting power by 9%.
And that's before counting the benefit to failure recovery. In a traditional EPS architecture, a 1,024-TPU slice at 99.9% server availability achieves only 25% effective throughput; with a reconfigurable OCS Superpod, the same slice size at the same availability reaches 75%.
Case D: ACTINA — Taking OCS Down to the GPU Port Level
If TPUv4 is "dynamic reconfiguration between racks," the ACTINA paper [5] presented by Columbia University and NVIDIA at SC 2025 pushes the idea to its limit — reconfiguration directly at the GPU port level.

This figure shows ACTINA's core innovation: each GPU directly integrates a "multi-port, wavelength-reconfigurable silicon photonics transceiver" (implemented with MZIs plus a DWDM comb laser), allowing optical-path bandwidth to be reallocated dynamically within a single training iteration.
How strong are the results? Compared with an OCS-enabled 3D Torus (i.e., Google TPUv4's topology), ACTINA's OCS-BCube delivers 1.84× faster iteration time while maintaining comparable energy efficiency; compared with Google Jupiter's OCSFT-2L architecture, it cuts energy consumption by 1.72× and improves tokens-per-joule by 1.75×.
ACTINA's key assumption is that silicon photonics OCS switches fast enough to complete the next stage's topology reconfiguration during the GPU compute phase. That assumption holds only if the earlier call — "silicon photonics OCS is where the decisive battle will be fought" — proves right. If yield and cost of silicon photonics plus SOA integration line up before 2028, this path will take the high end of the market straight from MEMS.
Case E: MixNet — A Hybrid Electro-Optical Network Built for MoE Training
Finally, MixNet [7], presented by an HKUST team at SIGCOMM 2025, is another signal pointing toward "silicon photonics OCS + direct GPU attachment."

This figure shows MixNet's hybrid architecture: local scale-up (NVSwitch) handles intra-server TP traffic, global scale-out (EPS) handles DP and PP, and a Regional OCS Fabric is dedicated to all-to-all communication for EP (Expert Parallelism). MixNet exploits two observations: EP communication is localized (only 8–32 GPUs participate at a time), and changes between iterations are gradual (so the previous round's pattern can predict the next).
Measured result: training time for the Mistral 8×7B MoE model is 1.6× shorter than on a static fat-tree.
MoE is now the dominant direction in LLM design (GPT-4, Mixtral and DeepSeek-V3 are all MoE architectures), which means the MixNet approach has real, immediate, at-scale market demand.
5. Industry Verdict: Three Calls
Having read the entire white paper, here are three clear industry calls:
Verdict 1: OCS has gone from "crying wolf" to "the wolf is at the door". For the past three years the industry took a wait-and-see stance on OCS — after all, only Google was using it. But this OCP white paper, the ACTINA and MixNet papers, plus NVIDIA's work on OCS resiliency (Case C in the white paper comes from NVIDIA's latest publication [4]) mean that at least Microsoft, Meta, Amazon and NVIDIA are evaluating or piloting OCS architectures internally. The silicon photonics supply chain has now entered the "lock capacity, pay deposits, fight for allocation" phase (consistent with the view we laid out in Earnings Highlights: Tower Semiconductor (TSEM) | Q1 2026), and OCS is one of the key drivers of this wave of demand.
Verdict 2: Silicon photonics OCS is the ultimate answer, but MEMS will hold on for another 3–5 years. MEMS is in production, reliable and cheap — today's money-maker. But once data center topologies need "intra-iteration dynamic reconfiguration" (the ACTINA and MixNet direction), millisecond-class MEMS won't be enough and nanosecond-class silicon photonics MZIs will be required. This inflection point falls roughly in 2028–2030, closely overlapping the Coherent Lite ramp timeline analyzed in ZR Is the Meat, Coherent Lite Is the Bone: The Next Decade of the Optical Transceiver Market — no coincidence; the entire optical communications industry will go through a generational shift in this window.
Verdict 3: The next battlefield is "GPU port-level reconfiguration". The ACTINA and MixNet papers point to the same thing: integrating OCS directly into GPU transceivers. What does this mean for Taiwan's supply chain? Three opportunities: (a) silicon photonics PIC foundry — Tower already has a queue but not enough capacity, opening the door for other foundries; (b) DWDM comb lasers and EML chips — volume will scale, and Taiwanese vendors already have a foundation here; (c) integrated packaging and test — know-how from CPO / NPO / XPO packaging formats will carry straight over. Taiwan's positioning is still in upstream components, not DSPs or system integration.
When Google unveiled Jupiter Evolving in 2022, most of the industry saw it as "Google using something nobody else can do." Three years later, in 2026, OCP is telling us with a white paper: others are getting ready to do it too. For Taiwanese vendors, the window between "watching this technology" and "preparing capacity for it" may be only 18 to 24 months.
This article is for technology and industry trend analysis only and does not constitute investment advice.




Comments