top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

2026 OCP APAC Summit | CPO Didn't Take Off for a Decade Because It Wasn't a Real Problem Yet: How NVIDIA Is Pushing Co-Packaged Optics into Volume Production

2 days ago
9 min read

At the OCP 2026 APAC Summit, Gilad Shainer said the line most worth writing down: co-packaged optics (CPO) failed in many past attempts not because the technology was lacking, but because "it wasn't a real problem"—it was not yet solving the most painful problem of the moment. Today three things have crossed the line at once: scale-out bandwidth doubles every generation, a single data center now packs hundreds of thousands of GPUs, and optical network power is starting to eat a "percentage-level" share of the compute budget. Only now has CPO become a must-do. NVIDIA's definition of CPO also differs from every player before it: not just building it, but building it in volume. That dividing line determines which Taiwanese suppliers get in.

1. An AI Factory Is Not One Network, but Four Separate Infrastructures

Shainer opened by dismantling a common misconception: building an AI factory is not about laying down one infrastructure that covers everything. What they are designing is a supercomputer with multiple networks built for different purposes, and the number keeps growing.

  • Scale-up: He said plainly that "this isn't really a network; it's more like compute infrastructure." Its job is to connect a large number of GPU ASICs into "one GPU"—NVLink 72 treats 72 ASICs as one. The key is that NVLink itself participates in compute: it used to handle only reductions, but now runs reduce-scatter, gather, and other collective operations, and is re-optimized for new workloads every generation.

  • Scale-out: The layer he spent the most time on and the most mature, where Spectrum-X lives.

  • Scale-across: Connecting multiple data centers to run the "same" workload, which is relatively new.

  • Context memory infrastructure: A brand-new layer built for inference KV cache and context—keeping data outside the GPU server, but close enough, accessible enough, and power-efficient enough. It is being rolled out with partners.

Add the north-south access network, plus his prediction that "more infrastructure will be built," and the AI factory's network bill will only keep splitting into finer pieces.

The underlying reason is that compute units are no longer homogeneous. A few years ago there were only GPUs; now there are at least three: GPUs, standalone CPUs for agentic workloads, and LPUs. Workloads move between devices, and every move creates new interconnect demand. We covered the other half of this story in Ethernet Moves into Scale-Up: Broadcom's Keynote Takes Aim at NVLink—when everyone wants a piece of scale-up, the interconnect layer becomes the most crowded battlefield of the next three years.


2. Scale-Across: The Real Enemy Is Not Distance, but Deep Buffers

This was the most technical part of the talk, and the least covered by the media.

To run a single workload across data centers, the traditional approach relies on DCI deep-buffer switches. Shainer said they still work today, but the cost is hard to swallow: long distances already carry a physical latency of "5 nanoseconds per meter"—that's physics, accept it. The problem is that once deep buffers start filling up, the drain penalty can exceed the distance itself, reaching 2 to 3 times the latency. You pay once to cross sites, and the deep buffer makes you pay twice more.

So NVIDIA's approach is: no deep buffers. If long distances cannot be lossless, go lossy and put the effort into algorithms—"distance-aware" adaptive routing and congestion control.

An implementation detail worth noting: when a connection is set up, a probe first measures the distance, and parameters are then set for each RDMA connection accordingly. Tiers fall roughly at 500 meters, 2 km, and 10 km, each with its own settings.

And scale-across starts at 500 meters, not tens of kilometers. This matters far more to Taiwan's optical communications supply chain than it sounds: cross-building links within a campus and multiple sites in the same city are now within scale-across range. For related context, see NTT Brings AICC to OCP: AI's Next Bottleneck Is Not Compute, but Connecting Distributed Compute with Light.

Another easily missed point: he stressed this is not only about switches and paths. When a GPU sends data down a long path, the data should not keep occupying GPU memory; it should move to another memory region to wait—so software, frameworks, and the network must be designed as a single unit. NVIDIA claims that, done this way, long-distance performance can double.

3. Why CPO Now: Three Things Crossed the Line at Once

"There were many attempts and they didn't succeed. I think the reason is that it wasn't a real problem."

Everyone working on CPO should copy this onto a whiteboard. Shainer said bluntly that CPO is not new; there were many previous attempts, and it never got real attention because the pain it solved wasn't painful enough. So why is it a real problem today? He gave three simultaneous conditions:

Condition 1: Scale-out bandwidth doubles every generation, and scale has reached hundreds of thousands of GPUs. Optical network power has gone from "a line item" to "a percentage"—he said the scale-out layer alone can consume a percentage-level share relative to compute.

Condition 2: Scale-up is leaving the rack. Today, raising in-rack density still allows copper—he even said "use copper when you can; it's the best thing ever"—but once you go across racks, optics is mandatory, and that is yet another power bill. This is exactly the moment when the yardstick in After Copper Runs Out for AI: Seven Paths to Scale-Up Optical Interconnect starts to matter.

Condition 3: The unit of value has changed. The most important metric now is "how many tokens per unit of energy." So every possible watt must be saved: liquid cooling, and specifically warm-water liquid cooling (eliminating heater power), plus 800V racks (raising compute density while optimizing power distribution). Under this math, every watt optical networking saves is more compute you can run.

Stack these three together, and CPO shifts from "cool technology" to "unavoidable cost engineering."

4. What NVIDIA Got Right: Micro-Rings, Packaging, Lasers, Fiber Arrays

Shainer said they "did it differently from past attempts," in four specific areas:

  1. Shrinking the optical engine with micro-ring modulators. The official abstract further states this is the world's first 1.6T CPO chip based on a new micro-ring modulator, paired with a 3D-stacked silicon photonics engine.

  2. Packaging: Co-packaging the photonic IC (PIC) and electronic IC (EIC) with high confidence and high reliability—the explicit goal being volume production with strong yields.

  3. Custom laser sources: Reducing laser count while still delivering enough optical power to the CPO. This directly affects both BOM cost and reliability.

  4. Fiber array attachment: He specifically noted that "there are different ways to handle the fiber array," and the approach they chose allows full verification of finished parts. Together with the detachable fiber connector mentioned officially, this points to serviceability and production yield.

The official abstract's quantified claims versus traditional approaches: 3.5x power efficiency, 63x signal integrity, 10x network resiliency, and 1.3x faster deployment, deliberately designed to fit OCP mechanics and be managed by SONiC.

A caveat: these multiples are vendor claims, and the baselines and measurement conditions were not explained on stage—fine for comparison, but cite them as hard metrics with caution.

5. 10x MTBI: CPO's Real Selling Point May Not Be Power Savings

Power savings are the motivation, but the hardest number presented was elsewhere.

Shainer said they accumulated millions of hours across three labs—tens of millions of hours in total—collecting MTBF and MTBI data, and that "the numbers other CSPs got all fall into the same range." The result: 10x better MTBI than pluggable optical modules—meaning, on average, 10 times longer running time before an interruption.

This is more persuasive to buyers than "3.5x less power." In a training job spanning hundreds of thousands of GPUs, the cost of one interruption is not a reconnect—it is an entire fleet of idle GPUs and a checkpoint rollback. What CPO is really selling is "no interruptions"; power savings are a bonus.

He also made the status clear: Spectrum-6 is in volume production, the Spectrum-6 CPO version is also in volume production and has begun shipping, with customers running it. Going from "feasibility demo" to "MTBI statistics" is a completely different stage—to see where each player stands today, compare with CPO Players' Roadmaps: Beyond the Big Four, a Whole Line of Challengers.

6. Digital Twins: Compressing Validation from Months to One Week

This section is easy to skip as a demo, but it is actually NVIDIA's other hand for locking in the entire supply chain.

They have turned the entire AI factory—from how much land and power it needs and how to cool it, to all networking equipment—into a reference architecture plus a digital twin. This simulation environment can run the whole infrastructure, configure every component, and validate it before physical installation.

The figure from CSP partners: network validation compressed from months to one week. Even software updates can be run in the twin before being applied to the physical data center, pushing validation down to "days." He said the platform is open, and access can be requested on the website.

The implication for component makers is direct: as customer validation moves into digital twins, whether your component models are included in that reference architecture will directly affect whether you get designed in.

7. Two Key Signals from the Q&A: SerDes Path Undecided, ELS Not Locked In

Signal 1: The SerDes path for next-generation scale-up is not settled. Asked whether scale-up would follow a particular MSA, he answered that scale-up and scale-out both use 200G-class SerDes today; NVIDIA participates in multiple MSAs (including the OCI MSA and 400G-related MSAs); on options like 50G, "I'm not sure"; and PWDM "has potential" at higher speeds. The bottom line: multiple paths are being explored in parallel, and the best one will be chosen when the next generation is finalized.

In plain terms: it is too early to bet on a single SerDes or single wavelength scheme. The supply chain should keep capabilities on both paths rather than go all-in. This aligns with SemiAnalysis saying at the same event that "copper vs. optics is a false dichotomy".

Signal 2: ELS and optical engines are deliberately not locked in. Asked whether the external laser source (ELS) would create vendor lock-in, his answer was quite open: if you don't like CPO, use pluggables, NPO, or CPX; the optical engine can sit on the board, on an interposer, or you can use someone else's CPO; ELS can be bought from different vendors. His words were roughly, "There's always a choice; pick the one that makes sense for you."

This is not purely goodwill. Opening up ELS and packaging options is itself a way of outsourcing supply chain risk to the ecosystem—and it also leaves a crack in the door for second-tier optical component makers.

8. What It Means for Taiwan's Supply Chain

In the final question, Shainer spelled out NVIDIA's bar very bluntly, and this was the most useful part of the talk for Taiwanese vendors:

NVIDIA cannot launch something just because it's "cool and interesting." To launch it, it must be manufacturable at scale, with good yields and massive capacity.

He added that this is precisely why NVIDIA has made its public investments—investing in specific companies to ensure they have the right technology, the right products, the ability to ramp capacity, and complete production lines.

Breaking this down into three practical propositions for Taiwan's supply chain:

1. The ticket into CPO is a production line, not a spec sheet. Showing samples with great yields is not enough; NVIDIA wants to see "massive capacity with stable yields." This favors packaging and test houses, fiber array assemblers, and laser makers that already have volume-manufacturing discipline; it is bad news for startups with only lab data—unless they get investment and are tied into the supply chain.

2. Getting investment = getting onto the roadmap. He effectively admitted the purpose of investment is to lock in capacity and technology. Conversely, reading NVIDIA's investment list is reading its BOM for the next two generations, more reliable than any spec rumor.

3. The genuinely new market is 500 meters to 10 km. With scale-across starting at 500 meters and tuned in three tiers (500m / 2km / 10km), a set of product needs will emerge that is neither traditional DCI nor short-reach in-hall. Vendors of optical modules, fiber infrastructure, and connectors for this distance range should put it into their 2027 plans.

The risks should be stated too: once CPO rolls out, pluggable module makers will see their value on the scale-out switch side compressed, with value concentrating in optical engines, laser sources, packaging, and fiber arrays. Shainer himself said "if you don't like it, use pluggables"—the flip side of which is "we no longer need you." At the same event Arista argued that XPO gives pluggables another generation; both views can hold at once—meaning the transition will be longer than either side claims, but the end direction is not in dispute.

Conclusion

The most important thing in this keynote is not the multiples, but that NVIDIA has changed the definition of CPO: CPO is no longer a technology that must be proven feasible, but a product that must be proven manufacturable at volume.

The criteria have changed accordingly. From today, when evaluating any CPO player, stop looking at pJ/bit and demo videos and look at three things: MTBI statistics, a production line with massive capacity, and inclusion in customers' reference architectures and digital twins. Those with none of the three are still in the previous era.

As for Taiwan's supply chain—the opportunity this time is not "can you build it" but "can you get on that list." The door is closing, but it hasn't closed yet.

This article is for technology and industry trend analysis only and does not constitute investment advice.

Related Reading



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page