top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

2026 OCP APAC Summit | Lightmatter | Bijan Nowroozi | Half of AI Compute Is Waiting on the Network: CPO Goes from Proprietary to Open, Spec Due in Q4

2 days ago
9 min read

  • Bijan Nowroozi, head of ecosystem development at Lightmatter, built this talk around a single number: Model FLOPS Utilization (MFU) in training clusters is only 38–43%. In other words, at any given moment more than half the XPUs are sitting idle. This isn't a scheduling problem — interconnect simply can't keep up.

  • Two walls are closing in at once. On the chip side, Rent's Rule means compute grows with area squared while usable bandwidth grows only linearly with edge length — and HBM eats that edge. On the link side, 448G PAM4 squeezes copper's unrepeated reach from 2 m in the 112G era down to 25 cm.

  • The real news is not "is CPO feasible" — 2.6M-hour MTBF has already been validated in volume at hyperscalers — but that every working implementation is proprietary. On August 13, OCP formally launched Open Silicon Photonics for AI Systems as a workstream: 19 companies, a roughly 300-page white paper, and the first spec version to be delivered in Q4 2026.

  • The targets are spelled out plainly: grow a single-tier domain from 72 to 1,024+ XPUs, raise MFU to 55–65%, reach below 5 pJ/bit, and cut the network's share of power from about 31% to 15%.

  • What it means for Taiwanese vendors: FIT, QCT and GUC are already on the list, but in manufacturing and design-service seats. The genuinely new part numbers this white paper opens up are in passive optical infrastructure — faceplates, ODFs, fiber raceways and external laser source (ELS) enclosures. This is one of the few segments where connector and mechanical vendors can still stake a claim from zero.

1. Start with the uncomfortable number: 38–43%

Nowroozi skipped the entire "why do we need optics" section up front. In his words, "I assume you've already figured that out," and he went straight to the question itself: why can't we just keep adding compute?

The answer is MFU. He cited public paper data from Llama training systems, where MFU lands at 38–43%. The metric's definition is brutal — it is the floating-point work actually achieved divided by what those chips could theoretically run. 38% means that of the hundred accelerator cards you bought, only thirty-eight cards' worth of compute actually turns into model parameter updates; the other sixty-two are waiting on data.

The reason this is only now becoming the main argument is that it has finally grown large enough to measure in dollars. For a $500M-class training cluster, the idle compute is worth roughly $285M to $310M. You can build more data centers or scale across, but every unit of compute you add is stranded at the same ratio.

Adding compute isn't broken — its marginal return is just being held down by interconnect.

We break down the "is copper vs. optics even the right question" debate from another angle in 2026 OCP APAC Summit | SemiAnalysis | Dan Nishball | Scale-Up Sophistry: Copper vs. Optics Is a False Dichotomy; reading the two talks together gives a fuller picture.

2. The first wall: Rent's Rule and the edge length eaten by HBM

Nowroozi used Rent's Rule to explain the first bottleneck — the most memorable part of the talk.

The rule itself is simple: compute resources grow with a chip's area, while I/O bandwidth grows with its edge length. Area grows quadratically, edge length linearly. So each generation that makes the die bigger explodes the logic in the middle, but the channels for moving data in and out can never keep pace. This scissors gap isn't something process technology can fix — it's geometry.

What really widens the gap is HBM. He gave three scenarios:

  • Full edge available, no HBM attached: roughly 122 Tb/s can flow out. This is why today's switch ASICs reach the 100T class — all four edges are used for SerDes.

  • XPU with HBM attached: HBM directly consumes a large share of pins and perimeter, leaving only about 33 Tb/s that can be driven out directly. This is the status quo.

  • Adding an I/O die for fan-out (he cited the Rubin-generation approach): packaging "unfolds" the effective edge, roughly doubling it to about 63 Tb/s.

Put the three numbers side by side and the conclusion is direct: within the same process generation, your I/O ceiling is set almost entirely by package geometry, not by transistors. That also explains why advanced packaging kept coming up at OCP — at the same Summit, TSMC went as far as saying "the era of component-level validation is over" (TSMC Says It Plainly: Advanced Packaging Is the Gatekeeper of AI Compute).


3. The second wall: beyond 448G, copper only reaches 25 cm

The second wall is the link itself. Nowroozi showed a loss curve with frequency on the x-axis and insertion loss on the y-axis, with three lines:

  • Megtron 7-class laminate: the steepest drop — already struggling at the 224G/lane dashed line, and essentially out of the question at 448G.

  • Megtron 9-class laminate (he stressed this is only an example; other suppliers have equivalent grades): somewhat better, but how far you can run on the motherboard is still not encouraging.

  • Twinax copper / flyover cable: a noticeably flatter curve.

This explains why "bypass the PCB" solutions like CPC, CPX and flyover have proliferated over the past two years — not because everyone loves making cables, but because electrons can't travel far enough on the board.

Converting to distance makes it more intuitive. Without retimers, running directly on laminate, unrepeated reach is about 2 m at 112G/lane, about 1 m at 224G/lane, and about 25 cm at 448G/lane.

What does 25 cm mean? It's shorter than the depth of a 1U chassis, and shorter than the distance from an accelerator card to the rack backplane. As Nowroozi put it, "physics is pushing back" — you can move data, but if you need to move a lot of it, physics won't let you move it very far.

Coherent's message at the same Summit was complementary: what's holding back 400G/lane is actually the electrical side, not the optics (Coherent: What Limits 400G/lane Isn't Optics — It's Electrical). Together the two talks make the full conclusion: copper isn't exiting because optics got cheaper, but because copper hit the wall first.

4. So why hasn't this been solved yet? Three very non-technical reasons

Nowroozi himself admitted this part was "more subjective," but I think it was the most valuable part of the talk — because it's about business, not physics.

First, reliability is a function of scale. Pluggable module MTBF has long been criticized, but in a rack with only 18 servers you barely notice. Once you reach NVL72 scale and need to run five thousand links, MTBF goes from a number on paper to troubleshooting frequency. The pain grows super-linearly with pod size — per public data, pluggables sit around 0.55M hours MTBF versus 2.6M hours for CPO-class, a gap of nearly 5x.

Second, too many options is itself a cost. Front-panel modules, XPO, NPO, OBO, CPC, CPX — multiplied by the choice of 224G or 448G — every combination needs its own toolchain and know-how. Each additional option a system vendor supports means another team to staff. It's a hidden tax that locks down speed.

Third, early movers fragmented the market. Large players picked their routes early and bet on their own silicon, splitting the market into two or three camps while everyone else stands outside unsure which side to join. Common protocols like UALink and UEC have helped, but Nowroozi was blunt — they are "point patches" that don't solve the end-to-end system-level problem.

We lay out all seven routes side by side — and which wall each one hits — in After Copper Can't Carry AI: Seven Paths for Scale-Up Optical Interconnect, and Two Ways to Live With the Walls. Lightmatter's position in this talk: don't choose — define the interface so all seven paths can plug into the same rack.

5. Lightmatter's answer: not selling an optical engine, but an "interface contract"

The positioning needs to be clear here, or it's easy to misread this as Lightmatter pitching its own Passage.

Open Silicon Photonics for AI Systems is technology-neutral. The white paper explicitly supports silicon photonics, VCSELs and micro-LEDs, and even multimode with microlens arrays within a single tray. What it defines is not "whose optical engine," but four things:

  1. Interface contract: extending OCP's MHS and ORv3 to lock down the tray, waveguide, power and cooling interfaces. Once fixed, whether the XPU sits in the green box or the blue box, or moves forward toward the faceplate, becomes a reconfigurable choice.

  2. Passive optical infrastructure: faceplates, ODFs (optical distribution frames), raceways (fiber routing trays) and ELS (External Laser Source) enclosures. The key is that this layer is deliberately decoupled from the SerDes generation — you don't have to re-pull the fiber plant at each upgrade.

  3. Link profiles: from DR (single wavelength) and OCI (4+4) all the way to DWDM 16.

  4. OE taxonomy: CPO, NPO, XPO and CPX coexist, with no one forced to pick a side.

On top of that is OCP's modular spec process — first establish a base spec, then others layer on top like extending an API, achieving compatibility through inheritance rather than round after round of interop testing.

Commercially, this is a smart move. Lightmatter isn't fighting NVIDIA and Broadcom over "whose scale-up protocol wins"; it's going after "what the shared mechanical and optical layer looks like." Protocol wars are zero-sum; infrastructure wars are positive-sum — and the latter happens to be exactly what Taiwan's supply chain does best.


6. Targets and timeline: not a vision, but a delivery date

What makes this proposal least like a slide deck is that it gives numbers it can be held to:

  • Single-tier domain size: 72 XPUs today, target 1,024+ XPUs (the white paper mentions up to 2,048, and 100K-class in the long run).

  • MFU: 38–43% today, target 55–65%.

  • Energy efficiency: CPO today is about 5–7 pJ/bit; target below 5 pJ/bit (3 pJ/bit is aspirational).

  • Network share of power: from about 31% down to about 15%.

  • 1U switching capacity: design target of 200T and above, which the speaker called the only current path to a 1U 400T.

The timeline is tight too: announced in March with 10 founding members; on August 13, the day of OCP APAC, a roughly 300-page white paper, "Architecture Vision: Open Silicon Photonics for AI Systems," was released and membership grew to 19; the first formal specs are targeted for Q4 2026. The list includes Celestica, Corning, Dell, Flex, Foxconn Interconnect Technology, Hyve Solutions, Keysight, Qualcomm, Quanta Cloud Technology and GUC.

Nowroozi also noted on stage that more than 50 APAC suppliers are already building AI MHS systems — a line aimed at the audience, which translates to: you're already building the hardware; what's missing is the document that defines the interfaces.

7. Counterarguments: three things that could stall this

It can't all be upside. The proposal carries three clear risks.

First, the missing names matter more than the ones present. The list has no NVIDIA, no Broadcom, and none of the hyperscalers themselves. With the two companies that account for the vast majority of today's scale-up shipments absent, there's a high chance this spec becomes "the common language of the non-NVIDIA camp" rather than an industry standard. That's nothing new in the history of networking standards.

Second, test methodology stands between the white paper and volume production. At the same OCP APAC, representatives from ASE and TSMC both pointed out that wafer-level optical test standards and shared simulation methodology remain unresolved. Without common test methods and models, the interface contract is only compatibility on paper. I think closing this gap will take at least another year.

Third, the window is shorter than it looks. Pluggables can still hold up for the 49.6T switch generation, but the industry broadly expects to hit the density ceiling in about 18 months. Spec in Q4 2026, multi-vendor products to buy no earlier than 2027, a real market in 2027–2028. A one-quarter slip in the spec effectively cuts a third off that window.

Conclusion

What this keynote really announced is not "optics will replace copper" — that was settled the moment 448G left copper with only 25 cm. What it announced is: CPO's technical risk has come down far enough to talk about an ecosystem; the next competition is over who defines the interface document.

For Taiwan's supply chain, three things are worth looking at separately:

First, the question of position. FIT, QCT and GUC are all on the list, but in ODM, system integration and ASIC design-service seats. These seats come with volume and margin, but no voice — someone else writes the spec and you build to it. Joining the workgroup and joining the supply chain are two different things; the former only costs person-hours.

Second, the window for new part numbers. Passive optical infrastructure (faceplates, ODFs, raceways, ELS enclosures) is the segment this white paper opens up that has no established winner yet. It needs no silicon photonics process and no CW lasers; it needs mechanical design, fiber management, connector precision and serviceability — exactly Taiwan's strengths, and the only segment where vendors can get involved from the spec-definition stage. For connector, enclosure and cable makers, sending people in now costs a few person-months.

Third, the hidden thread of test. If wafer-level optical test and shared simulation methodology really are the bottleneck, the beneficiaries won't just be module makers. From probe cards and alignment equipment to optical metrology, demand in this segment will be forced out by the spec — and it will arrive a year ahead of CPO's own ramp.

Whether that Q4 2026 spec ships is the only thing to watch this coming quarter. If it does, multi-vendor CPO will be available to buy in 2027; if not, pluggables get one more generation — and the price of that generation is continuing to leave half the compute waiting on the network.

This article is for technology and industry trend analysis only and does not constitute investment advice.



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page