top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

2026 OCP APAC Summit | Broadcom | Bhaskar Chinni | CPO's Third-Generation Report Card: Broadcom Shifts the Pitch from "Saves Power" to "Doesn't Break"

2 days ago
8 min read

In its CPO session at OCP 2026 APAC, Broadcom did something very un-"marketing": it spent almost no time explaining how co-packaged optics (CPO) works, and instead laid out reliability data from three product generations running at customer sites. The 65% power saving was still on the slides, but the numbers speaker Bhaskar Chinni really wanted to drive home were three others: MTBF in the millions of hours, more than 4x the time between failures compared with pluggable optical modules, and a 10x reliability improvement in large data-center deployments.

This marks a turn in the CPO narrative. For the past three years CPO was sold on "power savings," because that's the easiest pitch to grasp; but what has actually kept CPO out of the data hall was never whether it saves power, it was how you fix it when it breaks. When Broadcom starts telling its story with failure data reported back by customers, the main battleground of the debate has moved from "is it worth switching" to "is it easy to live with after switching."

At the same event, another Broadcom keynote covered how Ethernet is starting to eat scale-up, aimed squarely at NVLink. The two sessions only make full sense together: one is about whose turf the protocol wants to invade, and this one is about how the physical layer holds up once it gets there.

1. Opening Jab: CPO Wasn't Invented Yesterday

Chinni put it bluntly: "There's a lot of buzz about CPO in the industry right now, but Broadcom has already shipped three generations of CPO products, and there will be a fourth in 2027. Some companies talk about CPO as if they invented it."

That line is worth pausing on. It's not just vendor trash talk; it's a timeline declaration: Broadcom wants to close the question of "is CPO a volume-production technology" outright. Three generations shipped and a fourth slated for 2027, corresponding to its two CPO platforms, Tomahawk 5 Bailly and Tomahawk 6 Davisson. When a company starts talking about a technology in terms of "which generation" rather than "when," it's pulling the conversation out of the lab and onto the BOM.

For Taiwan's supply chain, this timeline matters more than any spec: if the fourth generation lands in 2027, the design freeze for the corresponding optical engines, external light sources, FAUs and connectors happens in 2026.

2. Three Reliability Numbers More Worth Writing Down Than 65% Power Savings

The power number on the slide is "65% power savings versus retimed optics." That figure has been repeated to the point of numbness over the past two years, and Chinni himself breezed past it with "the previous slide explained why." What he really spent time on were these:

  • MTBF in the millions of hours, and he stressed that "this data isn't from our own testing, it's reported by partners and customers"

  • More than 4x MTBF improvement compared with pluggable optical modules

  • An overall 10x reliability improvement in large data-center deployments

  • Even more striking: one partner with a large-scale deployment reported that failure events were so rare there weren't "enough to calculate a meaningful MTBF"

That last line is the point. Reliability statistics require enough failure samples to compute; when a customer says "I can't collect enough samples," that's not a statistics problem, it means product failures are no longer on the radar.

One limitation must be flagged honestly: the live transcript had recognition errors on several numbers (for example, passages with clearly anomalous magnitudes like "50 million hours"), so this article uses only numbers the speaker repeated and that matched the Q&A. Before publishing, we recommend checking hard numbers against Broadcom's official slides or later public OCP materials.


3. Why Reliability Is a Financial Problem for AI Training, Not an Operations Problem

The logic Chinni used to connect reliability to training efficiency is the single most valuable takeaway from the session.

Collective communications in AI training use a go-back-N-style retransmission mechanism. Once a link flaps or drops packets, the cost isn't retransmitting "that one packet," but rolling back and resending the whole batch. In synchronous training, that means every other GPU in the cluster sits idle: you're paying for a full rack's machine time to absorb the penalty of one jittery optical link.

Fewer link errors → fewer retransmissions → fewer global go-back-N rollbacks → higher training efficiency.

Broadcom's figure is roughly 90% training efficiency observed under the CPO architecture, compared against large clusters with a mix of various pluggable optics. An audience member pressed in the Q&A on where the 90% came from, and Chinni's answer was simple: it's converted directly from the 4x MTBF improvement. Fewer failures, and efficiency naturally rises.

In other words, the 90% isn't an independently measured new metric, but another way of expressing the reliability number. This is important to understand: it's a derived value, not a measured one, so don't treat them as two independent pieces of evidence when reading the data.

Derived or not, the direction is right. This conversion of "optical link stability = training cost" is turning optical interconnect from a networking-team procurement line item into a performance variable for training teams. We discussed the trade-offs of each path in After Copper Can't Keep Up with AI: Seven Paths for Scale-Up Optical Interconnect, and reliability is the most underrated column in that table.

4. The Real Turning Point: Scale-Up Goes from One Rack to Many

This is where the talk's technical pivot lies. Chinni's chain of argument is concise: the scale-up domain is growing from a single rack to multiple racks. Once you cross racks, copper's reach isn't enough; and if you force active copper cables to cover it, rising cost and cable density eat away copper's cost advantage over optics. In his words, it "becomes comparable to optical solutions."

As scale-up domains grow from hundreds of XPUs to thousands, optical interconnect goes from "optional" to "mandatory." This is the flip side of SemiAnalysis's argument at the same event that copper vs. optics is a false dichotomy: it's not about which replaces which; once scale passes a certain threshold, the choice simply disappears.

Broadcom's answer is called OCI (Optical Compute Interconnect), a new specification body it co-founded with other partners, which has already published a white paper. The goal is stated clearly: extend scale-up from a single rack to multiple racks, supporting thousands of GPUs/XPUs.

OCI's technical choices are interesting; nearly every one leans toward "cheap, rugged, low-latency":

  • NRZ modulation: chosen for efficiency and power; the Q&A confirmed no FEC is needed, eliminating FEC latency and power outright

  • WDM + standard lasers: uses the industry's existing standard light sources, avoiding the supply and cost risks of custom lasers

  • Bidirectional fiber: runs both directions on the same fiber, cutting fiber count in half

Dropping FEC is a big message for anyone building scale-up. FEC is the most annoying piece of the latency budget; scale-up demands memory-semantic-level low latency, and being able to remove FEC shows OCI's link budget is deliberately designed for "short reach, high quality, zero error correction." That also explains why it dares to use NRZ rather than squeezing into PAM4 or higher-order modulation.

As for reach, when pressed by the audience Chinni said the standard had just launched and reach wasn't finalized, but Broadcom's CTO, who was present, chimed in directly: 1 to 2 km. For Taiwan's connector and fiber-cable makers, that number is more useful than any architecture diagram: in one stroke it extends scale-up's physical boundary from inside the rack to between data-center floors.

5. The 1K-XPU Reference Design and Optical Backplane: The Shape of the Rack Is About to Change

Chinni presented a reference architecture: a scale-up cluster of about 1,000 XPUs, made up of 16 XPU groups linked through 4 ultra-high-capacity switching devices.

This needs careful citation. On stage he used a switch chip with far higher capacity than the current generation for the illustrative calculation, while emphasizing that Broadcom is currently the only vendor shipping a 100T-class switch chip in volume (Tomahawk 6, 102.4 Tbps). In other words, the 1K-XPU diagram is drawn with next-generation capacity and is not a SKU shippable today. The transcript's capacity figures for this section were unclear, so this article does not cite specific Tbps values; we recommend relying on the official slides.

More noteworthy than the reference design is what he mentioned in passing: the optical backplane.

In the architecture he described, both compute racks and switch racks "plug into" an optical backplane. The implication is heavier than it sounds: if the connection interface moves from fiber patch cords on the rack front panel to blind-mate optical connections on the backplane, the entire connector and passive component value chain gets repriced. What's needed is blind-mate alignment, high-density MT/MPO interfaces and mechanical tolerance control, not just a patch cord.

This is also why the pluggable camp is still working to extend its life: Arista's XPO proposal of immersing optical modules in liquid cooling essentially says "don't change the rack shape yet." Broadcom's session stands on the opposite side: the rack shape should change, and the optical backplane is where it's heading.

6. What It Means for Taiwan's Supply Chain: This Talk Opens Three Orders

Translated into terms Taiwanese vendors understand, the talk boils down to three things.

First, reliability will become a hard metric on the spec sheet. Optical module MTBF used to sit on the last page of the datasheet, unread; once customers start back-calculating link failure rates from "training efficiency," MTBF moves to the front of purchasing decisions. That's a plus for Taiwanese vendors capable of long-duration reliability qualification, failure analysis (FA) and lot-to-lot consistency control. This is a business of manufacturing discipline, not design capability, which is exactly Taiwan's home turf.

Second, OCI's bill of materials is relatively friendly. NRZ, no FEC, standard WDM lasers, bidirectional fiber, 1–2 km reach: none of it requires betting on new materials or new processes. It needs CW laser sources, low-cost WDM filtering, high-density connectors and fiber management, all within Taiwan's existing capabilities. By contrast, the 400G/lane path is limited by electronics, not optics, and the barrier there is much higher.

Third, the optical backplane is a new strategic position, and the window is 2026. If fourth-generation CPO lands in 2027, the mechanical and alignment specs for backplane-level optical connections must be settled within 2026. This layer has no clear winner yet; it's one of the few segments where "it's not too late to get in now."

The risks must be stated clearly too: CPO's biggest unsolved problem is still field serviceability. When a pluggable fails you swap the module; when CPO fails you're touching the whole switch. Broadcom's answer is "reliability so high you won't need to fix it." That answer holds up in the data, but it's a statistical argument, not an operational process. Until there is large-scale, long-duration independent third-party data, this remains a gap customers have to fill with trust.

Conclusion

The real signal of this talk isn't the 65% or the 10x, but that Broadcom no longer needs to convince you CPO is viable; it's convincing you CPO is easy to live with.

A technology's narrative usually shifts from "performance advantage" to "operational fitness" at only one point in time: when it's certain to enter volume production. Add OCI, a spec framework that extends scale-up to multiple racks, 1–2 km and no FEC, plus a 2027 fourth-generation timeline, and Broadcom effectively handed the market both a technology roadmap and a timetable.

For Taiwan's supply chain, there are only three things to watch over the next 12 months: when OCI white paper reach and connector specs are finalized, when fourth-generation CPO optical engines and external light sources start sampling, and who leads the mechanical specs for the optical backplane. The first two decide who gets on board; the third decides who gets a front-row seat.

This article is for technology and industry trend analysis only and does not constitute investment advice.

Related Reading



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page