top of page

📢 STT 訂閱專區已上線

免費文章會照常更新,一篇都不會少。訂閱是「加強版」——每週深度週評、財報法說的完整判讀、所有長篇深度報告全包。

免費讓你跟上,訂閱讓你看懂、能做判斷。

月訂 NT$199|年訂 NT$2,000(約 NT$167/月)
👉 立即訂閱: vocus.cc/salon/simpletechtrend

OIF Draws the Official Map for AI Interconnect: Three Networks, One pJ/bit Battlefield, and CPO as the Written Endgame

2 days ago
8 min read

  • OIF's "System Vendor Requirements for Energy Efficient Interfaces" formally splits AI data center interconnect into three networks: front-end (FEI), back-end (BEI), and compute (CI). Their priorities, reach, and interoperability requirements are completely different, so they can no longer be measured with the same yardstick.

  • The real battlefield is energy. FEI holds at 10 pJ/bit, while short-reach, high-density scenarios like BEI and CI are being pushed by system vendors to 5 or even 3 pJ/bit. With the industry sitting at roughly 15 pJ/bit today, that means cutting per-bit power by two-thirds to four-fifths.

  • OIF's next steps are already locked in: BEI moves first in the short term (already ramping in AI clusters, riding the Ethernet supply chain, solved with bookended links), CI follows in the mid term (scale-up, CXL/PCIe domains, a cost-driven knife fight), and the form-factor endgame points to linear and co-packaged optics (CPO).


1. Why a "Requirements Document" Deserves a Careful Read

First, be clear about what this document is: not a specification, but a requirements list in which system vendors spell out exactly what they want. The technical editor is from Huawei and the working group chair from Juniper, but the references that actually shaped the content come from Microsoft, Meta, NVIDIA, and HPE. In other words, this is the target the buyers (system vendors and hyperscalers) drew for the optical supply chain after eight months of interviews.

What gives it weight: OIF will use it to launch Implementation Agreements (IAs) — the standards that future optical engines, optical modules, and laser sources will have to align with. Whoever reads the bullseye first gets into position first.

The pain point driving all of this is simple: AI clusters have blown up interconnect demand. Once a cluster has to scale out to 1,000–10,000 nodes, the era of "one optical module for everything" is over. The document opens by naming the three networks — exactly the thread STT has been following.

Figure 1: The three networks of modern high-performance networking — FEI, BEI, CI | Source: OIF-EEI-Requirements-RD-01.0 - Figure 1
Figure 1: The three networks of modern high-performance networking — FEI, BEI, CI | Source: OIF-EEI-Requirements-RD-01.0 - Figure 1

2. Three Networks, Three Yardsticks: What FEI / BEI / CI Each Need

The difference between these three networks isn't a matter of "a bit faster or slower" — their priorities are flipped entirely.

FEI (front-end, intra-datacenter): Covers server-to-server, server-to-storage, and server-to-edge traffic — how the data center talks to the outside world. It's already deployed (NOW), with the market growing from US$2.14 billion in 2023 to US$4.39 billion in 2027 (about 2x). Its top priority isn't power but standardized interoperability with forward and backward compatibility; energy efficiency ranks only second. The reason is pragmatic: FEI has already been well polished by IEEE, and there's limited room left to squeeze out efficiency.

BEI (back-end, GPU-to-GPU scale-out): The dedicated accelerator-to-accelerator network inside an AI cluster, supporting 1,000–10,000 nodes. The market jumps from US$1.32 billion in 2023 to US$3.88 billion in 2027 — about 3x, the fastest of the three. Its top priority is latency/throughput and low error rate (directly tied to training efficiency); energy efficiency ranks only third. Today it's dominated by proprietary implementations like InfiniBand and HPE Slingshot, and the document explicitly says a major overhaul is coming in 2024–2025.

CI (compute, scale-up, intra-node): Packs a large number of GPUs into a single node (possibly spanning racks), interconnected over a "single CXL domain" network such as CXL/PCIe. The most counterintuitive part: its top priority is cost (it must support N×10 Tbps of aggregate bandwidth per accelerator), and optical interfaces won't arrive until 2025–2027.

In one line: BEI is scale-out, CI is scale-up. We broke down this framing in The Great Optical Packaging Transition (Part 2): CPO's Three-Stage Evolution — Why OBO Died, Scale-Out Moves First, and Scale-Up Is the Endgame; this OIF document is effectively an official endorsement of that framework.


3. The Real Battlefield: From ~15 Down to 5, Even 3 pJ/bit

If you remember only one number from this document, make it energy.

FEI's consensus baseline is 10 pJ/bit. It can hold this relatively loose threshold because reach varies widely (0–500 m, in some cases up to 2,000 m), so margin is needed.

But in short-reach (typically 10–100 m), high-density scenarios like BEI and CI, the target drops straight to 5 pJ/bit, with some operators calling for 3 pJ/bit. Compared with the "Today" mark in Figure 8 — roughly 15 pJ/bit — BEI/CI must cut per-bit power by two-thirds to four-fifths.

Figure 8: Relative energy-efficiency requirements by application — FEI ~10 pJ/bit, BEI/CI target 5 pJ/bit, today ~15 pJ/bit | Source: OIF-EEI-Requirements-RD-01.0 - Figure 8
Figure 8: Relative energy-efficiency requirements by application — FEI ~10 pJ/bit, BEI/CI target 5 pJ/bit, today ~15 pJ/bit | Source: OIF-EEI-Requirements-RD-01.0 - Figure 8

Why is energy the lifeline rather than just an ESG slogan? Because lower power directly buys higher density — only by pushing power down can you fit more on the faceplate and dare to move optics deep into the system to capture system-level benefits. That is exactly why CPO exists.

Appendix C is even more aggressive than the main text: FEI targets <10 pJ/bit (OE + laser), while BEI targets <<4 pJ/bit (OE + laser). That "<<" is the system vendors' real wish list. Reaching that level hinges on bringing the light source into the package and slashing opto-electronic conversion energy — NVIDIA has demonstrated this path with DWDM laser arrays, which we broke down in Light Shouldn't Be Forced Into Electrons and Back: Why NVIDIA's DWDM Laser Array Rewrites the Economics of Optical Interconnect.


4. Latency, Interoperability, Reach: Why BEI Is the Sweet Spot for Trading Relaxation for Performance

Stack latency, interoperability, and reach together, and BEI emerges as a deliberately open sweet spot.

Latency: The document only addresses PMD (physical medium dependent) latency and gives formulas — FEI is PMD Latency < 20ns + d×5ns/m + FEC delay, BEI is PMD Latency < 5ns + d×5ns/m + FEC delay, cutting the baseline by 15ns. Why is latency so critical for BEI? NVIDIA data (Reference 6) shows that in an AI cluster, every additional microsecond of latency cuts throughput by 25%. On reach, BEI tops out at 300 m (matching the feasible size of an AI cluster machine), and CI is even shorter at just 7–10 m (within a rack or adjacent racks).

Interoperability: Three tiers — Type-1 (single-vendor bookended), Type-2 (multi-vendor, same generation), Type-3 (multi-vendor, cross-generation with backward compatibility). The key insight: FEI demands the highest interoperability, CI is moderate, and BEI is actually the lowest. That's because BEI is typically deployed rack-by-rack by a single hyperscaler or system vendor, so bookended links are enough — no need to interoperate with anyone else.

Put these together and you get BEI's sweet spot:

Short reach and loose interoperability aren't constraints — they're chips BEI can trade for power and latency.

The document is blunt: although IEEE's FEI baseline can work for BEI, users expect the industry to "leverage the reduced reach requirements to gain power and latency benefits for BEI." In plain terms — don't just copy FEI specs over; be aggressive where it pays.


5. The Form-Factor Endgame: Pluggable → Linear → CPO, and the Roadmap Is Already Written

The document measures density in Tb/s/mm² (area) or Tb/s/mm (linear, along the faceplate edge), with a target of >1 Tb/s/mm.

For now, system vendors still prefer FPP (Faceplate Pluggable) form factors — OSFP, QSFP-DD — which are modular, hot-swappable, and backed by a mature, cheap supply chain. But once fan-out approaches 200G/lane and beyond, signal integrity and power force a move to NPO (near-packaged optics) and CPO. CPO's value lies in shortening the chip-to-module path and eliminating several connectors, trading distance for signal integrity and efficiency — and it makes linear technology easier to adopt at high data rates.

The document even lays out the evolution roadmap: server-to-switch goes "electrical → pluggable optics → pluggable on one end, CPO on the other → CPO on both ends," possibly with a linear phase in between; switch-to-switch goes "pluggable → CPO."

But don't mistake a roadmap for a timetable. Large-scale CPO ramp keeps getting pushed out, which in turn extends the shelf life of pluggable optical modules — we broke down this timing gap in CPO Volume Ramp Pushed to 2028, But It's Not Bearish: What Gets Extended Is the Shelf Life of Pluggable Optical Modules. OIF's roadmap is a steering wheel, not tomorrow's itinerary.


6. OIF's Next Step: BEI Moves First, CI Follows in the Mid Term

Chapter 4 spells out the priorities unambiguously.

BEI (short term): This is the piece that moves now — optical interconnect is already ramping fast in AI clusters, riding the existing Ethernet supply chain, and solved with bookended links. OIF will open IAs to drive common solutions, with four standardization priorities: lower power, lower latency, better link quality (lower raw BER → weaker FEC → higher throughput; shifting from margin-based link budgets to actively optimized links), and interface densification (first pushing module throughput in Tb/s/mm², then moving on-package).

CI (mid term): Scale-up, cost-sensitive, ultra-short reach — much of the technology will be carried over directly from BEI with shorter reach. Three priorities: lower cost (module architecture, lower packaging and test costs, consolidating multiple links into one module), energy efficiency, and lower latency (inside a rack-scale node, optical interconnect is the bus linking accelerators, memory, and CPUs).

The document also plants two threads worth tracking. First, photonic switches entering the network — millisecond-scale commercial switching already has precedents (Google Jupiter, TPU v4), but it eats into the optical link's loss budget, effectively adding an extra span of fiber and in turn demanding more optical power and FEC gain. Second, industry terminology is clashing: the electrical-bus world and the optical-interconnect world use different terms for the same concepts, and OIF wants to establish a common vocabulary. It looks trivial, but it's a prerequisite for standardization to converge.


7. Risks and Counterarguments: Where This Map Could Go Wrong

No matter how cleanly a requirements document is drawn, there's friction on the ground. A few counter-signals to watch:

First, the politics of interoperability. Low interoperability in BEI is convenient for a single hyperscaler, but it also means weak incentives to standardize. If every major GPU vendor runs its own bookended solution, will OIF's IA end up as a consensus document nobody follows?

Second, the physical bottleneck of memory disaggregation. CI wants memory pooling, but the document itself admits that heavy disaggregation slows memory access — and access speed is the lifeline of system throughput. Scale-up optical interconnect can't simply be switched on at will.

Third, the optimism baked into the pJ/bit targets. <<4 pJ/bit is an aspiration, not something achievable today. Laser source efficiency, coupling loss, and packaging yield are each betting against this target.

Fourth, timelines slip. CI optical interfaces are pegged at 2025–2027, and CPO on both ends is even further out ("beyond"). Real-world timelines (see the 2028 analysis above) keep moving later.


Conclusion

The value of this document isn't any single number, but one thing it does: formally divide AI interconnect into three networks and set priorities for each.

For anyone building optical engines, optical modules, or laser sources, the answers to which network your product targets, which pJ/bit bullseye to aim for, and which interoperability tier to support are all in this document. BEI is the first target with real money and volume — 3x growth, short-term priority, and tolerance for relaxing IEEE specs in favor of aggressive designs. CI is the cost battleground that comes next.

The takeaway: whoever can push pJ/bit from 15 down to single digits fastest on BEI — the network with short reach, loose interoperability, and room for aggressive design — wins first-mover position in AI scale-out optical interconnect.


References

  • OIF, "System Vendor Requirements Document for Energy Efficient Interfaces," OIF-EEI-Requirements-RD-01.0, 2024-05-09 (Technical Editor: Eric Bernier / Huawei; Working Group Chair: Jeffery Maki / Juniper)

  • C. Thompson, NVIDIA Motivation for Energy Efficient Interfaces, oif2023.270.00 (source for AI cluster latency impact on throughput and DGX H100 cluster data)

  • R. Huggahalli, Microsoft Use Cases, Optical Connectivity for AI Clusters, OCP 2023

  • D. Alduino et al., Energy Efficient Interfaces – Meta Perspective, oif2023.273.01

  • L. Poutievski et al., Jupiter Evolving, SIGCOMM'22; N. P. Jouppi et al., TPU v4, ISCA'23 (commercial photonic switch cases)


Related Reading


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page