OCP Global Summit 2025: Meta on Scaling AI Infrastructure to Data Center Regions
Introduction
At OCP Global Summit 2025, Meta shared its latest challenges and solutions in scaling AI infrastructure. As a platform serving 3.4 billion users, Meta must not only make AI deliver value across a wide range of applications but also absorb the infrastructure pressure that comes with that scale. From the evolution of its LLaMA models to building giant clusters that span data centers, the talk offered a full picture of how hard — and how inventive — infrastructure scaling has become in the AI era.
Key Takeaways
Meta began by stressing that AI is now woven into every part of its platform, from content integrity and ads ranking to recommendation systems. But with the explosion of large language models (LLMs), traditional ways of scaling no longer suffice; infrastructure has to advance at far greater speed and scale.
Its infrastructure evolution has been extreme — from small early GPU clusters to super-clusters spanning entire data centers. Every doubling of GPU count brought sweeping hardware and software challenges. Meta even introduced the concept of "server karma" to predict and screen out servers prone to failure, keeping training jobs stable.
On the networking side, Meta described distinct paths for scale-out, scale-in and scale-up. Scale-out can grow quickly thanks to mature Ethernet standards; scale-in is still dominated by proprietary protocols but will move toward openness; and scale-up has become critical for AI workloads, because dense compute and Mixture-of-Experts models both need large, low-latency domains. That is why Meta is driving standardization through Ethernet for Scale-up Networking (ESUN) and the UEC (Ultra Ethernet Consortium).
On the build-out side, Meta showcased two landmark data center projects:
Prometheus (New Albany, Ohio): a 1GW-plus cluster built for flexibility, even using temporary tent structures to add capacity quickly.
Hyperion (Richland Parish, Louisiana): a 5GW mega single-site data center designed from scratch, with a footprint comparable to the walk from lower Manhattan to the north end of Central Park.
Another challenge is hardware diversity. For supply chain resilience and performance optimization, Meta has had to bring in multiple types of accelerators and servers, which demands higher levels of abstraction from software developers. Meta also presented Open Rack Wide and its collaboration with AMD Helios, noting that going forward it will need optical disaggregation to break past copper limits and enable larger scale-up domains.
Finally, the talk highlighted three major pain points:
Insufficient power supply
Shortage of specialized talent
Sustainability challenges
Conclusion
The talk made it clear that AI infrastructure has moved from "massive" to "insane." Through real-world cases, Meta showed that AI training's demands on hardware, networking, power and people far exceed anything previously imagined.
STT Perspective
Technology impact: Meta's optical disaggregation is closely tied to silicon photonics. To break the copper bottleneck, future high-speed modules will inevitably rely on the low-power characteristics of SiPh.
Supply chain: From server karma to diverse accelerators, chip designers and module makers (Broadcom, Marvell, Credo) will play key roles in this transition.
Market trend: A single company investing in 5GW-class sites means data center CAPEX is entering the tens-of-billions-of-dollars range. AI is no longer a research application but the core driver of infrastructure build-out.




Comments