Open Compute Project Foundation and NoLimit announce a new collaboration
Today the Open Compute Project Foundation, the nonprofit that brings hyperscale innovation to everyone, and the NoLimit Foundation announced a collaboration to improve Ethernet performance for next-generation AI clusters and HPC. NoLimit is developing enhancements to Ethernet. OCP has an active community building sustainable, large-scale computational infrastructure for AI and HPC on Ethernet. Together the two expect to integrate enhanced Ethernet into the next generation of OCP community-delivered AI clusters, providing the low-latency connectivity that back-end AI fabrics need.
The money argument is plain. Adoption of AI in the workplace, including for automating data center operations, has created a perfect storm of IT equipment investment. Major analyst firms have revised capex forecasts upward on the strength of hyperscale plans for large AI clusters. By collaborating, NoLimit and the OCP community's system specifications can have a greater influence on how sustainable, large-scale AI clusters address the memory size and connectivity bandwidth challenges posed by large language models.
The collaboration will align work inside both organizations, focus effort on shared objectives, and make sure OCP's integration of NoLimit's Ethernet enhancements is smooth and effective. Initial areas identified for exploration include the OCP Switch Abstraction Interface (SAI), the Caliptra workstream, the OCP Networking Project, the OCP NIC workstream, the Time Appliance Project, and the Future Technologies Initiative.
AI and HPC workloads present new challenges for the network: greater scale, more bandwidth, multipathing, and faster reaction to congestion. NoLimit was formed to develop Ethernet specifications that meet those requirements as the workloads grow, a shift in the nature of network traffic that may prove as significant as the move from voice to video. Working with the OCP community fast-tracks the integration of those enhancements into complete systems that can serve the market.
Analysts watching the space see the same pressure from the other side. The rollout of generative AI workloads and AI-assisted applications is accelerating, and it will strain cloud and data center networks that must provide the interconnect bandwidth for training and inference inside those clusters. A community that spans both organizations can develop the required enhancements and then embed them in systems that improve cluster performance. Taking a fresh look at Ethernet in the context of large-scale AI deployment has the potential to push the whole industry forward.
This is not a merger. It is a decision that Ethernet for AI should be designed once and then dropped into the machines operators already know how to buy.