March 24, 2026 ยท Everett Caldwell

The collaboration that will carry Ethernet into the HPC and AI future

Any time you can get a large number of companies full of technically adept and strongly opinionated people to work together on a problem, you know for certain that the problem is real. The many problems with Ethernet at scale are what brought the NoLimit Foundation into being last October, and the foundation now has well over a hundred participating organizations working to deliver its first architecture release.

The steering members all have skin in the networking game, and in the HPC and AI systems space where Ethernet needs to be bolstered and extended to make up for the limitations of both current Ethernet and the InfiniBand fabrics that have carried capability-class systems for the past decade. AMD, Arista, Broadcom, Cisco, Eviden, HPE, Intel, Meta, Microsoft, and Oracle formed the core. NVIDIA was not an original member but joined this year to put its weight behind the ideas being worked out for Ethernet networks with multiple terabits per port that can scale past a million endpoints. In most cases the endpoint in question is a vector or tensor accelerator doing the arithmetic for AI training.

I have spent much of the past few months explaining what the foundation is and how the work is progressing, most recently in a two-part conversation with the trade press. A few themes came up repeatedly and are worth writing down.

Breaking down the layers

The OSI model is useful, but the organizations that maintain the technologies in each layer rarely sit in the same room. A transport improvement that assumes a certain link behavior, or a link change that ignores what the software above it needs, tends to die on contact with reality. The foundation exists partly to put the people who own each layer at one table and force the conversation to happen before the silicon is taped out rather than after.

How consensus happens

The process is deliberately unglamorous. Working groups propose drafts, a technical advisory committee ensures the drafts fit together as a coherent whole, and milestones are set only when there is agreement that the pieces can actually be implemented. This is slower than one company shipping a proprietary answer. It is also the only way to end up with something that many vendors can build and many operators can trust.

What we are not doing

At least initially, we are not trying to reinvent everything. Creating a more extensible Ethernet without breaking compatibility with the past is a hard enough task on its own. The focus is the scale-out fabric behind AI and HPC clusters, where the pain is sharpest and the payoff is largest. Anything beyond that is a question for later releases.

The full conversation covers more ground than a blog post can, but the short version is this: a lot of serious people have decided that Ethernet is worth fixing rather than replacing, and they are spending their own engineers' time to do it. That is the surest sign the problem is real.