Why Ethernet matters for AI networking
I have worked on Ethernet for most of my career. I saw it in its infancy and have watched, and occasionally helped shape, its growth. It has a simplicity and an openness that the world has embraced. Now it is being called on to meet the voracious networking demands of AI, and the foundation's members are answering that call.
AI's growing appetite: more compute, power, cooling, and a new network
Everyone agrees that AI is reshaping industries and daily life. The transformation has a price. For anyone providing AI services it means far more compute, far more power, far more cooling, and, crucially, a different approach to networking.
Whether you run a private AI cloud for an enterprise or sell GPU-as-a-Service, training and inference force you to rethink the data center network. Modern AI jobs run across hundreds, thousands, or tens of thousands of GPUs that must behave as one tightly coupled cluster. That happens in the back-end scale-out network, where zero packet loss and ultra-low latency are non-negotiable.
From InfiniBand to Ethernet
For years InfiniBand was the default for AI and HPC. Its protocol guarantees end-to-end delivery, providing the lossless fabric early clusters needed. Today the industry is shifting to Ethernet, because Ethernet has been proven for decades in the world's largest and most demanding networks and brings advantages that fit the evolving AI landscape:
- Broad ecosystem. Switches, NICs, test gear, optics, and open-source management platforms are all readily available.
- Fast evolution. New link speeds, optics, cabling, and protocol enhancements arrive regularly.
- Universal familiarity. Engineers everywhere understand Ethernet, which makes deployment and troubleshooting smoother.
- Scale with IP. Ethernet plus IP grows arbitrarily to support very large networks.
- Open, multivendor flexibility. No single-vendor lock-in; mix and match the components that fit your design.
These traits have made Ethernet a viable, often preferable alternative to InfiniBand for hyperscalers, cloud providers, and enterprises building AI-focused data centers.
How today's Ethernet powers AI back-end networks
The current Ethernet-based scale-out solution keeps the proven InfiniBand transport and encapsulates it in UDP, IP, and Ethernet: RDMA over Converged Ethernet, or RoCEv2. To keep loss at bay, RoCEv2 relies on DCQCN, which blends Explicit Congestion Notification, marking packets as queues fill to tell senders to slow down, with Priority Flow Control, which pauses traffic on specific priority lanes when thresholds are crossed. Both activate when leaf or spine queues pass predefined limits, preserving the lossless environment training demands.
Pushing Ethernet further
RoCEv2 works well for many deployments, but it struggles in massive AI clusters with complex topologies and bursty traffic. Head-of-line blocking from PFC and the absence of real-time congestion signaling can make the network feel fragile.
This is where the foundation comes in. NoLimit is modernizing RDMA into a performant, open, interoperable, Ethernet-based full-communications stack designed specifically for AI and HPC at scale. Architecture 1.0, released in May, adds advanced load balancing, refined congestion control, built-in security, and richer API support. The focus remains the scale-out portion of the AI back end, where the biggest gains are needed. In short, the foundation is addressing the exact pain points that keep operators awake when they watch an AI job stall.
Members putting it into practice
Several member companies with deep Ethernet heritage are already building on the 1.0 architecture, and early internal tests have shown NoLimit Transport traffic running successfully on production switching platforms. The same programmability that lets operators automate data center operations is what lets them experiment with new congestion-control schemes, which is precisely the flexibility the architecture asks for.
Closing thoughts
AI is rewriting the rules for what data center networks must deliver. By embracing Ethernet's ecosystem, its rapid innovation, and the forward-looking work of the foundation, we can build networks that keep pace with ever-growing AI demands. If you are wrestling with how Ethernet can accelerate your AI projects, we would like to hear about the challenges you are facing. Reach out.