AI networking startups race to replace Nvidia's NVLink
As Nvidia expands its influence through its NVLink Fusion tech, rival networking vendors are scrambling to bring alternative interconnects and switches to market.
At the AI Infra Summit this week, Delos Data and Cornelis Networks officially entered the scale up networking race.
Scale-up fabrics, like NVLink, are what have allowed Nvidia to make eight, 72, and now 576 GPUs behave as one enormous AI accelerator. To catch up, rivals like AMD have embraced emerging protocols like Ultra Accelerator Link.
Today, these protocols are largely being tunneled over standard Ethernet switches. For instance, AMD is using Broadcom’s 102.4 Tbps Tomahawk 6-based switches connecting to custom I/O dies on the MI455X.
Purpose-built UALink switches and physical interconnects remain elusive, but that won’t be the case for long if Cornelis and Delos have their way.
The two companies are approaching this challenge from a few different angles, including standardization, software optimization, and physical hardware.
Setting the standard for the Never-Nvidia network
At AI Infra on Monday, HPC-centric networking vendor Cornelis introduced the Active Compute Fabric (ACF), which seeks to establish an open architecture for scale up and scale out networking that integrates programmable compute into the fabric.
The standard signals Cornelis’ entry into the scale up networking arena. Spun out of Intel in 2020, Cornelis’ Omni-Path tech was originally designed as a scale-out interconnect for high-performance computing applications, including supercomputers like Trinity and Lynx.
With the imminent launch of the company’s 800 Gbps-capable CN6000-series switches and NICs, the company has its sights set not only on bringing its tech to a broader audience through the open Ultra Ethernet protocol, but also on scale-up networking.
ACF expands on the mission of technologies like UALink and the Ethernet for Scale-Up Networking, another scale up networking protocol, to set a baseline for in-network compute capabilities across a wide range of hardware, not just Cornelis’ own CN-series parts.
One of the key technologies behind Nvidia’s NVSwitch ASICs is support for SHARP, which allows for things like in-network collectives to be offloaded to the switch ASICs, freeing up GPU compute and cutting down on communication overheads.
Collective acceleration isn't new by any means, with major networking vendors from Broadcom to Cisco having implemented it in some capacity. As an open architecture, ACF appears to set a least common denominator so that customers deploying these systems know exactly what they’re getting.
However, it doesn’t stop there. Cornelis sees opportunities to accelerate a variety of other functions. We’ve detailed a few below:
KV cache offloading used to store and retrieve model state across multiple sessions.
MoE expert dispatch to reduce duplicate transfers and communication overhead for mixture-of-experts models during inference and training.
Message passing interface offload for HPC-centric applications.
In fabric checkpointing for failure recovery during training and other large workloads.
“Communications overhead and synchronization can leave expensive accelerators underutilized,” the company explained. These in-network accelerators “allow the network to operate on data as it moves through the system.”
According to Cornelis, the savings potential from reclaiming that idle compute is significant. “In a 100,000 GPU system, Cornelis modeling of public data shows that roughly half of all GPU hours are spent waiting for data, worth about $1.68 billion a year in wasted capacity and 500 GWh of power.”
As usual, take these claims with a grain of salt. However, reclaiming GPU idle time would equate to significant savings. Eliminating all the bottlenecks that contribute to them is easier said than done.
A new kind of NIC for the AI age
While Cornelis champions its ACF architecture, Delos Data, a startup founded by former Barefoot Networks and Intel execs, aims to tackle the physical layer with a series of new network reference designs.
Marketed under its Nonstop AI portfolio, the network interfaces span the full gamut of connectivity from co-packaged interconnects to more traditional NICs.
The idea, the company explains, is to provide customers with a system-level blueprint for speeding up data movement between endpoints and enabling larger compute domains scaling beyond a single rack.
Delos’ data interface will be offered in three form factors. The first is an I/O die capable of more than 30 Tbps of aggregate bandwidth or about 8 TB/s in either direction. That’s more than double the interconnect bandwidth of either Nvidia's or AMD’s latest accelerators, which cap out at 3.6 TB/s. Just how competitive that ends up being will depend on when Delos’ I/O chiplets actually see deployment.
That timeline is going to depend heavily on integration since those I/O dies need to be integrated directly into the accelerator package, which requires a high-degree of co-design.
Critically, Delos isn't trying to trap its customers in a walled garden to the same extent that Nvidia does with its NVLink Fusion I/O dies or IP. At least as of writing, NVLink Fusion still requires customers to buy NVSwitches for scale up networking.
Delos’ chiplets are protocol agnostic. They don’t care whether you’re using ESUN, UALink, or something else. Chip designers that'd rather focus their efforts and capital on the AI bits of their accelerators – which, it turns out, applies to most hyperscalers – could simply license Delos' chiplet design.
This protocol agnosticism also extends to Delos’ near-packaged optics (NPO) tech, which will integrate a 10-plus Tbps data interface — that’s 2.5 TB/s bidirectional bandwidth in Nvidia speak — with optics engines from leading optics suppliers.
Rack scale systems, like Nvidia’s NVL72, have largely relied on copper interconnects up to this point due to power constraints. As systems grow from 72 GPU rack systems to row-scale clusters with 576 and eventually 1,152 GPUs, optics become unavoidable.
While integrating optics directly into the accelerators is possible using tech available today, the blast radius of a failed optical module is considerable. NPO offers an alternative. Copper interconnects are used within the rack, with optics provided via user-serviceable NPO modules between racks.
Finally, Delos is working on a 400-plus Gbps NIC, which like the rest of its data interface offerings is protocol agnostic. This is arguably the least surprising entry in the entire lineup. Scale out networks are still going to be required for large training clusters as well as for front-end access and storage networks.
These interfaces build on Delos’ existing compute reference design, which we looked at during Computex, as well as its Nonstop AI software platform, which is designed to facilitate the configuration and monitoring of these switched fabrics or meshes in order to enable dynamic rerouting of traffic in the event of a link failure.
Telemetry gathered by its data interface offerings allows for even greater visibility into the network, allowing for faster rerouting and recovery.
Fueling competition
Investors appear more than eager for more competition in the scale up arena. Alongside its product announcements, Cornelis announced approximately $205 million in funding to bring its next generation of scale up and scale out networking products to market.
Meanwhile, Delos Data announced Tuesday that it’s now raised more than $100 million in funding with support from venture capital firms Matrix, Playground, and Socratic Partners among others. ®
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)