InfiniBand
A network built specifically for supercomputers. Lowest latency, and the default choice when GPUs in different servers must train one model together.
HDR 200Gb/s and NDR 400Gb/s switches, adapters, cables and optics.
Once you have more than one GPU server, the network decides whether extra accelerators actually make anything faster. This is the layer most first-time AI buyers under-specify.
A short answer, with no acronyms.
A single GPU server is self-contained: the accelerators inside it talk to each other over a direct internal link, and an ordinary office network is perfectly adequate for getting work in and out.
The moment a job is too large for one server, that changes. Training a large model across several machines means the GPUs exchange intermediate results constantly — thousands of times a second. On a normal network, every GPU spends most of its time waiting. You have bought eight accelerators and are getting the throughput of two.
An AI fabric is a second, dedicated network that exists purely for that traffic. It is faster, but more importantly it is predictable — no queueing behind somebody's file transfer. That predictability is what lets separate servers behave like one larger machine.
The practical rule. One server: you almost certainly do not need this. Two to four servers: worth costing. More than four, or any distributed training: budget for it from the start, because retrofitting a fabric into a live cluster is expensive and disruptive.
Every term used above is defined in the glossary, and Carl will explain any of it in whatever depth you want.
A network built specifically for supercomputers. Lowest latency, and the default choice when GPUs in different servers must train one model together.
HDR 200Gb/s and NDR 400Gb/s switches, adapters, cables and optics.
Ethernet engineered for AI traffic patterns, so your existing networking team can run it without learning a second technology.
NVIDIA Spectrum-X SN5600 and SN5610 up to 800GbE, SN2201 management switching.
The card inside each server that connects it to the fabric. On AI systems it does far more work than an ordinary network card.
ConnectX and BlueField adapters, OCP 3.0 and PCIe Gen5 form factors.
The boxes the servers plug into. Port count and speed decide how large the cluster can grow.
Leaf and spine switching for InfiniBand and Ethernet AI fabrics.
Used where AI traffic must cross between sites or datacentres rather than staying inside one room.
AI-native routing platforms for data-centre interconnect and distributed AI.
The unglamorous part that decides whether the cluster works on the day. Lengths, transceivers and topology all have to be right.
DAC, AOC and transceiver options matched to switch and adapter combinations.
Working out the topology before anything is ordered — how many switches, how many cables, and how the cluster grows later.
Creative Compute infrastructure design engagement.
Tell us how many servers and what you are training. We will size the switching, the adapters and the cabling, and tell you honestly if you do not need any of it yet.