AI infrastructure

AI networking

Once you have more than one GPU server, the network decides whether extra accelerators actually make anything faster. This is the layer most first-time AI buyers under-specify.

Request a network designAI storage
Start here

Why does AI need specialist networking?

A short answer, with no acronyms.

A single GPU server is self-contained: the accelerators inside it talk to each other over a direct internal link, and an ordinary office network is perfectly adequate for getting work in and out.

The moment a job is too large for one server, that changes. Training a large model across several machines means the GPUs exchange intermediate results constantly — thousands of times a second. On a normal network, every GPU spends most of its time waiting. You have bought eight accelerators and are getting the throughput of two.

An AI fabric is a second, dedicated network that exists purely for that traffic. It is faster, but more importantly it is predictable — no queueing behind somebody's file transfer. That predictability is what lets separate servers behave like one larger machine.

The practical rule. One server: you almost certainly do not need this. Two to four servers: worth costing. More than four, or any distributed training: budget for it from the start, because retrofitting a fabric into a live cluster is expensive and disruptive.

  1. GPU serverAccelerators
  2. AI networkDedicated fabric
  3. GPU serverAccelerators
  4. StorageTraining data
Every GPU server connects to the same dedicated fabric, and so does storage. The fabric is what makes the separate boxes usable as one pool of capacity.

Every term used above is defined in the glossary, and Carl will explain any of it in whatever depth you want.

Categories

What we source

Representative system image

InfiniBand

A network built specifically for supercomputers. Lowest latency, and the default choice when GPUs in different servers must train one model together.

HDR 200Gb/s and NDR 400Gb/s switches, adapters, cables and optics.

Representative system image

Ethernet AI fabrics

Ethernet engineered for AI traffic patterns, so your existing networking team can run it without learning a second technology.

NVIDIA Spectrum-X SN5600 and SN5610 up to 800GbE, SN2201 management switching.

Representative system image

SuperNICs and adapters

The card inside each server that connects it to the fabric. On AI systems it does far more work than an ordinary network card.

ConnectX and BlueField adapters, OCP 3.0 and PCIe Gen5 form factors.

Representative system image

Switches

The boxes the servers plug into. Port count and speed decide how large the cluster can grow.

Leaf and spine switching for InfiniBand and Ethernet AI fabrics.

Representative system image

Routers

Used where AI traffic must cross between sites or datacentres rather than staying inside one room.

AI-native routing platforms for data-centre interconnect and distributed AI.

Representative system image

Cabling and optics

The unglamorous part that decides whether the cluster works on the day. Lengths, transceivers and topology all have to be right.

DAC, AOC and transceiver options matched to switch and adapter combinations.

Representative system image

Network design

Working out the topology before anything is ordered — how many switches, how many cables, and how the cluster grows later.

Creative Compute infrastructure design engagement.

Not sure what fabric your cluster needs?

Tell us how many servers and what you are training. We will size the switching, the adapters and the cabling, and tell you honestly if you do not need any of it yet.