Solutions

LLM infrastructure

Sizing, hardware and networking for training, fine-tuning and serving large language models — specified around your model, context length and user count.

Start with the model, not the GPU

The model you intend to run decides memory, and memory decides the GPU. As a guide, allow roughly 2 GB of GPU memory per billion parameters at 16-bit precision, plus space for context.

  • 8B model: 24–48 GB
  • 70B model: 141–192 GB
  • Mixture-of-experts: memory heavy

Training versus inference

Training needs interconnect bandwidth between GPUs. Inference needs memory capacity and memory bandwidth. Buying the wrong one is the most common and most expensive mistake.

Networking matters

Multi-node training is limited by the fabric. We specify InfiniBand or Ethernet fabrics appropriate to the cluster size.

  • 400 Gb InfiniBand
  • Spectrum-X Ethernet
  • Storage network sizing

Buy, rent or both

Rent while you experiment, buy once utilisation is steady. We can model the crossover point for your workload.

Reference configurations

Departmental inference, team fine-tuning, and multi-node training reference builds, all quotable as complete projects.

Software stack

Container runtime, schedulers, inference servers and monitoring, configured before delivery.

  • Kubernetes or Slurm
  • vLLM or TensorRT-LLM
  • Prometheus and Grafana

Related products

Shop everything
GPU Servers
Representative

Creative Compute Rack — 8 × H200 GPU Server

A large training server. It needs three-phase power and proper cooling, so it belongs in a data centre or a purpose-built room.

Memory

Configurable

Deployment

Rack server

Multi-GPUProduction Inference

Price on request

Available to order

Refurbished AI Infrastructure
Representative

Refurbished 8 × A100 GPU Server

A previously deployed training server, fully tested and warranted. A cost-effective way to obtain serious training capacity.

Memory

Configurable

Deployment

Rack server

RefurbishedMulti-GPUProduction Inference

£58,500

ex VAT

Low stock

AI Clusters and Rack-Scale
Representative

NVIDIA GB300 NVL72 Rack System

An entire rack delivered as one computer. It is bought for frontier model training and very large inference platforms, and requires liquid cooling and high-capacity power.

Memory

20736 GB

Deployment

Datacentre

High-MemoryMulti-GPU

Price on request

Allocation only