AI & GPU Colocation UK

Large Language Model Hosting in the UK

NVIDIA A100 and H100 bare metal for LLM inference and fine-tuning - fixed monthly pricing

NVIDIA A100 GPU server for LLM hosting in UK data centre

About This Service

Large Language Model Hosting in the UK

Running large language model inference at production scale requires GPU memory: LLaMA 3 70B in BF16 needs approximately 140GB of VRAM, which means two A100 80GB GPUs or a single H100 80GB running in quantised form. Hosting these workloads on a Glitch Servers GPU dedicated server costs a fraction of per-token cloud API pricing once your inference volume reaches meaningful scale. Our UK infrastructure at Telehouse North provides low-latency connectivity to UK API consumers, and our bare-metal model means your model weights stay on hardware you control - no co-tenanted inference serving, no data leaving your GPU environment.

GPU compute at Glitch is built on NVIDIA data centre hardware - A100, H100, and RTX configurations available for AI inference, model training, and rendering workloads. Unlike shared cloud GPU instances where VRAM and compute resources are time-sliced among tenants, dedicated GPU colocation and GPU server rental provide the full card allocation to a single workload, eliminating the throughput variability that degrades training runs on shared infrastructure.

The network operates through AS212868 with LINX and LONAP peering and NTT/GTT Tier-1 transit. High-bandwidth connectivity supports the data ingestion rates that large model training pipelines require, with 1Gbps ports with a 10Gbps burst upgrade available.

UK data residency is maintained throughout - training data, model weights, and inference outputs remain on UK infrastructure, supporting compliance requirements under UK GDPR and sector-specific data governance frameworks. Full GPU server and colocation specifications are on the GPU colocation page.

External resources: NVIDIA Data Centre and the UK AI Safety Institute.

Technical Capabilities

Infrastructure Built for AI Workloads

High-density power, low-latency networking, and carrier-neutral connectivity for demanding GPU clusters.

Power & Cooling

GPUs: NVIDIA A100 80GB PCIe, H100 SXM5 80GB, L40S 48GB

Network Connectivity

Multi-GPU: NVLink available on A100 and H100 pairs

GPU Compatibility

CUDA: 12.x, cuDNN 8.x, NCCL pre-installed on Ubuntu 22.04

Security & Compliance

Inference: vLLM, llama.cpp, TGI (Text Generation Inference) supported

Ideal Customers

Who Uses Our AI/GPU Colocation

We work with a range of organisations running compute-intensive AI workloads on owned hardware.

🔬

Research & Academia

AI product teams running self-hosted LLM APIs

🏢

Enterprise AI Teams

Researchers fine-tuning open-source models on proprietary datasets

🚀

AI Startups

Enterprises requiring UK data residency for AI workloads

Infrastructure

Our Data Centre Capabilities

Our GPU servers are located in Telehouse North, London, with 100G uplinks via NTT and GTT. Intel Xeon host CPUs provide the I/O bandwidth needed to feed A100 and H100 GPUs at full PCIe 5.0 throughput.

LLM inference running on GPU server

FAQ

Frequently Asked Questions

We offer configurations with NVIDIA A100 80GB PCIe (single and dual), NVIDIA H100 SXM5 (single), and NVIDIA L40S 48GB. For models requiring more than 80GB VRAM, dual-A100 NVLink configurations provide 160GB pooled VRAM.
We offer Ubuntu 22.04 with CUDA 12.x, PyTorch 2.x, vLLM, Hugging Face Transformers, and llama.cpp pre-installed as an optional deployment template.
Yes. Fine-tuning with LoRA/QLoRA via PEFT, Axolotl, or LLaMA-Factory works on our A100/H100 configurations. Full GPU access - no shared GPU quota.

Deploy an LLM server from £299/mo

NVIDIA A100/H100. Full VRAM access. UK data residency.

Enquire Now →