NVIDIA A100 and H100 bare metal for LLM inference and fine-tuning - fixed monthly pricing
About This Service
Running large language model inference at production scale requires GPU memory: LLaMA 3 70B in BF16 needs approximately 140GB of VRAM, which means two A100 80GB GPUs or a single H100 80GB running in quantised form. Hosting these workloads on a Glitch Servers GPU dedicated server costs a fraction of per-token cloud API pricing once your inference volume reaches meaningful scale. Our UK infrastructure at Telehouse North provides low-latency connectivity to UK API consumers, and our bare-metal model means your model weights stay on hardware you control - no co-tenanted inference serving, no data leaving your GPU environment.
GPU compute at Glitch is built on NVIDIA data centre hardware - A100, H100, and RTX configurations available for AI inference, model training, and rendering workloads. Unlike shared cloud GPU instances where VRAM and compute resources are time-sliced among tenants, dedicated GPU colocation and GPU server rental provide the full card allocation to a single workload, eliminating the throughput variability that degrades training runs on shared infrastructure.
The network operates through AS212868 with LINX and LONAP peering and NTT/GTT Tier-1 transit. High-bandwidth connectivity supports the data ingestion rates that large model training pipelines require, with 1Gbps ports with a 10Gbps burst upgrade available.
UK data residency is maintained throughout - training data, model weights, and inference outputs remain on UK infrastructure, supporting compliance requirements under UK GDPR and sector-specific data governance frameworks. Full GPU server and colocation specifications are on the GPU colocation page.
Technical Capabilities
High-density power, low-latency networking, and carrier-neutral connectivity for demanding GPU clusters.
GPUs: NVIDIA A100 80GB PCIe, H100 SXM5 80GB, L40S 48GB
Multi-GPU: NVLink available on A100 and H100 pairs
CUDA: 12.x, cuDNN 8.x, NCCL pre-installed on Ubuntu 22.04
Inference: vLLM, llama.cpp, TGI (Text Generation Inference) supported
Ideal Customers
We work with a range of organisations running compute-intensive AI workloads on owned hardware.
AI product teams running self-hosted LLM APIs
Researchers fine-tuning open-source models on proprietary datasets
Enterprises requiring UK data residency for AI workloads
Infrastructure
Our GPU servers are located in Telehouse North, London, with 100G uplinks via NTT and GTT. Intel Xeon host CPUs provide the I/O bandwidth needed to feed A100 and H100 GPUs at full PCIe 5.0 throughput.
FAQ
Deploy an LLM server from £299/mo
NVIDIA A100/H100. Full VRAM access. UK data residency.
Enquire Now →