How We Reduced PyTorch Model Inference Time by 40% on the Glasgow Uni GPU Cluster

Around 11 AM on Saturday, I typed nvidia-smi into the terminal and sighed. A computer vision model was chewing through about 210ms per image for inference. The entire batch pipeline was crawling because of this bottleneck, and allocation time on the University of Glasgow GPU cluster is finite. If I couldn't meaningfully reduce that by Sunday evening, I'd have to book another slot the following week.

Long story short, I got it down to about 125ms. A reduction of just over 40%. But getting there involved burning the entire Saturday afternoon on a dead end.

Profile First, Optimise Second

My first instinct was "just slap on quantisation and it'll speed up." Classic mistake. I should have run the PyTorch Profiler first.

When I captured a trace with torch.profiler, a significant chunk of the total inference time was being eaten in unexpected places. The main culprits were the NMS (Non-Maximum Suppression) operation in post-processing and unnecessarily frequent tensor copies in the model's intermediate layers. It wasn't the GPU computation itself — data movement was the real bottleneck.

If you open the profiling results in the Chrome trace viewer, you can see short gaps appearing repeatedly between GPU kernels that should be packed tightly on the timeline. That's CPU-GPU synchronisation overhead.

GPU profiling timeline visualization showing kernel execution gaps caused by CPU-GPU synchronization overhead

Quantisation, and the Saturday Afternoon Dead End

After profiling, I applied dynamic quantisation. I used torch.quantization.quantize_dynamic to convert the Linear layers to INT8, and there was almost no drop in accuracy. So far, so good.

The problem was getting greedy. I tried to push further with static quantisation — prepared a calibration dataset, set up the prepare → convert pipeline — but the model's custom attention blocks clashed with the quantisation backend, and errors came flooding out. After wrestling with it for over three hours, I accepted it simply couldn't be cleanly applied to this model architecture without separate modifications. Saturday afternoon, completely gone.

Dynamic quantisation alone reduced inference time by about 15%. I let go of the ambition and moved on.

Operator Fusion Was More Effective Than Expected

I converted the model to TorchScript using torch.jit.trace, then applied Conv-BatchNorm fusion. Calling torch.jit.optimize_for_inference handles this kind of fusion automatically under the hood.

This step shaved off an additional 10–12% or so. When I checked with the profiler again, the gaps between kernels I'd seen before were noticeably smaller.

The rest came from adding up smaller optimisations:

  • Applied mixed precision inference by combining torch.no_grad() with torch.cuda.amp.autocast() during inference
  • Swapped the torchvision built-in NMS for a batched version processed directly on the GPU
  • Removed unnecessary .cpu() calls — the direct cause of the tensor copy issue caught during profiling

Before and after comparison of model inference latency showing 40 percent reduction

All of that combined brought it to about 125ms per image. I finished on Sunday evening, roughly two hours before the cluster allocation ran out. After getting home, the IPA I had at a pub over in Finnieston tasted impossibly good.

One thing worth remembering. Run the profiler first. The reason I wasted Saturday afternoon was that I started optimising before measuring. Fight the urge to reach for fancy techniques, and look at the numbers first. Numbers don't lie — though interpreting them is on you.

Comments