8 PyTorch Profiling Tools I Actually Pull Out When My Training Loop Slows Down
3 AM. A model was clocking 47 minutes per epoch on an A100. A paper running the same architecture reported 12 minutes. When GPU time feels as painful to waste as Gangnam rent, reaching for a profiler isn't a choice—it's survival.
Over the past two years training various models in Daegu, I've narrowed things down to the tools I actually use repeatedly. I'll be honest about when each one helps and when it's a waste of time.
torch.profiler (+ TensorBoard Plugin)
This is the first thing I reach for. Wrap your code with the torch.profiler.profile context manager, specify warmup and active steps with schedule, then load the trace in TensorBoard—you get kernel-level timing. GPU kernel names, per-CUDA-stream timelines, and memory allocation trends all visible at a glance.
The limitations are clear, though. Profiling itself introduces overhead, so measurements can be distorted in experiments with small batch sizes. The TensorBoard plugin also tends to freeze the browser when loading large traces. I usually capture just 5–10 steps and cut it off.
You can also export JSON via export_chrome_trace from the PyTorch Profiler and view it in Chrome's chrome://tracing, but when there are many kernels, TensorBoard actually works better.
py-spy
I use this when I suspect a bottleneck at the Python level—to quickly check whether data loader preprocessing is slow or pure Python logic is holding things back. A single py-spy top --pid <PID> shows in real time which functions are consuming time.
It can't see anything on the GPU side, so it's useless for CUDA kernel optimization. But it can distinguish a CPU preprocessing problem from a GPU problem in 30 seconds, so I often pull it out as a first diagnostic step.

NVIDIA Nsight Systems
I use this when torch.profiler has given me a rough idea of where things are slow but I need to dig deeper at the kernel level. The timeline view is overwhelmingly detailed. CPU-GPU synchronization points, per-CUDA-stream kernel scheduling, and PCIe data transfer segments are all visualized.
The downside is the steep learning curve. The GUI is feature-rich but not intuitive, and the first time you use it, you may spend more time learning the tool than actually profiling. I lost half a day on my first attempt. Now I run nsys profile from the command line to generate an .nsys-rep file first, then open only the parts I need in the GUI.
Manual Timing with torch.cuda.Event
I use this when a full profiler is overkill and I just want to quickly measure a specific section. Create start/end markers with torch.cuda.Event(enable_timing=True) and call elapsed_time. You can isolate how many milliseconds each of the forward pass, backward pass, and optimizer step takes.
One thing to watch out for: if you forget torch.cuda.synchronize(), the numbers come out completely wrong due to asynchronous execution. I once forgot this and got an absurd result of "forward only takes 0.3 ms," then wasted half a day chasing it.
Memory Tools
A large portion of training speed issues are actually memory issues. Reduce batch size due to OOM and throughput drops. Let unnecessary tensors linger in memory and fragmentation kicks in.
torch.cuda.memory_stats / memory_summary
Calling torch.cuda.memory_summary() shows currently allocated memory, cached memory, and peak memory in a table. Print this right before an OOM and you can get a sense of where memory is leaking. I log this summary at the end of every epoch in my training scripts.
This alone won't tell you which tensors are eating the memory, though. That's why I use it alongside the next tool.
torch.cuda.memory._record_memory_history
A relatively recent addition that records stack traces for tensor allocations and deallocations. You can track which line of code allocated a tensor of what size and when it was freed. Decisive for catching memory leaks.
I once used this to catch a bug where a custom loss function was accumulating intermediate tensors in a list without .detach(), causing memory to grow linearly every epoch. Without it, I would have been stuck for a long time.

Data Pipeline Diagnostics
torch.utils.bottleneck
Pass a script path as an argument and it runs cProfile and CUDA profiling simultaneously, then outputs a summary. Lowest barrier to entry when you just want a rough idea of where things are slow.
The output is coarse, though, and in complex training loops there's often so much noise that it's hard to extract anything meaningful. I usually run it once when first setting up a new project, then move on to more precise tools. Fine for initial screening.
DataLoader num_workers Tuning + Simple Time Measurements
More of a habit than a tool, but simply increasing num_workers from 0 and measuring batch loading time is enough to quickly identify data pipeline bottlenecks. Time next(iter(dataloader)) with time.perf_counter. I also experiment with combinations of prefetch_factor and persistent_workers.
When using a local NVMe SSD in our Daegu office versus mounting remote storage, the optimal num_workers value was completely different. Whenever the environment changes, I redo this from scratch.
I never use all eight of these at once. Typically I start with py-spy or torch.utils.bottleneck to get a rough location, narrow it down with torch.profiler, and go to Nsight Systems if needed. If a memory issue is suspected, I start with memory_summary. Having a set order noticeably cuts down on wasted time.
In the next post, I plan to document the numerical instability issues I ran into during mixed precision conversion.
Comments