Analyze PyTorch CUDA memory snapshots from torch.cuda.memory._dump_snapshot(), torch.cuda.memory._snapshot(), or memory_viz artifacts. Use when the user wants to understand GPU memory usage, top allocation sites, deallocation timing, OOMs, allocator fragmentation, inactive cached memory, memory leaks, retained tensors, or Python reference cycles in PyTorch workloads. Supports diffing two snapshots from the same run to pinpoint growing allocation sites.