This paper studies warp divergence across NVIDIA Pascal, Ampere, Hopper, and datacenter and consumer Blackwell GPUs. Using cycle-accurate microbenchmarks, hardware counters, and static analysis of compiler-generated SASS, it reports that divergent paths serialize approximately linearly with the number of paths k, while warp execution efficiency falls toward 32/k. The observed cost is independent of occupancy and predication removes it. Pascal already shows the same programmer-visible cost despite predating Independent Thread Scheduling. In contrast, reconvergence machinery changes substantially, including barrier-register instructions, uniform branches, and partial-mask synchronization on Blackwell.
No heat snapshots are available in the last 24 hours.