GPU-CFR: Accelerating Counterfactual Regret Minimization up to 80x via Static Dataflow and CUDA Graphs
Original title:GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay
Counterfactual Regret Minimization (CFR) has long run faster on CPUs because millions of tiny gather and scatter kernel dispatches dominate GPU runtimes. GPU-CFR circumvents this bottleneck by observing that game tree topology is completely fixed ahead of solving. By compiling the game into a static dataflow of flat arrays and depth-batched passes executed through CUDA Graph Replay, it reduces framework operations by up to 18.1x. Evaluated on a single A100 across eight games, it outperforms prior GPU solvers by 29.8–80.4x and surpasses leading CPU frameworks by up to 258x without altering numerical update rules.
Why it's worth reading
It resolves the long-standing dispatch bottleneck of tree-based CFR on GPUs through static dataflow compilation and graph replay, achieving orders-of-magnitude speedups without modifying the solver's mathematical formulation.