GPU-CFR: Accelerating Counterfactual Regret Minimization up to 80x via Static Dataflow and CUDA Graphs
First seen · 9/11/2026, 01:58 AMLatest activity · 9/11/2026, 01:58 AM
Counterfactual Regret Minimization (CFR) has long run faster on CPUs because millions of tiny gather and scatter kernel dispatches dominate GPU runtimes. GPU-CFR circumvents this bottleneck by observing that game tree topology is completely fixed ahead of solving. By compiling the game into a static dataflow of flat arrays and depth-batched passes executed through CUDA Graph Replay, it reduces framework operations by up to 18.1x. Evaluated on a single A100 across eight games, it outperforms prior GPU solvers by 29.8–80.4x and surpasses leading CPU frameworks by up to 258x without altering numerical update rules.
Event heat · last 24 hours
There are 8 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 17:00; latest heat is 0.