The paper presents STEEL, an open-source FlashAttention implementation for XDNA-like NPUs. Its dataflow formulation maps prefill attention onto spatially parallel NPU resources and on-chip memory, while sparsity-aware pipeline placement addresses load imbalance caused by causal masking. On an AMD Ryzen AI 9 HX 370 SoC, the authors report 9.17x lower energy consumption than an optimized CPU baseline and 1.75x lower consumption than a GPU baseline on average. STEEL reportedly reduces latency by 9.6x over prior work on XDNA 1 and achieves 22.8x average speedup over layer-by-layer attention on XDNA 2.
No heat snapshots are available in the last 24 hours.