The paper introduces Atrex-Bench, a production-trace-driven benchmark covering 30 operators and 440 shapes sampled from compute-limited, memory-rich inference GPUs. Problems are weighted by observed GPU time and serving phase, with per-problem roofline ceilings. Across six frontier coding agents, the best vanilla model reaches only about 10% of the hardware roofline on production operators. The authors argue that correctness pass rates are misleading because many successes rely on PyTorch fallbacks. They also release Atrex-Kernel-Agent (AKA), combining profile-driven measure-revise search, optimization dropout, and a layered knowledge base to turn FlyDSL fallbacks into real kernels.
No heat snapshots are available in the last 24 hours.