The paper introduces Transfer-Aware Curriculum (TAC), an online bandit-style curriculum for multi-domain reinforcement learning with verifiable rewards (RLVR). TAC combines per-domain advantages, which estimate local learnability, with projected gradients from the ongoing GRPO update to estimate whether training on one domain will benefit others. The method reportedly adds less than 1% wall-clock overhead. On a six-domain reasoning suite covering mathematics, programming, and science, TAC achieves the best macro-averaged accuracy on Qwen3-1.7B and Llama3.2-3B, improving over a learnability-only bandit by up to 2.8 percentage points, or 10% relatively.
No heat snapshots are available in the last 24 hours.