This paper identifies repetitive copying as a widespread failure mode in long-context reasoning: models copy large portions of the prompt into their reasoning traces instead of solving the task, and the behavior worsens as context length increases. The authors attribute the problem to insufficient grounding in task-relevant evidence. They propose GEAR, a reward-shaping method that rewards overlap with key evidence and penalizes overlap with distractor context. Across multiple model scales and benchmarks, GEAR reportedly improves average performance by up to 4.6 points over standard accuracy-based reinforcement learning, while reducing copying and reasoning length.
No heat snapshots are available in the last 24 hours.