This paper proposes Self-Guided Test-Time Training (S-TTT) for improving long-context utilization. Before adapting the model, S-TTT selects evidence spans relevant to the question, then applies the standard language-modeling objective only to those spans. The authors report that random-span TTT can reduce LongBench-v2 performance, while oracle-span adaptation improves it. On LongBench-v2 and LongBench-Pro, S-TTT improves both Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct, with gains of up to 15% relative. The abstract does not provide absolute scores, compute costs, or selection-error analyses.
No heat snapshots are available in the last 24 hours.