SWE-Touch evaluates coding agents in shared workspaces where users inspect or modify task-relevant code during an ongoing task. The framework introduces validated Counter-Edits: plausible user changes that conflict with completing the target task. A separate User Patch Generator creates these edits, which are injected with contextual messages when agents reach relevant code. Across nine coding models on SWE-bench Verified, Counter-Edits reduce average resolve rate by 7.7 percentage points. Similar degradation appears on longer-horizon tasks from SWE-Bench Pro and DeepSWE. Trajectory analysis attributes failures to weak awareness of evolving workspace state, insufficient reconciliation, and inadequate targeted testing after changes.
No heat snapshots are available in the last 24 hours.