FocusMem proposes a compact latent-memory interface for GUI agents that separates three functions: what evidence is stored, what view is exposed for the current decision, and whether that memory should be trusted. A role-aware content basis targets reusable episodic experience and unfinished working progress. A state-conditioned readout produces decision-specific views, while a lightweight trust gate can suppress seemingly irrelevant retrieved trajectories. The authors report consistent gains over a matched action-only fixed-memory baseline and prior latent-memory adaptations across five GUI-agent benchmarks, with the policy frozen during training. Analyses examine complementary semantic and functional supervision, context-length robustness, and resistance to irrelevant episodic evidence.
No heat snapshots are available in the last 24 hours.