SeerGuard is a consequence-aware safety framework for mobile GUI agents. It combines instruction-level screening before task execution with action-level risk assessment before individual GUI actions are carried out. The framework uses multitask learning to build a safety-augmented world model (SAWM), jointly predicting semantic next states and safety risks. On Qwen3-VL-8B-Instruct, the reported safety-utility score rises from 0.191 to 0.596 at ω=0.8, while the risk-cost score falls from 0.347 to 0.130 at α=0.8. The paper also reports generalization across diverse mobile GUI agents.
No heat snapshots are available in the last 24 hours.