SafeRelBench evaluates process-level safety in VLM-driven embodied agents by testing spatial relations such as support, containment, and proximity before risk-prone actions. The benchmark contains 507 executable samples: 248 spatial-relation cases and 259 non-spatial controls. Evaluations across seven open- and closed-source VLM agents reportedly reveal a substantial gap between task success and safety compliance, showing that an agent can complete a requested task while violating constraints during execution. The paper argues that embodied safety requires reasoning about how object relations evolve during interaction, not only recognizing static hazards or checking the final state.
No heat snapshots are available in the last 24 hours.