The paper introduces SIS-Bench, a benchmark for embodied spatial intelligence in UAVs formulated around “self in space.” It evaluates perception, memory, and reasoning across 13 tasks, using 4,856 question-answer pairs derived from 1,646 real-world UAV videos and expert verification. Evaluations report that current multimodal large language models struggle with dynamic, agent-centered processes, showing an imbalance between spatial cognition and self-awareness and degradation across cognitive levels. The authors also investigate a motion-aware representation that fuses optical flow with visual features. According to the paper, modeling self-related motion consistently improves perception and memory, including self-awareness, and transfers to downstream UAV decision-making tasks.
No heat snapshots are available in the last 24 hours.