Most current evaluations of Vision-Language-Action (VLA) models rely on simple task-completion metrics under constrained settings. RoboSPA introduces a diagnostic benchmark designed to examine where embodied reasoning falters as spatial ambiguity and procedural depth scale up. Spanning 10 task categories, 56 base tasks across five difficulty tiers, and 527,000 trajectories across diverse embodiments, the benchmark moves beyond binary success rates. Evaluations across representative VLA architectures demonstrate that existing models still struggle significantly with intricate spatial relations, fine motor execution, and memory-intensive long-horizon planning.
No heat snapshots are available in the last 24 hours.