Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
arXiv·Zhenxuan Fan·Sep 4, 2026, 4:19 PM

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

Papers78

Most current evaluations of Vision-Language-Action (VLA) models rely on simple task-completion metrics under constrained settings. RoboSPA introduces a diagnostic benchmark designed to examine where embodied reasoning falters as spatial ambiguity and procedural depth scale up. Spanning 10 task categories, 56 base tasks across five difficulty tiers, and 527,000 trajectories across diverse embodiments, the benchmark moves beyond binary success rates. Evaluations across representative VLA architectures demonstrate that existing models still struggle significantly with intricate spatial relations, fine motor execution, and memory-intensive long-horizon planning.

Why it's worth reading

It shifts VLA evaluation from coarse success rates to granular diagnostic metrics, systematically probing where robotic spatial reasoning and multi-step execution fail.

Tags

RoboticsVLAEmbodied AIBenchmarkManipulationDatasetComputer Vision

Score breakdown

  • Novelty76
  • Impact78
  • Practicality82
  • Credibility80
  • Timeliness75