SceneActBench introduces a unified agent-environment benchmark for evaluating whether vision-language model agents can act on complete, multi-object 3D scenes rather than merely describe them or manipulate a single object. Agents receive PNG images or sampled video frames and, where relevant, supplied 3D assets, then operate through one fixed loop. The benchmark contains five tasks built from 210 source instances, producing 520 task cases with paired input conditions. Hidden ground truth and task-specific geometric metrics evaluate final outputs. Across eleven proprietary VLM configurations, overall scores range from 38.6 to 50.2, with no configuration performing consistently across all tasks.
No heat snapshots are available in the last 24 hours.