Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

SceneActBench: Can Agents Act on the 3D Scenes They See?

First seen · 7/27/2026, 12:00 PMLatest activity · 7/27/2026, 12:00 PM

SceneActBench introduces a unified agent-environment benchmark for evaluating whether vision-language model agents can act on complete, multi-object 3D scenes rather than merely describe them or manipulate a single object. Agents receive PNG images or sampled video frames and, where relevant, supplied 3D assets, then operate through one fixed loop. The benchmark contains five tasks built from 210 source instances, producing 520 task cases with paired input conditions. Hidden ground truth and task-specific geometric metrics evaluate final outputs. Across eleven proprietary VLM configurations, overall scores range from 38.6 to 50.2, with no configuration performing consistently across all tasks.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/27, 12:00 PMnot independentRepresentative
    SceneActBench: Can Agents Act on the 3D Scenes They See?