This paper introduces visual prompt engineering (VIPE), an approach that automatically modifies a task image before it is given to a video model. For example, an abstract scene in a visual physics problem can be transformed into a photorealistic image using an image-editing model. The authors report that VIPE improves video reasoning across tasks and can be more effective than conventional text prompt engineering or test-time scaling. They frame it as a simple, compute-efficient method for eliciting stronger visual reasoning, but the supplied abstract does not provide task-level scores, model names, or detailed experimental settings.
No heat snapshots are available in the last 24 hours.