PixelEyes targets a failure mode in multi-turn visual reasoning: inaccurate localization causes MLLM agents to revisit wrong crops, producing long and redundant trajectories. The system separates the reasoner’s decision about what evidence to seek from a specialized referring-segmentation tool that determines where it is. It adds semantic-region breadth-first search to reduce repeated exploration and trains on PixelEyes-6K, a dataset synthesized from expert trajectories. The authors also introduce Pinpoint-Bench, a zero-hint benchmark with instance masks and boxes for separating localization failures from reasoning failures. The abstract states that code and models are open-sourced, but provides no quantitative results.
No heat snapshots are available in the last 24 hours.