MindTopo: Can Foundation Models Reason in Topological Space?
While spatial evaluations for foundation models typically measure metric properties like distance and bounding boxes, real-world understanding relies heavily on topological relations invariant to continuous deformation. MindTopo introduces a benchmark of 11,030 procedurally generated instances testing five core topological concepts—continuity, separation, order, enclosure, and knots—across static reasoning and closed-loop agent planning. Evaluating 14 multimodal models revealed a uniform shortfall: models consistently perform worse at planning than reasoning, lagging significantly behind human baselines. Even when paired with video generation, rollouts struggle to preserve physical dynamics and topological integrity across transitions.
Why it's worth reading
Highlights a fundamental blind spot of multimodal models beyond metric geometry, demonstrating that topological invariance and closed-loop planning remain deep bottlenecks for visual agents.