While spatial evaluations for foundation models typically measure metric properties like distance and bounding boxes, real-world understanding relies heavily on topological relations invariant to continuous deformation. MindTopo introduces a benchmark of 11,030 procedurally generated instances testing five core topological concepts—continuity, separation, order, enclosure, and knots—across static reasoning and closed-loop agent planning. Evaluating 14 multimodal models revealed a uniform shortfall: models consistently perform worse at planning than reasoning, lagging significantly behind human baselines. Even when paired with video generation, rollouts struggle to preserve physical dynamics and topological integrity across transitions.
There are 7 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 17:00; latest heat is 0.