The paper introduces TokAG, a zero-shot affordance grounding framework that localizes image regions relevant to a specified action without external supervision. It uses token-level semantic-spatial signals from large vision-language models and proposes spatial-aware token selection to identify output tokens whose attention maps focus on the target object rather than background regions. The resulting attention maps are converted into affordance heatmaps. According to the abstract, TokAG consistently outperforms prior weakly supervised methods, improving NSS by 10.7% on the unseen split of AGD20K and by 29.7% on HICO-IIF.
No heat snapshots are available in the last 24 hours.