This paper introduces ParseFIxLIP, an extension of FIxLIP that integrates dependency-tree-based Tree-Gram Parsing into a weighted Banzhaf interaction framework. Its smart_depth strategy groups related spaCy tokens into semantically coherent explanation players, addressing fragmentation of clinical concepts such as “saddle embolus” and reducing the interaction space for long captions. The approach is qualitatively evaluated with BiomedCLIP on medical images from ROCOv2 and on general examples. The abstract reports improved semantic coherence and statistical robustness, but provides no concrete benchmark values, ablations, or detailed evaluation protocol.
No heat snapshots are available in the last 24 hours.