This paper audits attribution faithfulness for two echocardiography video regressors fine-tuned on EchoNet-Dynamic: a self-supervised VideoMAE transformer with Chefer relevance propagation and a Kinetics-pretrained R(2+1)D model with Grad-CAM. Across the full 1,276-study test split, both models are strongly anatomically grounded, reaching 3.04x and 3.76x chance, respectively. Temporal reliance on the clinically decisive end-systolic and end-diastolic frames is much weaker: 1.05x chance for VideoMAE and 1.15x for R(2+1)D. The reported results indicate that temporal faithfulness is architecture-dependent and is not fixed by additional training.
No heat snapshots are available in the last 24 hours.