The paper introduces Counterfactual Vision Action Analysis (CVAA), which removes detected objects from front-camera images through photorealistic generative inpainting and measures changes in an autonomous driving model’s planning behavior. Applied to the Alpamayo 1 trajectory predictor across 210 nuScenes scenes, the authors create Counter-nuScenes. Vehicles and pedestrians within the predicted path show dominant causal influence, while traffic lights have disproportionate impact relative to their image area. The study also reports strong responses to objects humans may consider irrelevant, and compares intermediate representations across model layers to investigate whether these effects reflect interpretable scene objects or other internal features.
No heat snapshots are available in the last 24 hours.