This paper introduces responsibility distribution estimation for ego-view traffic accident videos, asking models to predict the percentage of responsibility assigned to each involved agent. The authors build an LLM-assisted annotation pipeline and fine-tune multimodal large language models under several input conditions: raw video frames, segmentation-enhanced inputs, and textual descriptions. The study presents an initial benchmark for reasoning about accident avoidability and responsibility from the driver’s visual perspective, extending traffic-video understanding beyond classification and narrative explanation.
No heat snapshots are available in the last 24 hours.