Salesforce AI Research introduces Evidence-Backed Video Question Answering (E-VQA), a task that requires Video LLMs to produce both semantic answers and precise spatio-temporal evidence: temporal segments plus dense tracked object segmentation masklets. The paper presents ST-Evidence, described as the first human-verified benchmark covering discriminative and generative pixel-level grounding, and ST-Evidence-Instruct, a 160k-scale automatically generated training set. Fine-tuning grounded Video LLMs reportedly improves a 7B model over a size-matched UniPixel baseline by 27.2 t-mean and 13.8 J&F.
No heat snapshots are available in the last 24 hours.