MedStreamBench (arXiv:2607.01751) combines 22 medical datasets and 5,419 QA instances to evaluate time-aware video understanding across retrospective, present, future, and proactive settings. Models receive only temporally bounded evidence rather than unrestricted full-video access. The benchmark supports both single-turn and streaming evaluation, and adds responsiveness and post-evidence stability metrics. Experiments with general-purpose and medical vision-language models report a substantial gap between offline recognition and temporally grounded decision-making, with marked degradation in streaming and proactive scenarios.
No heat snapshots are available in the last 24 hours.