VIABench is a video benchmark for evaluating multimodal large language models in real-world visual impairment assistance. It uses first-person videos recorded or shared by visually impaired individuals and covers three tasks: Proactive Reminder, Visual Question Answering, and Vision-Guided Interaction. The benchmark supports both online real-time and offline evaluation. According to the supplied abstract, current MLLMs remain weak at comprehensive assistance, with Proactive Reminder being particularly difficult because it requires anticipating navigation-critical events and responding promptly. The authors say the code and data will be released.
No heat snapshots are available in the last 24 hours.