Vinci2 reframes proactive assistance in continuous egocentric video as a context-dependent decision problem: an assistant must determine not only what is happening, but whether intervention is warranted. The paper introduces EgoServe, a benchmark with more than 3,000 service instances across four temporal memory horizons and 10 service categories. Its training-free EgoMemo agent combines multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives for retrieval-augmented reasoning. The authors report strong baseline performance on EgoServe and competitiveness on existing egocentric benchmarks.
No heat snapshots are available in the last 24 hours.