GROVE is a training-free framework for assistants that continuously observe video and need to answer questions about the past while also recognizing when past experience is relevant to the present. It grows one causal memory from streaming input, preserving fine-grained perceptual evidence and progressively consolidating it into timestamped moments, coherent episodes, and recurring patterns across days. Each temporal stratum has a scale-specific retrieval skill. The authors report that GROVE achieves the best results among compared methods on benchmarks including MM-lifelong and EgoServe, while ablations indicate that the strata and their access skills are complementary.
No heat snapshots are available in the last 24 hours.