Light-Omni is a lightweight multimodal agent framework for continuous, long-horizon video understanding. It replaces iterative detective-style search and evidence aggregation with two coupled contextual states: a finite multimodal global script consolidated from episodic memory, and a parametric latent state that drives actions and retrieval embeddings in a single forward pass. The paper reports a 2.4% average accuracy gain over M3-Agent, a 12.1x speedup, and 2.6x better GPU memory efficiency. It also presents the framework as a memory system that can improve the efficiency and performance of existing multimodal large language models.
No heat snapshots are available in the last 24 hours.