Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
First seen · 9/11/2026, 01:53 AMLatest activity · 9/11/2026, 01:53 AM
Processing multi-hour video under tight edge and bandwidth budgets often forces a compromise between temporal continuity and fine-grained visual details. The CFD (Caption-once, Frames-on-Demand) framework resolves this through an edge-cloud division of labor. Edge devices perform a single offline captioning pass to create a dual-track narrative index of story skeletons and clip logs. A cloud MLLM then answers temporal-structural queries within text space, activating a lightweight Visual-Need Router to fetch keyframe pixels only when perceptual disambiguation is strictly required.
Event heat · last 24 hours
There are 8 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 11:00; latest heat is 0.