Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
Processing multi-hour video under tight edge and bandwidth budgets often forces a compromise between temporal continuity and fine-grained visual details. The CFD (Caption-once, Frames-on-Demand) framework resolves this through an edge-cloud division of labor. Edge devices perform a single offline captioning pass to create a dual-track narrative index of story skeletons and clip logs. A cloud MLLM then answers temporal-structural queries within text space, activating a lightweight Visual-Need Router to fetch keyframe pixels only when perceptual disambiguation is strictly required.
Why it's worth reading
It offers an actionable architectural compromise for edge-cloud long-video analysis by decoupling macro temporal structure into cached text while reserving raw pixel retrieval for targeted perceptual verification.