Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
arXiv·Weitong Cai·Sep 10, 2026, 5:53 PM

Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

Papers76

Processing multi-hour video under tight edge and bandwidth budgets often forces a compromise between temporal continuity and fine-grained visual details. The CFD (Caption-once, Frames-on-Demand) framework resolves this through an edge-cloud division of labor. Edge devices perform a single offline captioning pass to create a dual-track narrative index of story skeletons and clip logs. A cloud MLLM then answers temporal-structural queries within text space, activating a lightweight Visual-Need Router to fetch keyframe pixels only when perceptual disambiguation is strictly required.

Why it's worth reading

It offers an actionable architectural compromise for edge-cloud long-video analysis by decoupling macro temporal structure into cached text while reserving raw pixel retrieval for targeted perceptual verification.

Tags

Long Video UnderstandingMLLMEdge AIVideo CaptioningAgentic FrameworkComputer VisionEfficiency

Score breakdown

  • Novelty76
  • Impact74
  • Practicality82
  • Credibility75
  • Timeliness76