Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

First seen · 9/11/2026, 01:53 AMLatest activity · 9/11/2026, 01:53 AM

Processing multi-hour video under tight edge and bandwidth budgets often forces a compromise between temporal continuity and fine-grained visual details. The CFD (Caption-once, Frames-on-Demand) framework resolves this through an edge-cloud division of labor. Edge devices perform a single offline captioning pass to create a dual-track narrative index of story skeletons and clip logs. A cloud MLLM then answers temporal-structural queries within text space, activating a lightweight Visual-Need Router to fetch keyframe pixels only when perceptual disambiguation is strictly required.

Event heat · last 24 hours

There are 8 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 11:00; latest heat is 0.

There are 8 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 11:00; latest heat is 0.10.509/12, 11:00, event heat 09/12, 14:00, event heat 09/12, 17:00, event heat 09/12, 20:00, event heat 09/12, 23:00, event heat 09/13, 02:00, event heat 09/13, 05:00, event heat 09/13, 08:00, event heat 024 hours agoNow
  1. 9/12, 11:00, event heat 0
  2. 9/12, 14:00, event heat 0
  3. 9/12, 17:00, event heat 0
  4. 9/12, 20:00, event heat 0
  5. 9/12, 23:00, event heat 0
  6. 9/13, 02:00, event heat 0
  7. 9/13, 05:00, event heat 0
  8. 9/13, 08:00, event heat 0

Reporting Timeline

  1. AggregatorarXiv9/11, 01:53 AMnot independentRepresentative
    Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding